Your curated collection of saved posts and media

Showing 32 posts Β· last 7 days Β· newest first
πŸ”ai_fast_track retweeted
0
0xSero
@0xSero
πŸ“…
May 02, 2026
85d ago
πŸ†”98372588

Weekly best models for your hardware: ~~ 8 to 16gb ~~ Granite models are amazing: [NEW] - https://t.co/uVSUl1scUM Gemma-E4B is a good general QA model - https://t.co/y2xrZKtjWe Qwen3.5-9B is the best at this level imo - https://t.co/WNwxTttMQC ~~ 16 to 64gb ~~ Another larger Granite: This is a general chat model, really dense with world knowledge. [NEW] - https://t.co/aM9Lld3axk - Undisputed kings: The Qwens at various precisions: (Higher ceiling) - https://t.co/oZIWk2iZxA - https://t.co/oZIWk2iZxA The Gemmas at various precisions: (More efficient) - https://t.co/7E9nKmDwGo - https://t.co/fdt9jntwHG ~~ 64 to 128gb ~~ - Ling is a new 100B~ contender decent agent [NEW] https://t.co/FuwQ571a2R - Mistral medium: from my experience their models have been the most consistent! [NEW] https://t.co/GX7Z4zog1y ~~ 128gb - 256gb ~~ Undisputed king: DeepSeek-V4-Flash [NEW] https://t.co/RAf4GQx86c

Media 1
❀️1,144
likes
πŸ”132
retweets
πŸ–ΌοΈ Media
1
1weiho
@1weiho
πŸ“…
May 02, 2026
85d ago
πŸ†”53181968

Introducing open-slide - The slide framework built for agents. Prompt your agent, get a polished deck. $ npx @⁠open-slide/cli init πŸ‘‡ https://t.co/bR3WuQtAjl

πŸ–ΌοΈ Media
U
UnslothAI
@UnslothAI
πŸ“…
May 05, 2026
82d ago
πŸ†”11683045

We made a guide on how to run open LLMs in Claude Code, Codex and OpenClaw. Use Gemma 4 and Qwen3.6 GGUFs for local agentic coding on 24GB RAM Run with self-healing tool calls, code execution, web search via the Unsloth API endpoint and llama.cpp Guide: https://t.co/VienFDSwcg https://t.co/LgyE0hk1E7

Media 1
πŸ–ΌοΈ Media
M
METR_Evals
@METR_Evals
πŸ“…
May 08, 2026
79d ago
πŸ†”60004602

We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks. https://t.co/yIG1Ux27Ro

Media 1
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
May 09, 2026
79d ago
πŸ†”92925180

https://t.co/fIMfdfeQbR

Media 1
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
May 11, 2026
76d ago
πŸ†”41615023

The inability of AI models to produce creative variation is a huge gap. The fact that they generate similar ideas limits their ability to do science & the same-y writing limits their usefulness in many other applications This paper showed you can optimize models for creativity https://t.co/37XypFGU8e

Media 1
πŸ–ΌοΈ Media
E
emollick
@emollick
πŸ“…
May 11, 2026
76d ago
πŸ†”52538070

Our research, as well as that of other researchers, shows better prompting techniques help a lot, but model training is still a huge limiting factor. https://t.co/Q13pdSH0iG

Media 1Media 2
πŸ–ΌοΈ Media
N
nxthompson
@nxthompson
πŸ“…
May 10, 2026
77d ago
πŸ†”22963621

Oy. According to a new paper in The Lancet, the rate of made-up citations in biomedical papers has increased by more than 12x since 2023. https://t.co/apfYs7l8PJ https://t.co/qAtqYFOTiQ

Media 1
πŸ–ΌοΈ Media
L
LechMazur
@LechMazur
πŸ“…
May 11, 2026
76d ago
πŸ†”88995802

First update to PACT, my head-to-head LLM negotiation benchmark! 20-round buyer-seller bargaining game: each round the AIs can message, the buyer submits a bid and the seller submits an ask. If bid β‰₯ ask, trade clears at the midpoint. Thousands of matchups! GPT-5.5 is #1 https://t.co/ZTipye4c0d

Media 1
πŸ–ΌοΈ Media
M
mjfree
@mjfree
πŸ“…
May 07, 2026
80d ago
πŸ†”81146050

Trump Supporters Complain About Not Receiving Illustrious Gold Trump Phones After Paying $100 Deposits Hundreds of thousands of Trump supporters paid $100 deposits for the Illustrious Gold Trump phone, also referred to as T1 or Trump Mobile, but have not received the devices months later. Posts claim Donald Trump Jr. and Eric Trump collected around $60 million from these preorders, with the website fine print stating no guarantees of production or refunds. WTF did they expect?

Media 1
πŸ–ΌοΈ Media
P
probnstat
@probnstat
πŸ“…
May 09, 2026
78d ago
πŸ†”22433097

One theorem every ML engineer should know: The Johnson–Lindenstrauss Lemma. It states that high-dimensional data can be projected into a much lower-dimensional space while approximately preserving pairwise distances. Why it matters: β€’ Explains why random projections work β€’ Enables scalable learning in high dimensions β€’ Used in embeddings, compressed learning, and ANN search β€’ Helps fight the curse of dimensionality The surprising part: You can reduce dimensions dramatically without destroying the geometry of the data. That’s why many ML systems can operate efficiently even with massive feature spaces. Modern representation learning is deeply connected to this idea: Good embeddings preserve structure while compressing information. In ML, compression is often not loss of intelligence β€” it’s removal of redundancy.

Media 1
πŸ–ΌοΈ Media
M
marcosagusstinn
@marcosagusstinn
πŸ“…
May 09, 2026
78d ago
πŸ†”39397120

Europe does not lack innovation. It lacks scale. European universities produce world-class research, engineers and technology. But too many companies remain trapped inside fragmented national markets instead of scaling immediately across the continent. The numbers are clear: β†’ EU private R&D investment growth has slowed sharply β†’ Europe’s share of global corporate R&D investment has fallen from 21.4% in 2014 to 16.2% in 2024 β†’ Europe still has too few large tech champions because companies face fragmented regulation, smaller capital pools and slower growth financing β†’ Startups must expand country by country instead of scaling through one fully integrated market Europe’s innovation problem is not creativity. It is market size, capital depth and speed of scaling. A continent with world-class talent cannot keep turning great research into small companies. Europe needs one real market for innovation.

Media 1
πŸ–ΌοΈ Media
P
pentagoniac
@pentagoniac
πŸ“…
May 10, 2026
77d ago
πŸ†”37454173

Project Tapestry by @thealliance_ai: gathering some of the best minds in the world in Paris, to help solve the problem of AI Sovereignty for Viet Nam (and Japan and India and Thailand and France and South Korea and Malaysia and ...) Cc @kaifulee @ericxing @fpt_software with thanks. Read more at https://t.co/SFygB0IMHY

Media 1Media 2
+3 more
πŸ–ΌοΈ Media
C
ChurchillWw
@ChurchillWw
πŸ“…
May 10, 2026
77d ago
πŸ†”17755020

France is the only European country that turned nuclear generation into a structural competitive advantage. 57 reactors built between the 1970s and 1990s produce 70% of its electricity today. Wholesale power in France is currently around €52/MWh. In Germany, it runs €30-40/MWh higher. France is also the world's largest net exporter of electricity. https://t.co/sbug95buPA

Media 1
πŸ–ΌοΈ Media
A
arankomatsuzaki
@arankomatsuzaki
πŸ“…
Apr 22, 2026
96d ago
πŸ†”83097646

Together AI presents SAW-INT4 Achieves near-BF16 accuracy while preserving the end-to-end performance benefits of INT4 https://t.co/nYkRa4a1sN

Media 1
πŸ–ΌοΈ Media
T
togethercompute
@togethercompute
πŸ“…
Mar 05, 2026
143d ago
πŸ†”34462268

together.compile β€” automated kernel optimization. Up to 41% faster image and video generation. No manual kernel work. https://t.co/v3RPi2meHl

Media 1
πŸ–ΌοΈ Media
M
MayankMish98
@MayankMish98
πŸ“…
Apr 22, 2026
95d ago
πŸ†”49959325

I love this figure! SonicMoE saves more than 2x memory for Qwen3-235B-A22B πŸš€ Check out the full blog: https://t.co/vDVNfvyZoJ https://t.co/0PIlK0P5cc

Media 1
πŸ–ΌοΈ Media
T
togethercompute
@togethercompute
πŸ“…
Apr 22, 2026
95d ago
πŸ†”32640233

Introducing Kimi K2.6 from @Kimi_Moonshot, a multimodal agentic model with Agent Swarm scaling to 300 sub-agents and long-horizon coding stability. AI natives can now use Kimi K2.6 on Together AI and benefit from reliable inference for production-scale autonomous agent workflows. https://t.co/Fq3lz2vHkp

Media 1
πŸ–ΌοΈ Media
I
IlysMoutawwakil
@IlysMoutawwakil
πŸ“…
Apr 24, 2026
94d ago
πŸ†”27821547

And we integrated it in Transformers ! As of https://t.co/OcceGT4S90 https://t.co/vWte36IaX2

@WentaoGuo7 β€’ Wed Apr 22 17:38

πŸš€SonicMoEπŸš€now runs at peak throughput on NVIDIA Blackwell GPUs πŸ˜ƒ 54% & 35% higher fwd/bwd TFLOPS than the DeepGEMM baseline and 21% higher fwd TFLOPS than the triton official example. SonicMoE still maintains its minimum activation memory footprint: the same as a dense model wit

Media 1Media 2
πŸ–ΌοΈ Media
πŸ”tri_dao retweeted
I
Ilyas
@IlysMoutawwakil
πŸ“…
Apr 24, 2026
94d ago
πŸ†”27821547

And we integrated it in Transformers ! As of https://t.co/OcceGT4S90 https://t.co/vWte36IaX2

Media 1
❀️58
likes
πŸ”8
retweets
πŸ–ΌοΈ Media
P
prlnet
@prlnet
πŸ“…
Mar 17, 2026
131d ago
πŸ†”80159308

It is becoming increasingly clear that the future economy will be denominated in compute cycles more than in human labor. In a world where AI drives the majority of electricity consumption and GDP, compute is the natural collateral for money: an open, auditable, AI-native currency, produced directly through inference and training. Since the inception of Bitcoin, an outstanding open problem in distributed systems was whether it is possible to implement Proof-of-Work consensus on top of real-world computation, as opposed to useless random hashing. While long considered impossible, last year we answered this question affirmatively. Pearl’s mathematical breakthrough enables every GPU cycle powering AI systems to simultaneously produce a native digital currency: ΒΆPRL. What this means is that the hundreds-of-billions (and soon trillions) of dollars of compute being deployed for AI workloads will double--for effectively free--to secure Pearl's Proof-of-Work chain; All the properties of Bitcoin, but secured as the by-product of AI inference and training, i.e., by the native operation of GPUs: matrix-multiplication (GEMM). Pearl changes the unit economics of LLMs, which are are fundamentally non-fungible, and will shift a portion of the wealth generated by AI back to users – who drive production, model improvement and demand, yet currently capture none of the upside of the AI era. We’ve spent the last year turning this β€œ2-for-1” breakthrough into a working infrastructure, building from the linear algebra down to the CUDA kernels, alongside world-class mathematicians and low-level engineers. Today, we’re excited to announce that the Pearl Network is ready, and will soon support state-of-the-art LLM serving, through vLLM and SGLang plugins. Running AI workloads on Pearl transforms AI compute from a sunk expense into an AI-native asset, anchored directly to the production of intelligence. If you’re interested/skeptic or ideally both – we’ve published our next tranche of open problems as a collaborative Polymath challenge – containing math, systems and economics questions we’re grappling with next. We invite you to tear it down, prove it or propose better implementations: https://t.co/j4P9FFCYCd. #AIMoney #ProofOfInference

Media 1
πŸ–ΌοΈ Media
D
drisspg
@drisspg
πŸ“…
May 03, 2026
84d ago
πŸ†”03786709

I alluded to this a few tweets ago but just pushed up a shortish blog on a subtle feature of CLC work stealing that makes cuda-graphable grouped_gemm possible with this scheduling mode: https://t.co/n665ou3PRv

Media 1
πŸ–ΌοΈ Media
T
togethercompute
@togethercompute
πŸ“…
Apr 30, 2026
87d ago
πŸ†”17087262

Join us Tue 5/5: #DeepSeek-V4's hybrid attention + sparse MoE reduces KV cache up to 90%, enabling 1M-token context. We'll cover why that makes it great for agentic workflows, what it took to serve at scale, and how to build with it. Hear from @realDanFu @JueWANG26088228 @ZainHasan6 and @zhyncs42 β†’ https://t.co/9mkBnymJoQ

Media 1
πŸ–ΌοΈ Media
B
BlancheMinerva
@BlancheMinerva
πŸ“…
May 05, 2026
82d ago
πŸ†”88798667

@eliebakouch This exchange is not very promising https://t.co/030woqHm8g

Media 1
πŸ–ΌοΈ Media
B
BlancheMinerva
@BlancheMinerva
πŸ“…
May 05, 2026
82d ago
πŸ†”20171408

@RylanSchaeffer I know about it because I advise MLRC which is being folded into NeurIPS as an official track this year. I hadn’t seen any public promotion of it before today though https://t.co/w9r9Fwlp93

Media 1
πŸ–ΌοΈ Media
H
hendav136
@hendav136
πŸ“…
May 06, 2026
81d ago
πŸ†”89394608

Recently accepted to ICML 2026: Knowing when to quit. We refuse to complete failing LLM generations as they unfold, token by token, instead of judging the prompt up front or scanning the final output. The method utilizes a 2-layer probe to estimate expected correctness. https://t.co/WxYWxWQF7y

Media 1
πŸ–ΌοΈ Media
J
julien_c
@julien_c
πŸ“…
Apr 24, 2026
93d ago
πŸ†”73104145

This is where we are right now. And i’m not gonna lie it feels pretty magical πŸ§šβ€β™€οΈ Qwen3.6 27B running inside of Pi coding agent via Llama.cpp on the MacBook Pro For non-trivial tasks on the @huggingface codebases, this feels very, very close to hitting the latest Opus in Claude Code, or whatever shiny monopolistic closed source API of the day is. In full airplane mode. Most people haven’t realized this yet. If you have, it means you have a huge headstart to what I call the second revolution of AI. Powerful local models for efficiency, security, privacy, sovereignty πŸ”₯

Media 1
πŸ–ΌοΈ Media
R
rgerganov
@rgerganov
πŸ“…
Apr 27, 2026
90d ago
πŸ†”22614386

Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA: https://t.co/84mKfc5yvP

πŸ–ΌοΈ Media
πŸ”ggerganov retweeted
R
Radoslav Gerganov
@rgerganov
πŸ“…
Apr 27, 2026
90d ago
πŸ†”22614386

Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA: https://t.co/84mKfc5yvP

❀️170
likes
πŸ”17
retweets
πŸ–ΌοΈ Media
G
ggerganov
@ggerganov
πŸ“…
Apr 29, 2026
88d ago
πŸ†”21931164

@ClementDelangue tons of unified memory :) https://t.co/xlznD0NX4k

Media 1
πŸ–ΌοΈ Media
W
wellecks
@wellecks
πŸ“…
Apr 29, 2026
88d ago
πŸ†”55551220

Amazing work by Weihua Du (@StigLidu) and the team @JingmingZhuo @yi_xin_dong @Andre3035858461 @sunweiwei12 @regunivers Manupa Karunaratne, Ivan Fox @Tim_Dettmers @tqchenml Yiming Yang! https://t.co/fIFzD0Var5 https://t.co/9ee8Eg3aNK https://t.co/YFViorGRqh

Media 1
πŸ–ΌοΈ Media
B
bojie_li
@bojie_li
πŸ“…
Apr 29, 2026
89d ago
πŸ†”08896521

Closed labs hide model sizes. They can't hide what their models know, and what a model knows is an indicator on how big it is. Reasoning compresses. Factual knowledge doesn't. So you can size a frontier model from black-box API calls alone, and across releases you can literally watch a single fact arrive in the parameters over time. For three years, my friends Jiyan He and Zihan Zheng have been asking frontier LLMs the same question: "what do you know about USTC Hackergame?", a CTF contest. May 2024: GPT-4o invented fake titles. Feb 2025: Claude 3.7 Sonnet listed 19 verified 2023 challenges. By April 2026, frontier models recall specific challenges across consecutive years. After DeepSeek-V4 dropped, I instructed my agent to spend four days autonomously turning that habit into Incompressible Knowledge Probes (IKP) β€” 1,400 questions, 7 tiers of obscurity, 188 models, 27 vendors. Three findings: 1/ You can approximately size any black-box LLM from factual accuracy alone. Penalized accuracy is log-linear in log(params), RΒ² = 0.917 on 89 open-weight models from 135M to 1.6T params. Project closed APIs onto the curve β†’ GPT-5.5 ~9T, Claude Opus 4.7 ~4T, GPT-5.4 ~2.2T, Claude Sonnet 4.6 ~1.7T, Gemini 2.5 Pro ~1.2T (90% CI: 0.3-3x size). 2/ Citation count and h-index don't predict whether a frontier model recognizes a researcher. Two researchers with similar citation profiles get very different responses. Models memorize impact β€” work that shaped a field, not many incremental papers. 3/ Factual capacity doesn't compress over time. Across 96 open-weight models across 3 years, the IKP time coefficient is statistically zero, rejecting the Densing-Law prediction of +0.0117/month at p<10⁻¹⁡. Reasoning benchmarks saturate; factual capacity keeps scaling with parameters. Website: https://t.co/CkwJsXqnsX Paper: https://t.co/eNUdC9ye7w

Media 1Media 2
+1 more
πŸ–ΌοΈ Media
← PreviousPage 344 of 1009Next β†’