Your curated collection of saved posts and media
First update to PACT, my head-to-head LLM negotiation benchmark! 20-round buyer-seller bargaining game: each round the AIs can message, the buyer submits a bid and the seller submits an ask. If bid β₯ ask, trade clears at the midpoint. Thousands of matchups! GPT-5.5 is #1 https://t.co/ZTipye4c0d
Trump Supporters Complain About Not Receiving Illustrious Gold Trump Phones After Paying $100 Deposits Hundreds of thousands of Trump supporters paid $100 deposits for the Illustrious Gold Trump phone, also referred to as T1 or Trump Mobile, but have not received the devices months later. Posts claim Donald Trump Jr. and Eric Trump collected around $60 million from these preorders, with the website fine print stating no guarantees of production or refunds. WTF did they expect?
One theorem every ML engineer should know: The JohnsonβLindenstrauss Lemma. It states that high-dimensional data can be projected into a much lower-dimensional space while approximately preserving pairwise distances. Why it matters: β’ Explains why random projections work β’ Enables scalable learning in high dimensions β’ Used in embeddings, compressed learning, and ANN search β’ Helps fight the curse of dimensionality The surprising part: You can reduce dimensions dramatically without destroying the geometry of the data. Thatβs why many ML systems can operate efficiently even with massive feature spaces. Modern representation learning is deeply connected to this idea: Good embeddings preserve structure while compressing information. In ML, compression is often not loss of intelligence β itβs removal of redundancy.
Europe does not lack innovation. It lacks scale. European universities produce world-class research, engineers and technology. But too many companies remain trapped inside fragmented national markets instead of scaling immediately across the continent. The numbers are clear: β EU private R&D investment growth has slowed sharply β Europeβs share of global corporate R&D investment has fallen from 21.4% in 2014 to 16.2% in 2024 β Europe still has too few large tech champions because companies face fragmented regulation, smaller capital pools and slower growth financing β Startups must expand country by country instead of scaling through one fully integrated market Europeβs innovation problem is not creativity. It is market size, capital depth and speed of scaling. A continent with world-class talent cannot keep turning great research into small companies. Europe needs one real market for innovation.
Project Tapestry by @thealliance_ai: gathering some of the best minds in the world in Paris, to help solve the problem of AI Sovereignty for Viet Nam (and Japan and India and Thailand and France and South Korea and Malaysia and ...) Cc @kaifulee @ericxing @fpt_software with thanks. Read more at https://t.co/SFygB0IMHY

France is the only European country that turned nuclear generation into a structural competitive advantage. 57 reactors built between the 1970s and 1990s produce 70% of its electricity today. Wholesale power in France is currently around β¬52/MWh. In Germany, it runs β¬30-40/MWh higher. France is also the world's largest net exporter of electricity. https://t.co/sbug95buPA
Together AI presents SAW-INT4 Achieves near-BF16 accuracy while preserving the end-to-end performance benefits of INT4 https://t.co/nYkRa4a1sN
together.compile β automated kernel optimization. Up to 41% faster image and video generation. No manual kernel work. https://t.co/v3RPi2meHl
I love this figure! SonicMoE saves more than 2x memory for Qwen3-235B-A22B π Check out the full blog: https://t.co/vDVNfvyZoJ https://t.co/0PIlK0P5cc
Introducing Kimi K2.6 from @Kimi_Moonshot, a multimodal agentic model with Agent Swarm scaling to 300 sub-agents and long-horizon coding stability. AI natives can now use Kimi K2.6 on Together AI and benefit from reliable inference for production-scale autonomous agent workflows. https://t.co/Fq3lz2vHkp
And we integrated it in Transformers ! As of https://t.co/OcceGT4S90 https://t.co/vWte36IaX2
πSonicMoEπnow runs at peak throughput on NVIDIA Blackwell GPUs π 54% & 35% higher fwd/bwd TFLOPS than the DeepGEMM baseline and 21% higher fwd TFLOPS than the triton official example. SonicMoE still maintains its minimum activation memory footprint: the same as a dense model wit

And we integrated it in Transformers ! As of https://t.co/OcceGT4S90 https://t.co/vWte36IaX2
It is becoming increasingly clear that the future economy will be denominated in compute cycles more than in human labor. In a world where AI drives the majority of electricity consumption and GDP, compute is the natural collateral for money: an open, auditable, AI-native currency, produced directly through inference and training. Since the inception of Bitcoin, an outstanding open problem in distributed systems was whether it is possible to implement Proof-of-Work consensus on top of real-world computation, as opposed to useless random hashing. While long considered impossible, last year we answered this question affirmatively. Pearlβs mathematical breakthrough enables every GPU cycle powering AI systems to simultaneously produce a native digital currency: ΒΆPRL. What this means is that the hundreds-of-billions (and soon trillions) of dollars of compute being deployed for AI workloads will double--for effectively free--to secure Pearl's Proof-of-Work chain; All the properties of Bitcoin, but secured as the by-product of AI inference and training, i.e., by the native operation of GPUs: matrix-multiplication (GEMM). Pearl changes the unit economics of LLMs, which are are fundamentally non-fungible, and will shift a portion of the wealth generated by AI back to users β who drive production, model improvement and demand, yet currently capture none of the upside of the AI era. Weβve spent the last year turning this β2-for-1β breakthrough into a working infrastructure, building from the linear algebra down to the CUDA kernels, alongside world-class mathematicians and low-level engineers. Today, weβre excited to announce that the Pearl Network is ready, and will soon support state-of-the-art LLM serving, through vLLM and SGLang plugins. Running AI workloads on Pearl transforms AI compute from a sunk expense into an AI-native asset, anchored directly to the production of intelligence. If youβre interested/skeptic or ideally both β weβve published our next tranche of open problems as a collaborative Polymath challenge β containing math, systems and economics questions weβre grappling with next. We invite you to tear it down, prove it or propose better implementations: https://t.co/j4P9FFCYCd. #AIMoney #ProofOfInference
I alluded to this a few tweets ago but just pushed up a shortish blog on a subtle feature of CLC work stealing that makes cuda-graphable grouped_gemm possible with this scheduling mode: https://t.co/n665ou3PRv
Join us Tue 5/5: #DeepSeek-V4's hybrid attention + sparse MoE reduces KV cache up to 90%, enabling 1M-token context. We'll cover why that makes it great for agentic workflows, what it took to serve at scale, and how to build with it. Hear from @realDanFu @JueWANG26088228 @ZainHasan6 and @zhyncs42 β https://t.co/9mkBnymJoQ
@eliebakouch This exchange is not very promising https://t.co/030woqHm8g
@RylanSchaeffer I know about it because I advise MLRC which is being folded into NeurIPS as an official track this year. I hadnβt seen any public promotion of it before today though https://t.co/w9r9Fwlp93
Recently accepted to ICML 2026: Knowing when to quit. We refuse to complete failing LLM generations as they unfold, token by token, instead of judging the prompt up front or scanning the final output. The method utilizes a 2-layer probe to estimate expected correctness. https://t.co/WxYWxWQF7y
This is where we are right now. And iβm not gonna lie it feels pretty magical π§ββοΈ Qwen3.6 27B running inside of Pi coding agent via Llama.cpp on the MacBook Pro For non-trivial tasks on the @huggingface codebases, this feels very, very close to hitting the latest Opus in Claude Code, or whatever shiny monopolistic closed source API of the day is. In full airplane mode. Most people havenβt realized this yet. If you have, it means you have a huge headstart to what I call the second revolution of AI. Powerful local models for efficiency, security, privacy, sovereignty π₯
Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA: https://t.co/84mKfc5yvP
Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA: https://t.co/84mKfc5yvP
@ClementDelangue tons of unified memory :) https://t.co/xlznD0NX4k
Amazing work by Weihua Du (@StigLidu) and the team @JingmingZhuo @yi_xin_dong @Andre3035858461 @sunweiwei12 @regunivers Manupa Karunaratne, Ivan Fox @Tim_Dettmers @tqchenml Yiming Yang! https://t.co/fIFzD0Var5 https://t.co/9ee8Eg3aNK https://t.co/YFViorGRqh
Closed labs hide model sizes. They can't hide what their models know, and what a model knows is an indicator on how big it is. Reasoning compresses. Factual knowledge doesn't. So you can size a frontier model from black-box API calls alone, and across releases you can literally watch a single fact arrive in the parameters over time. For three years, my friends Jiyan He and Zihan Zheng have been asking frontier LLMs the same question: "what do you know about USTC Hackergame?", a CTF contest. May 2024: GPT-4o invented fake titles. Feb 2025: Claude 3.7 Sonnet listed 19 verified 2023 challenges. By April 2026, frontier models recall specific challenges across consecutive years. After DeepSeek-V4 dropped, I instructed my agent to spend four days autonomously turning that habit into Incompressible Knowledge Probes (IKP) β 1,400 questions, 7 tiers of obscurity, 188 models, 27 vendors. Three findings: 1/ You can approximately size any black-box LLM from factual accuracy alone. Penalized accuracy is log-linear in log(params), RΒ² = 0.917 on 89 open-weight models from 135M to 1.6T params. Project closed APIs onto the curve β GPT-5.5 ~9T, Claude Opus 4.7 ~4T, GPT-5.4 ~2.2T, Claude Sonnet 4.6 ~1.7T, Gemini 2.5 Pro ~1.2T (90% CI: 0.3-3x size). 2/ Citation count and h-index don't predict whether a frontier model recognizes a researcher. Two researchers with similar citation profiles get very different responses. Models memorize impact β work that shaped a field, not many incremental papers. 3/ Factual capacity doesn't compress over time. Across 96 open-weight models across 3 years, the IKP time coefficient is statistically zero, rejecting the Densing-Law prediction of +0.0117/month at p<10β»ΒΉβ΅. Reasoning benchmarks saturate; factual capacity keeps scaling with parameters. Website: https://t.co/CkwJsXqnsX Paper: https://t.co/eNUdC9ye7w

Today weβre bringing new NSF OMAI compute online with NVIDIA Blackwell Ultra-powered systems, turning a $152M national investment from @NSF & @NVIDIA into a foundation for truly open AI research. π§΅ https://t.co/qFgtiibgAK

Happy 80th birthday John Waters! β‘οΈ β¨[π·: Rajendra Roy] https://t.co/G35qVTHOP4
Happy 80th birthday John Waters! β‘οΈ β¨[π·: Rajendra Roy] https://t.co/G35qVTHOP4
lena dunham reportedly sold 60k copies of her book in the first week. she did a ton of press on substackβguest essays, q&as, etc. thought this was an interesting insight that likely also applies to products other than books. https://t.co/JY3FO0tux7
100 years old and still the coolest person alive. Happy birthday, Sir David! https://t.co/dgxQb6LPZ6
100 years old and still the coolest person alive. Happy birthday, Sir David! https://t.co/dgxQb6LPZ6
Python, MCP, A2A, and a whole lot about how to build multi-agent systems. This goes well beyond building a single agent: it focuses on how multiple agents can collaborate towards a goal. Best part: The book will show you how to build everything from scratch. You start with a simple agent and then add complexity as the book progresses. Amazon link: https://t.co/GIIiIOKxgw
@Santandave1 βHeβs known for changing a key like John Coltraneβ https://t.co/6kjL98wfvE