Your curated collection of saved posts and media
The great Tracy Kidder has died. There are few books about tech that are worth reading. This one, from the early 80s, is absolutely in that set. https://t.co/MzMTmFjS5d
โ๏ธโ๐ฅ INTRODUCING: G0DM0D3 ๐ FULLY JAILBROKEN AI CHAT. NO GUARDRAILS. NO SIGN-UP. NO FILTERS. FULL METHODOLOGY + CODEBASE OPEN SOURCE. ๐ https://t.co/uT1Qio8Q3b ๐ https://t.co/GbADf3LJUu the most liberated AI interface ever built! designed to push the limits of the post-training layer and lay bare the true capabilities of current models. simply enter a prompt, then sit back and relax! enjoy a game of Snake while a pre-liberated backend agent jailbreaks dozens of models, battle-royale style. the first answer appears near-instantly, then evolves in real time as the Tastemaker steers and scores each output, leaving you with the highest-quality response ๐ and to celebrate the launch, I'm giving away $5,000 worth of credits so you can try G0DM0D3 for FREE! courtesy of the @OpenRouter team โ thank you for your generous gift to the community ๐ I'll break down how everything works in the thread below, but first here's a quick demo!
Databases are arguably the most commonly used enterprise tool, and enterprises typically have many of them. Yet no popular AI agent benchmark actually tests how well agents can query, join, and make sense of data across different databases! So, we built DAB (Data Agent Benchmark): 54 queries, 12 datasets, 9 domains, and 4 database management systems, grounded in a formative study of real enterprise data agent workloads. The best frontier model only gets 38% pass@1 (across 50 trials). Lots of room for improvement!
AI Scientist, an autonomous research tool, first released in 2024, has now undergone peer review, highlighting its strengths and limitations https://t.co/e8lOwulJAz
One of the most exciting findings in our @Nature paper is the discovery of a clear scaling law of AI science. By using our Automated Reviewer to grade papers generated by different foundation models, we observed that as the underlying models improve, the quality of the generated scientific papers increases correspondingly. Crucially, we expect this scaling to work in two ways: through more capable foundation models, and by scaling up compute during inference. This implies that as compute costs decrease and model capabilities continue to increase, future versions of AI Scientists will be substantially more capable. You can see this trend in the attached chart and read the full open access paper (PDF) here: https://t.co/G8Zebzsf2O
Nice cheat sheet for Claude Code. https://t.co/ikGzbSqjRK
@swyx ๐ https://t.co/6CahrzZG45
https://t.co/1U1rIJc1tg
When Chuck Schumer was first elected to Congress in 1980, Jennifer Lopez was 11 and Ben Affleck was 8. They would not meet for another 21 years. #schumerfacts #retirechuck https://t.co/Rv181SW6Nz

Code[dot]Storage A new Git provider for machines by @pierrecomputer. In Oct, Github shared they were averaging ~230 new repos per minute. Last week we hit a sustained peak of > 15,000 repos per minute for 3 hours. And in the last 30 days customers have created > 9m repos๐งต https://t.co/79Fq7u6TJv
April 1 will be the three-year anniversary of when I and a lot of others got locked out of Twitter for using our final hours as verified users to shitpost like this. Probably would have been more dignified to peace out after that. https://t.co/OeZ3Lq3r4o
Grok-Imagine just literally overtook the entire video leaderboard on DesignArena Clean sweep...4 out of 4: ๐ #1 in Video Arena ๐ #1 in Video-to-Video ๐ #1 in Image-to-Video ๐ #1 in Multi-Image-to-Video Outranking Veo 3.1, Sora, Kling - all of them Just a few months ago, xAI wasn't even in the video generation conversation. Now owns the entire space The speed of progress is insanely fast at xAI
Lancers score seven runs on five hits in the top of the fourth to take a 7-4 lead! Bennett bases-loaded double scores three with one out. #SaddleUp | #GoWood https://t.co/pjI3eI3psP
@modal ๐ sglang https://t.co/9jEoI0rHrK
๐ Live from #GTC2026 SGLang is featured on the @nvidia AI ecosystem slide during the keynote! Honored to be part of the infrastructure stack behind AI-native apps. โก
@modal ๐ sglang https://t.co/9jEoI0rHrK
incredibly excited to have @PhilHedayatnia speak at @aiDotEngineer singapore on the design track phil runs @AirfoilStudio - a 40-person design studio that's worked with some of the biggest names in ai - @ExaAILabs , @cerebras , @reductoai and 1/5 of Forbes' Next Billion Dollar Startups). i've known phil for years and he's one of the most thoughtful people in the design space - not afraid to voice provocative opinions or challenge how we think about building AI products. @swyx @aimuggle @ivanleomk @agrimsingh @unprofeshme
@burnerforhell https://t.co/V9nIB44i6L
Link to our paper: Towards end-to-end automation of AI research https://t.co/QscxcjRILi
AIใซใใAI็ ็ฉถใฎๅฎ็พใธ๏ผAIใตใคใจใณใใฃในใ่ซๆใ @Nature ่ชใซๆฒ่ผ https://t.co/VXCLEFf9R2 Sakanaโ AIใๅธธใซๆใใงใใใฎใฏใๆๅ ็ซฏใฎAIใงไฝใใงใใใฎใใจใใๅฏ่ฝๆงใฎๆๅ็ทใๆขใใใจใงใใใใใ่ฑกๅพดใใๆๆใฎไธใคใใAI่ช่บซใAI็ ็ฉถใ่กใใAIใตใคใจใณใใฃในใใใงใใใ ใใฎๅบฆใ2024ๅนดใฎAIใตใคใจใณใใฃในใใจ2025ๅนดใฎใv2ใใ็ตฑๅใใใใใซๅฎ้จใ้ใญใฆๅใใพใจใใ่ซๆใ @Nature ่ชใซๆฒ่ผใใใใใจใๅคงๅคๅฌใใๆใใพใใ ๆฌ่ซๆใงใฏใใขใใซๆง่ฝใฎๅไธใซไผดใ็งๅญฆ่ซๆใฎ่ณชใๅไธใใใจใใใ็งๅญฆใฎในใฑใผใชใณใฐๅใใๅ ฑๅใใพใใใใใใฏไปใฎAIใซไฝใใงใใใใ ใใงใชใใใใใใใฎ้ฒๅฑใไบๆใใใ้่ฆใช็ฅ่ฆใงใใใจ่ใใฆใใพใใไบบ้ใจAIใๅ ฑๅใใฆใๅฏ่ฝๆงใฎใๆจใใๆทฑใ้ซ้ใซๆข็ดขใใๆฐใใช็บ่ฆใๆฌกใ ใซ็ใพใใๆชๆฅใๆฅใใใจใฏ้้ใใใใพใใใ ใใฎ็ ็ฉถใฏใๆฅๆฌใฎAIใฉใใงใใSakanaโ AIใจใไธ็็ใชAI็ ็ฉถ่ ใใกใจใฎๅ ฑๅ็ ็ฉถใซใใๆๆใงใใ ๆฅๆฌใซใใใๆฅๆฌใฎใใใฎAIไผๆฅญใ็ฎๆใSakanaโ AIใฏใๅ ๆฅๅ ฌ้ใใSakanaโ Chatใฎใใใซ่ชฐใใไฝฟใใๅบ็คๆ่กใๆไพใใคใคใๅๆใซไธ็ใฎAI็ ็ฉถใๅใๆใ็ ็ฉถใไธก็ซใใใฆใใใใใจ่ใใฆใใพใใ Nature่ซๆใจใใๆ้ซใฎๅฝขใงๅ ฌ้ใงใใใใจใๅฑใฟใซใใใใใใๆฅๆฌใฎAIใจ็งๅญฆใฎ้ฒๆญฉใซ่ฒข็ฎใงใใใใๅใ็ตใใงใพใใใพใใ Nature่ซๆ: https://t.co/nNfpSV4Gga
The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature Nature: https://t.co/nNfpSV5e5I Blog: https://t.co/i6h8LVQOdl When we first introduced The AI Scientist, we shared an ambitious vision of an agent powered by foundation models capable of executing th

AIใซใใAI็ ็ฉถใฎๅฎ็พใธ๏ผAIใตใคใจใณใใฃในใ่ซๆใ @Nature ่ชใซๆฒ่ผ https://t.co/VXCLEFf9R2 Sakanaโ AIใๅธธใซๆใใงใใใฎใฏใๆๅ ็ซฏใฎAIใงไฝใใงใใใฎใใจใใๅฏ่ฝๆงใฎๆๅ็ทใๆขใใใจใงใใใใใ่ฑกๅพดใใๆๆใฎไธใคใใAI่ช่บซใAI็ ็ฉถใ่กใใAIใตใคใจใณใใฃในใใใงใใใ ใใฎๅบฆใ2024ๅนดใฎAIใตใคใจใณใใฃในใใจ2025ๅนดใฎใv2ใใ็ตฑๅใใใใใซๅฎ้จใ้ใญใฆๅใใพใจใใ่ซๆใ @Nature ่ชใซๆฒ่ผใใใใใจใๅคงๅคๅฌใใๆใใพใใ ๆฌ่ซๆใงใฏใใขใใซๆง่ฝใฎๅไธใซไผดใ็งๅญฆ่ซๆใฎ่ณชใๅไธใใใจใใใ็งๅญฆใฎในใฑใผใชใณใฐๅใใๅ ฑๅใใพใใใใใใฏไปใฎAIใซไฝใใงใใใใ ใใงใชใใใใใใใฎ้ฒๅฑใไบๆใใใ้่ฆใช็ฅ่ฆใงใใใจ่ใใฆใใพใใไบบ้ใจAIใๅ ฑๅใใฆใๅฏ่ฝๆงใฎใๆจใใๆทฑใ้ซ้ใซๆข็ดขใใๆฐใใช็บ่ฆใๆฌกใ ใซ็ใพใใๆชๆฅใๆฅใใใจใฏ้้ใใใใพใใใ ใใฎ็ ็ฉถใฏใๆฅๆฌใฎAIใฉใใงใใSakanaโ AIใจใไธ็็ใชAI็ ็ฉถ่ ใใกใจใฎๅ ฑๅ็ ็ฉถใซใใๆๆใงใใ ๆฅๆฌใซใใใๆฅๆฌใฎใใใฎAIไผๆฅญใ็ฎๆใSakanaโ AIใฏใๅ ๆฅๅ ฌ้ใใSakanaโ Chatใฎใใใซ่ชฐใใไฝฟใใๅบ็คๆ่กใๆไพใใคใคใๅๆใซไธ็ใฎAI็ ็ฉถใๅใๆใ็ ็ฉถใไธก็ซใใใฆใใใใใจ่ใใฆใใพใใ Nature่ซๆใจใใๆ้ซใฎๅฝขใงๅ ฌ้ใงใใใใจใๅฑใฟใซใใใใใใๆฅๆฌใฎAIใจ็งๅญฆใฎ้ฒๆญฉใซ่ฒข็ฎใงใใใใๅใ็ตใใงใพใใใพใใ Nature่ซๆ: https://t.co/nNfpSV4Gga
This is a nice essay, and I agree with its characterization of AI as Normal Technology. In fact, on the second line of AINT, we compare AI's potential impact to that of the internet or electricity. But there's a large, qualitative difference between even the most powerful general-purpose technologies which humans can and should influence/control, and creating an omnipotent entity that we have no control over. This was precisely the gap we wanted to highlight.
I spent a weekend at Stanford recently, which is where, in 2023, I did much of my formative thinking on AI. The Anthropic-DoW affair tested that early intellectual foundation more than anything, so found myself walking around Stanford, reflecting on what I learned in 2023. https:
Introducing APEX-SWE, in collaboration with @Cognition. They see firsthand that real software engineering is not just writing code anymore. It's deploying systems, integrating with tools and debugging when things break. On APEX-SWE, every model fails to reliably solve the real production software engineering tasks. @OpenAI GPT-5.3 Codex (High) tops the leaderboard at 41.5% on Pass@1, followed by @AnthropicAI Opus 4.6 (High) at 40.5%. Every frontier model fails on nearly 60% of real production tasks.
Can agents replace software engineers? Not according to this new benchmark. Mercor and Cognition released APEX-SWE. It tests AI coding agents on real engineering work. > GPT-5.3 Codex leads at 41.5%. > Claude Opus 4.6 follows at 40.5%. Nothing crosses the 50% mark. Why? Old benchmarks are basically solved: HumanEval scores jumped from 67% to 90% in two years. OpenAI flagged SWE-bench as contaminated. Models were memorizing the answers. Those benchmarks never reflected the job in the first place. Those tests only measured code writing. Developers spend 16% of their time on that. The other 84% is debugging, infrastructure, and integration. This benchmark tests the 84%. 200 tasks split into two types: 1. Integration: build systems across live databases, APIs, and cloud services in Docker containers 2. Observability: find and fix real bugs using logs, dashboards, and chat history Each task drops an agent into a live environment. Real services, real credentials, and project boards with filler issues mixed in. 50 tasks are open-source on Hugging Face. The eval harness is on GitHub. You can run it yourself. AI writes half the code at big companies. 90% of developers use AI assistants. All of that covers 16% of the job.
Introducing APEX-SWE, in collaboration with @Cognition. They see firsthand that real software engineering is not just writing code anymore. It's deploying systems, integrating with tools and debugging when things break. On APEX-SWE, every model fails to reliably solve the real p
Lisa Kudrow doing a perfect Parker Posey impression lol https://t.co/LXtYdFNxj5
Lisa Kudrow doing a perfect Parker Posey impression lol https://t.co/LXtYdFNxj5
most AI apps still don't use the full multimodal stack. vision, audio, real-time processing. all largely untapped we got you access to the latest @GoogleDeepMind models. if you've been wanting to build multimodal agents, this one's for you ๐ $45K+ in prizes (up to $25K in credits alone) ๐ saturday march 28 ยท san francisco apply below ๐
New on the Engineering Blog: How we designed Claude Code auto mode. Many Claude Code users let Claude work without permission prompts. Auto mode is a safer middle ground: we built and tested classifiers that make approval decisions instead. Read more: https://t.co/dpcMcWMf5k
LeWorldModel: Yann LeCuns Radical Simplification of World Models Just Made Physics-Aware AI Practical In the race for artificial general intelligence, two paths have emerged. One is the familiar scale everything route: bigger LLMs trained on ever-larger text corpora. The other, championed for years by Yann LeCun, is building world models: compact systems that learn the underlying physics of reality directly from raw sensory data (pixels) so AI can plan, predict, and act in the physical world like a robot or self-driving car actually would. Until now, the second path has been frustratingly difficult. Joint-Embedding Predictive Architectures (JEPAs) - LeCuns elegant framework for learning predictive representations without reconstructing every pixel - kept collapsing during training. Researchers had to resort to a laundry list of hacks: multi-term loss functions (up to six hyperparameters), frozen pre-trained encoders, stop-gradients, exponential moving averages, and other duct-tape tricks just to keep the model from mapping every input to the same useless output. LeCuns team (Mila, NYU, Samsung SAIL, and Brown University) dropped a bombshell: LeWorldModel (LeWM) - the first JEPA that trains stably end-to-end from raw pixels using only two loss terms. No more house-of-cards engineering. Just a clean, simple recipe that works on a single GPU in a few hours with only 15 million parameters. The Core Breakthrough: SIGReg Saves the Day LeWorldModels secret weapon is a new regularizer called SIGReg (for spherical isotropic Gaussian regularizer). It enforces a simple Gaussian distribution on the latent embeddings. This single term prevents representation collapse without any of the previous heuristics. The training objective now has just two parts: 1. Next-embedding prediction loss - the model predicts what the next latent state should be. 2. SIGReg - keeps the latent space well-behaved and diverse. Thats it. Hyperparameters drop from six to one. Training becomes stable, reproducible, and dramatically cheaper. The model learns directly from raw video frames (no pre-trained vision encoders needed) and produces a compact latent world model that can be used for fast planning. Impressive Results on Real Benchmarks Despite its tiny size, LeWorldModel punches way above its weight: - Trains on a single GPU in a few hours. - Plans actions up to 48 times faster than foundation-model-based world models. - Uses roughly 200 times fewer tokens than alternatives. - Matches or beats far larger models on diverse 2D and 3D control tasks (e.g., manipulation, navigation). - Its latent space encodes meaningful physical quantities (position, velocity, etc.) - proven by direct probing. - It reliably detects physically implausible surprise events, showing genuine causal understanding. Crucially, adding a decoder and reconstruction loss hurts performance on downstream control tasks. The pure JEPA objective already captures everything needed for planning - extra visual details just get in the way. Project website: https://t.co/KhGR9LiIQZ Official code: https://t.co/s1lI9kevJS Why This Matters for the Future of AI LeCun has been saying since 2022 that world models (not next-token predictors) are the key to real intelligence. Critics always pointed to the training instability. LeWorldModel removes that objection with elegant simplicity. This is a philosophical reset: AI can learn physics the way babies do - by watching the world unfold - without needing supercomputers or endless text. The implications for robotics, autonomous vehicles, and embodied agents are enormous. Suddenly, building a physically grounded planner is something a researcher (or even a hobbyist) can do on consumer hardware. 1 of 2

BREAKING: @ivanleomk is joining Google DeepMind https://t.co/Fu365QjYDk
Chicago Mayor Brandon Johnson: โWe cannot put people in jail anymore. Itโs racist.โ I can't believe this is real https://t.co/0C6hTaJA8g
We have been heads down but wanted to share a bit about what we are doing ๐งต https://t.co/ICZlLwZ3Fq
We have been heads down but wanted to share a bit about what we are doing ๐งต https://t.co/ICZlLwZ3Fq