Your curated collection of saved posts and media
Cautious Weight Decay is a surprisingly simple technique that has been repeatedly validated in Modded NanoGPT. I expect it will gain serious traction as the default variant of decoupled weight decay https://t.co/BLFSeqrf2w
modded-nanogpt WRs ๐ค Cautious Weight Decay Pretty much all the @speedrun WRs since November have used CWD in some form (e.g., @varunneal's "CWD w/ schedule"). Huge thanks to everyone experimenting and sharing results โ shoutout to @varunneal, @classiclarryd, @ChrisJMcCormick,
Did some statistics. My productivity ~doubled with moving from Claude Code to codex. Took me a bit to figure out at first but then ๐ฅ https://t.co/cfyKg0E1hf
Someome made a morning report skill and I just gave @clawdbot the tweet and it set up the skill + cron job. https://t.co/CXo0xMGcFv
Cowork but with local models not to send all your data to a remote cloud! https://t.co/2OrBMMO3NJ
In Cowork, you give Claude access to a folder on your computer. Claude can then read, edit, or create files in that folder. Try it to create a spreadsheet from a pile of screenshots, or produce a first draft from scattered notes. https://t.co/GEaMgDksUp
Cowork but with local models not to send all your data to a remote cloud! https://t.co/2OrBMMO3NJ
Codex team is at it again with just another insanely useful feature. If you see your agent going off the rails, or needing some addt'l context, you no longer need to stop the agent. Follow up prompts while the agent is working now inserts the prompt at the next thinking step, allowing it to pivot or provide better outputs without actually stopping the agent. This is undoubtedly a token and time saver. You can still queue up prompts using Tab. Well played!
Weโre releasing Action100M: the largest open dataset of ~15 years of video with dense action + caption annotations. It's a key ingredient behind VL-JEPA, now open to fuel the next generation of VLMs, World Models, and Robotics policies. Dataset on HF: https://t.co/bUs67QyOdd https://t.co/IzoGNUXzbJ
We release Action100M, the hero behind VL-JEPA. It is a large dataset with O(100 million) dense action annotations on HowTo100M procedural videos. We hope it serves as a robust data foundation to advance physical world modeling research.

I wrote an interactive article explaining the geometric intuition behind Rectified Flows. I visually explain why flow-models tend to learn curved trajectories, why this is bad for sampling latency, and a relatively simple technique for mitigating it. Check it out! Link ๐ https://t.co/X7nPye8fuJ
Bun โ the blazing-fast JS runtime that's eating Node.js's lunch โ already ships great llms.txt docs to supercharge your AI coding assistant. Discover these llms.txt files effortlessly while you browse: โ llmsdottxt โ Open-source Chrome extension that auto-detects llms.txt files, shows a red badge, lets you copy URLs/content instantly, and keeps a history. Perfect for Cursor, Claude, Windsurf, etc. https://t.co/nIBqwIjs3T Give it a try & star if it helps your workflow! โญ

ScaledML is back! After a 6 year hiatus. Consistently 2-3 years ahead of where ML will be. Examples of foresight at SML: - OpenAI in 2016,2017,2019 announced GPT-2 & RL efforts - Turing award for Deep Learning announced by Turing award winner on morning of award - Groq chip (acquired for $20bn) released - All Google TPU versions detailed Join us on January 29th at the Computer History Museum!
ScaledML is back! After a 6 year hiatus. Consistently 2-3 years ahead of where ML will be. Examples of foresight at SML: - OpenAI in 2016,2017,2019 announced GPT-2 & RL efforts - Turing award for Deep Learning announced by Turing award winner on morning of award - Groq chip (acquired for $20bn) released - Google TPU arch Join on Jan 29 at Computer History Museum!
CTO and Founder of Cerebras will be at ScaledML https://t.co/WoH27Uxx6q
can we go back? https://t.co/ePsrW4P5XD
๐ค https://t.co/bXGfPfQtlE
Stoicism 101: You are not required to have an opinion on every controversy. In fact, youโll be much happier if you donโt. https://t.co/KXkPgZByU3
#reallifedataviz https://t.co/l415d8BEdB
'How Much Grams?' A mini eval inspired by that viral guy who would ask ChatGPT voice+video to guess the weight of things. Flash crushes it - both speed and accuracy. Post: https://t.co/HNJWhGQIEb https://t.co/191EqCtbnt
Data from my sleep tracking came in handy recently: can you see which 5 days my wife tried a new feather pillow? (Graph shows my snoring duration) A similarly stark positive change came from running an air purifier. Probably lots more places I'm unknowingly sub-optimal... :) https://t.co/PdjHZrxMyI
@ATinyGreenCell I'm trying a https://t.co/ZWqHLe1GOr for similar reasons
It is November 2022, you yap "LLM cant reason because they are autoregressive!! We need cat level intelligence, neuro-symbolic AI! Chatgpt is fun toy product!" It is year 2026, they solve conjectures in 40 min. You literally share proof of open problems with chatgpt links. Actually insane timeline.
I've solved a second Erdos problem (#281) using only GPT 5.2 Pro - no prior solutions found. Terence Tao calls it "perhaps the most unambiguous instance" of AI solving an open problem: https://t.co/TBiCwiSFzl
Saturday hobby fun: biolistics tests with different nozzles, sticking DNA to carriers, and making a computer-controlled pipette for moving microliters of liquid around ๐ (Also breakfast date, happy walks with toddler niece, fireside reading - perfect Saturday) https://t.co/eNM9l57uiZ
Going to make so much agar art with this bad boy ๐ https://t.co/iFYdZq64SG
Repo: https://t.co/PRh999W0kp
Must watch. Why Transformers are taking over CNNs in computer vision. https://t.co/rH6CV0qdAK
Video: https://t.co/JySWpaKtvF
> Be Maor Shlomo. > Fail for a decade. > See Lovable take off. > Grab Claude 3.5. > Build a competitor in weeks. > Add a twist to it. > Hit $230k MRR in 90 days. > Sell it for $80M in 4 months. https://t.co/iOiAqRIP5a
Google just gave language models real long-term memory. A new architecture learns during inference and keeps context across millions of tokens. It holds ~70 percent accuracy at 10 million tokens. ๐ง๐ต๐ถ๐ ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ ๐น๐ฒ๐ฎ๐ฟ๐ป๐ ๐๐ต๐ถ๐น๐ฒ ๐ถ๐ ๐ฟ๐๐ป๐ Titans adds a neural long-term memory that updates during generation. Not weights. Not retraining. Live learning. โข A small neural network stores long-range context โข It updates only when something unexpected appears โข Routine tokens get ignored to stay fast This lets the model remember facts from far earlier text without scanning everything again. ๐๐ ๐ธ๐ฒ๐ฒ๐ฝ๐ ๐๐ฝ๐ฒ๐ฒ๐ฑ ๐๐ต๐ถ๐น๐ฒ ๐๐ฐ๐ฎ๐น๐ถ๐ป๐ด ๐ฐ๐ผ๐ป๐๐ฒ๐ ๐ Attention stays local. Memory handles the past. โข Linear inference cost โข No quadratic attention blowups โข Stable accuracy past two million tokens ๐๐ ๐ฎ๐น๐น๐ผ๐๐ ๐๐ผ๐ ๐๐ผ ๐ฏ๐๐ถ๐น๐ฑ ๐ป๐ฒ๐ ๐ธ๐ถ๐ป๐ฑ๐ ๐ผ๐ณ ๐ฎ๐ฝ๐ฝ๐ You can process full books, logs, or genomes in one pass. You can keep state across long sessions. You can stop chunking context just to survive limits.
Project: https://t.co/8aKlYQ2gb6
Small models just beat giant LLM agents at their own job. Not by thinking harder, but by coordinating better. A new system just outscored GPT-5 on Humanityโs Last Exam, using far less compute. ๐ง๐ต๐ถ๐ ๐๐๐๐๐ฒ๐บ ๐ฟ๐ฒ๐ฝ๐น๐ฎ๐ฐ๐ฒ๐ ๐ผ๐ป๐ฒ ๐ฏ๐ถ๐ด ๐ฏ๐ฟ๐ฎ๐ถ๐ป ๐๐ถ๐๐ต ๐ฎ ๐ฐ๐ผ๐ป๐ฑ๐๐ฐ๐๐ผ๐ฟ Instead of one model doing everything, it assigns roles. โข Large models handle hard reasoning. โข Small models handle routine steps. โข A controller decides what to call, when. That controller is trained only to make decisions. ๐๐ ๐น๐ฒ๐ฎ๐ฟ๐ป๐ ๐ฐ๐ผ๐ผ๐ฟ๐ฑ๐ถ๐ป๐ฎ๐๐ถ๐ผ๐ป, ๐ป๐ผ๐ ๐ฝ๐ฟ๐ผ๐บ๐ฝ๐ ๐๐ฟ๐ถ๐ฐ๐ธ๐ It uses reinforcement learning, not hand rules. Rewards optimize three things at once: - Task success - Latency - Compute cost ๐ง๐ต๐ฒ ๐ฟ๐ฒ๐๐๐น๐๐ It scores 37.1% on HLE versus 35.1%. Runs about 2.5ร faster. Uses roughly 70% less cost. This lets you build agents that scale by coordination, not parameters.
https://t.co/GTBUA1RzVy https://t.co/sB5dE5xoiY
https://t.co/GTBUA1RzVy https://t.co/sB5dE5xoiY
Great paper on Agentic Memory. LLM agents need both long-term and short-term memory to handle complex tasks. However, the default approach today treats these as separate components, each with its own heuristics, controllers, and optimization strategies. But memory isn't two independent systems. It's one cognitive process that decides what to store, retrieve, summarize, and forget. This new research introduces AgeMem, a unified framework that integrates long-term and short-term memory management directly into the agent's policy through tool-based actions. Instead of relying on trigger-based rules or auxiliary memory managers, the agent learns when and how to invoke memory operations: ADD, UPDATE, DELETE for long-term storage, and RETRIEVE, SUMMARY, FILTER for context management. It uses a three-stage progressive RL strategy. First, the model learns long-term memory storage. Then it masters short-term context management. Finally, it coordinates both under full task settings. To handle the fragmented experiences from memory operations, they design a step-wise GRPO (Group Relative Policy Optimization) that transforms cross-stage dependencies into learnable signals. The results across five long-horizon benchmarks: > On Qwen2.5-7B, AgeMem achieves 41.96 average score compared to 37.14 for Mem0, a 13% improvement. > On Qwen3-4B, the gap widens: 54.31 vs 44.70. Adding long-term memory alone provides +10-14% gains. > Adding RL training adds another +6%. > The full unified system with both memory types achieves up to +21.7% improvement over no-memory baselines. The unified memory management through learnable tool-based actions outperforms fragmented heuristic pipelines, enabling agents to adaptively decide what to remember and forget based on task demands. Paper: https://t.co/twhfiEsnho Learn to build effective AI agents in our academy: https://t.co/JBU5beIoD0