Your curated collection of saved posts and media
Introducing Radar, real-time collective intelligence for web agents π§΅ In our tests, agents that get a cache hit on Radar avoid the need for a browser entirely, dramatically improving agent performance Built & tested with @convex, @daytonaio and @superset_sh, @browser_use https://t.co/bFaFj6aMAA
Xcode 26.3 with Claude Agent & Codex hits the Mac App Store today! With advanced reasoning capabilities in Xcode, you can streamline workflows and build faster. And MCP support lets you easily connect other compatible agents. https://t.co/88NjaznE6E
I asked Cursor to add Vim support to the Ladybird browser. It automatically set up the environment to run the browser, made the code changes, and sent me a recorded demo. Not just for web apps! https://t.co/qDxnOr6CHU
My new favorite tmux dev layout features @opencode (with Kimi K2.5 running on @FireworksAI_HQ) on top and Claude Code on the bottom. I start almost all agent tasks with Kimi (so fast!), then ask Claude if I need a second opinion/more advanced stuff. Great combo! https://t.co/cUxfPgHFlW
Claude renovated my GitHub homepage for me by automatically setting up a CRON that pulls in my latest blog posts, and found images and other details to make things a bit nicer :) https://t.co/42GrcQ4w05
@amankhan Yeah. Just for clarity what I'm talking about is that this is when I'm interacting with claude it makes a mini interface https://t.co/dIg883ww2t
@garrytan We are absolutely back! https://t.co/ruVRC4CGxD
1/5 Happy CNYπ Still bothered by RL off-policy instability in LLM? Introducing a new wayπ‘Adaptive Layerwise Perturbation (ALP)π‘, a simple but robust fix that outperforms GRPO/MIS/Bypass, achieves better stability (KL, entropy) and exploration! π Blog: https://t.co/0def1Nb7uI https://t.co/9epsd4xJNp

https://t.co/YNGWOvywHf for those donβt remember it
pewdiepie just trained his own llm, and it beats gpt-4o on coding benchmarks. an apocalyptic, civilization-ending catastrophe of laughably, cosmically disproportionate magnitude for the entire ml research job category https://t.co/loDJZPnwN5
Wait, what?! PewDiePie using @axolotl_ai for his project! π₯ https://t.co/vnXeDfMzcc
pewdiepie just trained his own llm, and it beats gpt-4o on coding benchmarks. an apocalyptic, civilization-ending catastrophe of laughably, cosmically disproportionate magnitude for the entire ml research job category https://t.co/loDJZPnwN5
For those following the DoW AI drama, I highly recommend reading this post explaining how @OpenAI approached the negotiations with the DoW.
Yesterday we reached an agreement with the Department of War for deploying advanced AI systems in classified environments, which we requested they make available to all AI companies. We think our deployment has more guardrails than any previous agreement for classified AI deploy

https://t.co/3luoY7bgyM

Meta presents VecGlypher Unified Vector Glyph Generation with Language Models paper: https://t.co/anAFlgLMMV https://t.co/Nh3OpUBwa9

π₯Tongyi Lab releases Mobile-Agent-v3.5οΌ20+SOTA GUI benchmarks: (1) GUI automation, 56.5OSWorld, 71.6AndroidWorld, and48.4WebArena; (2) Grounding, 80.3ScreenSpotPro; (3) tool-calling , 47.6OSWorld-MCP @_akhaliq #LLM #Agent #GUI https://t.co/xCbyL0JZLl
Gradio's new HTML component is crazy! 3D Camera Control designed as a game-pad style toggleπ€― Try it for free on @huggingface π https://t.co/xoGdbrej3F
PewDiePie using Hugging Face π₯ https://t.co/9D7PUA0sun
seeing Hugging Face and The Stack/BigCode on PewDiePie video wasn't in my 2026 bingo card https://t.co/AIIOGzwOZT
seeing Hugging Face and The Stack/BigCode on PewDiePie video wasn't in my 2026 bingo card https://t.co/AIIOGzwOZT
The Trinity of Consistency as a Defining Principle for General World Models paper: https://t.co/21cbl3hAdu https://t.co/9YmzIPmsBJ

From Statics to Dynamics Physics-Aware Image Editing with Latent Transition Priors paper: https://t.co/Duflv5VmKj https://t.co/da73QuSURs

Introducing Code Review Bench v0: https://t.co/iAZDURyqol The first independent code review benchmark. 200,000+ PRs. Unbiased. Fully OSS. Updated daily. Tool performance highlights π§΅π Featuring: @augmentcode @baz_scm @claudeai @coderabbitai @cursor @GeminiApp @github @graphite @greptile @kilocode @OpenAIDevs @propelcode @QodoAI
Introducing Code Review Bench v0: https://t.co/iAZDURyqol The first independent code review benchmark. 200,000+ PRs. Unbiased. Fully OSS. Updated daily. Tool performance highlights π§΅π Featuring: @augmentcode @baz_scm @claudeai @coderabbitai @cursor @GeminiApp @github @graphite @greptile @kilocode @OpenAIDevs @propelcode @QodoAI
βοΈΒ Hello? AI Selves now have phone numbers! Put them in your imessage or SMS to be there when youβre not, settle arguments in your group chats, and make talking to yourself more normal. More ideas ππ§΅ Plus, weβre letting more people in off of our waitlist! QRT to get your own early access code.
Imagination Helps Visual Reasoning, But Not Yet in Latent Space Causal mediation analysis reveals latent visual reasoning in MLLMs fails: latent tokens ignore inputs and barely affect answers. CapImagine, a text-based alternative, teaches explicit imagination and significantly outperforms latent baselines.
Top AI Papers of The Week (Feb 24 - Mar 2) - A Very Big Video Reasoning Suite: 200 tasks, 1M+ video clips for video reasoning research - Does Your Reasoning Model Implicitly Know When to Stop Thinking? Introducing SAGE paradigm - AgentFly: Fine-tuning LLM agents without fine-tuning LLMs - Microsoft rStar2-Agent: 80.6% on AIME24 with just 14B parameters - From Blind Spots to Gains: Diagnostic-driven iterative training for LMMs - VibeVoice: Synthesizing 90-minute multi-speaker conversational speech - Alibaba MobilityBench: Benchmarking real-world route-planning agents - NVIDIA's data engineering strategies for scaling LLM terminal capabilities - VESPO: Variational sequence-level soft policy optimization for stable RL training - Beyond Pass@1: Self-play with variational problem synthesis sustains RLVR Find them below:
Thanks AK for reposting our work! Here are all the links for anyone who wants to check out more! Paper:Β https://t.co/6PajZXj6V0 Project Website:Β https://t.co/5VTiCqTDhN EvalKit:Β https://t.co/lxhyzMaI8j Cloud Infra:Β https://t.co/QNJRfOKQN3 Training Set:Β https://t.co/DlzLojQjsR Eval Set:Β https://t.co/Tzs2jAN99C Leaderboard:Β https://t.co/peZ1XkelYY Model:Β https://t.co/gFFJofrlNR
A Very Big Video Reasoning Suite paper: https://t.co/3ZY56TfbwD https://t.co/ojn1cL8VVN

JavisDiT++ Unified Modeling and Optimization for Joint Audio-Video Generation https://t.co/bd8BlNZNEr
Top AI Papers of The Week (Feb 16-22) - Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs - SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise - GLM-5: from Vibe Coding to Agentic Engineering by @zhipuAI - Experiential Reinforcement Learning - MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs - Zooming without Zooming: Region-to-Image Distillation by @InclusionAI - Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines? - DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval - SLA2: Sparse-Linear Attention with Learnable Routing and QAT - SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Find them below:
Just shipped! @huggingface storage add-ons. Starting at $12/month per TB - 3x cheaper than regular cloud storage, with very fast uploads and downloads powered by Xet's deduplication. You can now buy, upgrade, and cancel storage plans directly from your billing settings. https://t.co/RDylcDjkb4
Iβm giving an agent control over Reachy Mini from @huggingface and letting it understand and share spatial data via @Spectacles AR is the human interface for robotics and physical AI imo. It feels like absolute magic to interact with this, both in voice/agent and βpuppeteeringβ mode. Iβll probably work on AR for either an arm (manipulation tasks) or some sort of drone (locomotion in 3D space) nextβ¦ Project is fully open source btw: https://t.co/pmkXJR0U7f Thank you @SensAIHackademy for sending me the robot!
TranslateGemma 4B by @GoogleDeepMind now runs 100% in your browser on WebGPU with Transformers.js v4. 55 languages. No server. No data leaks. Works offline. A 4B parameter translation powerhouse, right in your browser. Try the demo π https://t.co/YgYskHqBRm