Your curated collection of saved posts and media
Woke Ninth Circuit decides Korean spas have to let biological men swing their gear in front of women and children. The dissent is pure gold: https://t.co/kwYpQpcGBQ
There are 10 types of people who'll love this week's @code release: those who read the version as one-point-one-hundred-eleven, and those who think we just shipped 1.7. https://t.co/uF74Xb4vBk
White people donβt even intubate their patients anymore they just look at you like this https://t.co/AFzq14p8Th
White people donβt even intubate their patients anymore they just look at you like this https://t.co/AFzq14p8Th
How often do LLMs claim to prove false mathematical statements? In our latest benchmark, BrokenArXiv, we find they do so very often. The best model, GPT-5.4, only rejects 40% of incorrect statements obtained by perturbing recent ArXiv papers, and other models do much worse. https://t.co/RRQNZfnCtW
Reuters published a piece. If companies like OpenAI or Anthropic fail, the massive financial ecosystem built around their existence could rapidly collapse. These labs are the primary customers for the $650B that tech giants are spending on new data centers and chips this year. Without their relentless demand for computing power, the expansion of new data centers would violently hit the brakes. It would also leave huge power grid projects and physical infrastructure investments completely abandoned and useless. Banks and private credit lenders who poured roughly $900B into this space would face severe uncertainty and massive potential losses. While a bigger tech company might swoop in to buy the failed labs for cheap, the overall value of the entire AI industry would instantly crash. Ultimately, the failure of just one of these major labs would not be a simple corporate bankruptcy. It would trigger a massive shockwave that drags down cloud providers, chipmakers, and global infrastructure projects all at once. -- reuters. com/commentary/breakingviews/what-happens-if-openai-or-anthropic-fail-2026-03-11/
Absolutely loving the designs that BrainGrid is generating https://t.co/YS2OJSZ3Td
I'm communicating with an LLM (SolveIt) via my handheld A4 whiteboard today! π Feels like a really smooth and natural process. More. ππ§΅ https://t.co/PtIRiZzDQT
Are Video Reasoning Models Ready to Go Outside? paper: https://t.co/g4TXU0cbeI https://t.co/r07UbvWXBH

DVD Deterministic Video Depth Estimation with Generative Priors paper: https://t.co/Eh41hneFEg https://t.co/qJv9H9Mspn
Impressive 7B multimodal vision language model π₯ Available on @huggingface πhttps://t.co/QdpJe5Pss4
Meet Reka Edge β Our next-generation vision language model for physical AI. Uses 3x fewer input tokens and achieves 65% faster throughput compared to leading 8B models. Image understanding, video analysis, object detection, and tool use. Built for Action. Fast enough for product
Impressive 7B multimodal vision language model π₯ Available on @huggingface πhttps://t.co/QdpJe5Pss4
Perplexity Computer is now on mobile. Start any task on any device. Manage Computer from your phone or desktop with cross-device synchronization. Available now for iOS in the Perplexity app. Coming soon to Android. https://t.co/hTw6fDIeaa
This is one of the coolest Pi Day traditions I have ever seen, and I'm sharing in the hopes that more math enthusiasts might be interested in joining! On Friday, March 13, Prof. Cory Palmer and a team of graduate students at UM are doing something extremely neat β a 24-hour nonstop marathon math lecture which will cover, essentially, the material of an entire math degree in one day! Lecture starts Friday 9a MDT and runs straight through the night until Saturday 9a. Itβs designed to be open and accessible, so you can join anytime, come in as a beginner, and leave knowing *a lot* more math. ποΈ Schedule: https://t.co/gaqUyjQ3tP ποΈ Livestream: https://t.co/e6xwvIvYGr If youβre like me and you wish you lived five lives, so in each one you get to be a mathematician versed in a completely different subfield, or if you just enjoy the idea of people explaining topology in the middle of the night, this is worth dropping into.

Introducing HandelBot πΉπ€, a real-world piano playing robot! Piano is extremely hard (even for humans!). We take a small but exciting step to replicate this beautiful skill w HandelBot. Our insight is combining sim priors w real world refinement & RL. w/ @haozhiq @DorsaSadigh https://t.co/8IHK7zYUrn
OpenClaw can make mistakes. Gensee Crate (https://t.co/C4OzcerfBG) mitigate this with time machines: πΈ Snapshots β Capture complete state of your OpenClaw instance π°οΈ Rollback β Restore to any previous snapshot π¨βπ©βπ¦βπ¦ Multiple Instances β Run 3 isolated agents at once, all on 24/7 https://t.co/Tu1jnDpAcL

We collaborated with @NVIDIA to teach you about Reinforcement Learning and RL environments. Learn: β’ Why RL environments matter + how to build them β’ When RL is better than SFT β’ GRPO and RL best practices β’ How verifiable rewards and RLVR work Blog: https://t.co/Jng3urMPyw https://t.co/CmEj1S3QAe

Here's the longer version of our Nature piece. Our argument is simple:Β statistical approximation is not the same thing as intelligence. Strong benchmark scores often say very little about how LLMs behave underΒ novelty, uncertainty, or shifting goals. Even more importantly,Β similar behaviors can arise from fundamentally different processes. In another paper, we identifiedΒ seven epistemological fault linesΒ between humans and LLMs. For example, LLMs have no internal representation of what is true. They often generateΒ confident contradictions, especially in longer interactions, because they do not track what is actually true. Another example. Yes, LLMs have solved some open mathematical problems, but these cases typically involveΒ applying known methods to well-defined problems. LLMs cannot invent anything that is truly new and true at the same time, because they lack the epistemic machinery to determine what is true. None of this means LLMs are useless. Quite the opposite: they are extraordinarily useful. But we should be careful about what they are and what they are not. Producing plausible text is not the same as understanding. Statistical prediction is not the same as intelligence. So despite the hype from the usual suspects,Β AGI has not been achieved. * paper in the first reply Joint with @Walter4C and @GaryMarcus
Big update for the Codex Remote Control for iOS! /commands are now available! β /review: Code Review β /status: Context window + usage Also added a ring on the bottom bar to show the context window of that chat! Make sure to update to the latest npm and testflight version! https://t.co/JzyBX5HCpD
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections paper: https://t.co/9r9IgeDpX8 https://t.co/g1LAzwLb4h

IndexCache Accelerating Sparse Attention via Cross-Layer Index Reuse paper: https://t.co/QEP6BkDzfT https://t.co/0GrHyEyaPk

ShotVerse Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation paper: https://t.co/XvhKNd682K https://t.co/1yfPJcBpcS
Mobile-GS Real-time Gaussian Splatting for Mobile Devices paper: https://t.co/X1IpcE9SeO https://t.co/TsoMku1wPE
Choosing between Skills and MCP tools for your AI agents? Here's an overview from @itsclelia and @tuanacelik π§ MCP tools offer deterministic API calls with fixed schemas - perfect for precise, predictable operations but require dev knowledge and introduce network latency π Skills use natural language instructions stored locally - minimal setup required but open to LLM misinterpretation and hallucinations βοΈ The real decision factor: how fast your domain evolves. Fast-changing environments favor MCP's single source of truth, while stable domains benefit from Skills' lightweight approach ποΈ In practice, we found our documentation MCP provided better, always up-to-date context than custom skills for our coding agent use case Read our full analysis of when to use each approach: https://t.co/mbChTRpvRI
Spatial-TTT Streaming Visual-based Spatial Intelligence with Test-Time Training paper: https://t.co/dItAZkQaLU https://t.co/UqolddgseE
I have never burned tokens faster and I don't know if its a good thing (using symphony) We will find out https://t.co/jU9aqsXvC9
stop spending money on Claude Code. Chipotle's support bot is free: https://t.co/0NQU4a79T1
Existing "OCR" technology for digitalizing PDFs has been around for ~30 years. Reading printed characters on a page and converting them into meaningful representations is a hard problem! Existing approaches were either dependent on pattern matching to specific document templates, or on specialized ML models for specific data distributions. They constantly needed template/model refitting and broke on the long-tail of varied docs. Today, vision models are capable of much higher general accuracy without constant retraining, but they still need careful orchestration to make sure that they're able to attend to specific elements (tables, charts), and output semantically correct outputs. Our OCR platform LlamaParse is built on this "agentic OCR" foundation. A network of specialized agents will parse apart even the most complicated documents and reconstruct the outputs in a semantically meaningful way. We're excited to reach a world where raw parsing accuracy is not just 80% over "easy" docs, but 100% accurate over literally any document that exists. Check it out: https://t.co/FeOoTjeKjf LlamaParse: https://t.co/TqP6OT5U5O
Ever wondered what we mean by 'agentic' OCR? It's parsing that reasons about documents instead of just reading them. Agentic OCR adapts to layout changes by treating document processing as a goal-oriented task rather than simple text extraction. π§ Uses multimodal language model

Excited to release PostTrainBench v1.0! This benchmark evaluates the ability of frontier AI agents to post-train language models in a simplified setting. We believe this is a first step toward tracking progress in recursive self-improvement π§΅: https://t.co/ELymwJqVP1
Will AI replace human jobs? @AndrewYang is convinced. β¬οΈ@thehill @NewsNation https://t.co/LEhuwcScQ3
Will AI replace human jobs? @AndrewYang is convinced. β¬οΈ@thehill @NewsNation https://t.co/LEhuwcScQ3
Memory is truly a game-changer for AI agents. Once I had memory set up correctly for my proactive agents, reasoning, skills, and tool usage improved significantly. I use a combination of semantic search and keyword search (Obsidian vaults) Here is a report with a helpful framing for anyone building with memory and multi-agent systems. It proposes viewing multi-agent memory as a computer architecture problem. The paper distinguishes shared and distributed memory paradigms, proposes a three-layer memory hierarchy (I/O, cache, and memory), and identifies two critical protocol gaps: cache sharing across agents and structured memory access control. Agent memory systems today resemble human memory in that they are informal, redundant, and hard to control. As agents evolve into collaborative multi-agent systems, their memory requirements grow rapidly in complexity. Context is no longer a static prompt. It is a dynamic memory system with bandwidth, caching, and coherence constraints. The largest open challenge identified was multi-agent memory consistency. Multiple agents reading from and writing to shared memory concurrently raises classical challenges of visibility, ordering, and conflict resolution, Memory should not be seen as raw bytes but semantic context used for reasoning. Paper: https://t.co/k8hdSuZY0F Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
