Your curated collection of saved posts and media
Yeah okay, Lego bros, brodettes and brotheys are cooked with this one. GPT 2 Image can create full Lego sets! With actual Bricklink IDs so you can order the parts and build it. Whole new business opportunity here for the taking. https://t.co/4d4aJeEGTU

Side quest: trying to write custom firmware to a DVD drive and get low-level hardware control. This has been a long-term AI bench of mine. Not quite there yet, but 5.5 made a ton of progress - it feels like we're quite close to the fun + scary world where any sufficiently complex device can probably be reverse-engineered and hacked with a bit of effort! Rough write-up of the current state: https://t.co/yIbRLMDstA
ๆใใใฆDVDใฎใขใฌใๅ็พใใใใทใณไฝใฃใฆใใ้ฃไผ็ตใใฃใ.. ่งใซๅฝใใใจ็่ฃใฌใคใณใใผใซๅ ใ https://t.co/QlTobdhEFj
ๆใใใฆDVDใฎใขใฌใๅ็พใใใใทใณไฝใฃใฆใใ้ฃไผ็ตใใฃใ.. ่งใซๅฝใใใจ็่ฃใฌใคใณใใผใซๅ ใ https://t.co/QlTobdhEFj
Microbial traffic jam. (Ciliates in biofilm, seawater sample from the Oregon coast, polarized light microscopy) https://t.co/ylDP9RyTeF
More tests with duckweed transformation and regeneration https://t.co/8gkCsuQozp

Mildly obsessed with this bacteria growing on an old plate (cc @ContamClub ) https://t.co/lW2XDVSHJ6

New EQ-Bench results! Opus 4.7: clean sweep, still the king. Deepseek 4: very strong, near frontier on EQ-Bench & longform writing. Kimi k2.6: Strong in shortform but seems to suffer degradation in longform writing. GPT-5.5: performs ~identically to GPT-5.4. https://t.co/slwFfDoBVj

Newsletter update of last 2 weeks of AI/game/web/narrative links of interest... more splats and 360ies, web games, narrative and creativity, vis of Aaron Reed's Subcutanean procgen novel, and more! 1/2 https://t.co/vaHQM9gO48
@AhmedShahnab Some country files are being screwed by overseas territories https://t.co/8RRmMpROkO
Iโm looking for the first full-time developer for https://t.co/nprGVKUWBn! Youโd work with me on making new web projects and games for the site. Preferably in nyc If anyone is interested or has leads dm me or email hi@neal.fun! https://t.co/0bOvzn7UVc
๐จNew preprint! We find evidence of LLMs enabling people to file lawsuits without lawyers (filing "pro se") at historically unprecedented rates in federal courts.๐ 1/n https://t.co/JCj8oq5Jym
This meditation app is was invented, designed and coded, and then submitted to the app store by an AI model (it made a few mistakes along the way). In a new research paper, @sayashk and others say having AI take on this kind of messy open world task could offer a better way to measure progress. Very interesting! (paper https://t.co/ki0ymYzGkT) (app https://t.co/jbWeNcGMma)

AI evaluation is becoming its own compute bottleneck. We often talk about the cost of training frontier models, but the cost of evaluating them is starting to matter just as much, especially for agents, scientific ML systems, and training-in-the-loop benchmarks. In our new Evaluating Evaluations post, we look at how evals are crossing a threshold where cost changes who can participate. The Holistic Agent Leaderboard spent about $40K on 21,730 agent rollouts across 9 models and 9 benchmarks. A single GAIA run on a frontier model can cost $2,829 before caching. And once you care about reliability, repeated runs can multiply these costs many times over. This creates a real accountability problem. If only large labs can afford statistically credible evals, independent researchers, auditors, journalists, and public-interest organizations are left with partial visibility into frontier systems. The core issue is that benchmark design is changing. Static benchmarks could often be compressed aggressively while preserving rankings. Agent benchmarks are noisier and scaffold-sensitive. Training-in-the-loop benchmarks are expensive by construction. As evals move closer to real work, they also become harder to make cheap. Some takeaways: โ Leaderboards should report cost alongside accuracy. โ Reliability should not be treated as optional. โ We need reusable eval artifacts! Shared documentation formats, such as Every Eval Ever, can help the field stop paying repeatedly for the same measurements. Read the full post: https://t.co/sArlZMkytF Thanks for the insights @LChoshen , Yifan Mai, and @cgeorgiaw๐ค

Imagine a 19-year-old scrolling TikTok. She watches a creator list five "signs you have undiagnosed anxiety." She recognizes three in herself. By the end of the week, she's describing herself as anxious to her friends. A month later, she's avoiding situations she used to handle fine. What went wrong? In a new paper by my PhD student Dasha Sandra, titled "Why mental health awareness can harm: Converging explanations for a societal problem", we argue that well-meaning mental health awareness can backfire, and we identify how. Four separate literatures (concept creep, nocebo effects, prevalence inflation, and illness self-labeling) have been circling the same problem from different angles. We show they converge on three mechanisms: 1.Awareness lowers the threshold for what counts as a disorder. 2. It trains people to scan their inner lives for symptoms and reinterpret normal distress as pathology. 3. Once someone adopts an illness identity, they behave in ways that confirm and deepen it. The evidence is wide. Learning that loneliness is harmful makes solitude feel worse. Learning that stress is harmful worsens well-being and performance. Awareness videos about fake conditions like "wind turbine syndrome" produce real headaches. Trigger warnings raise anticipatory anxiety without reducing distress. This does not mean awareness should stop. It means awareness can have unintended consequences, including manufacturing the suffering it tries to prevent. Inoculating people against these mechanisms works, and we already have evidence it does. Link to paper: https://t.co/ucoGyhEuAj
1800: If Thomas Jefferson is elected "Murder, robbery, rape, adultery, and incest will all be openly taught and practiced." When it comes to politics we have a bad habit of romanticizing the past and imagining that today's politics are worse and coarser. To make this visceral, I built a little app that shows what the 1800 election would have felt like if X had been around. Scrolling through it really does give you a sense that vicious, indecorous politics long pre-dates present day. Check it out here: https://t.co/eo2vFOf6TF
Experts have three views on the future of work, each credible but sharply opposed. Whoโs right? In a new paper for @CarnegieEndow & @CEIPTechProgram, I lay out the best arguments made by the alarmed, patient, and excited groups. ๐งตOn the most important points and what policymakers can do today
A recent viral paper claims to reverse-engineer the parameter counts of frontier models: GPT-5.5 = 9.7T, Opus 4.7 = 4.0T, o1 = 3.5T, etc. @ben_sturgeon and I investigated and found serious issues in the paper; fixing them gives GPT-5.5 as ~1.5T (90% CI: 256B-8.3T). https://t.co/MbWQyVlmsE
Some people are being way too alarmist about Mythos. 80,000 Hours called Mythos "an AI that can break into almost any computer on Earth". Zvi Mowshowitz said, "If given to anyone with a credit card, Claude Mythos would give attackers a cornucopia of zero-day exploits for essentially all the software on Earth". But these descriptions are unfounded. @natalia__coelho shows in her latest blog post that the cyber capabilities of Mythos are nearly tied with GPT-5.5, across practically every public benchmark we have available. This includes both narrow and broad cyber evaluations. It is not way ahead of trend. The same holds for general capability evaluations. Except for a somewhat impressive score on SWE-bench Pro, Claude Mythos is nearly on-trend, at most a few months ahead. In other words, it's barely better than models that millions of people already have access to. It's an impressive model, but I'm very skeptical that Mythos is going to take down our digital infrastructure or cause a cyber catastrophe.

Excited to give this talk at the Stanford Digital Economy Lab on May 18! I will do three things: discuss my group's recent research, identify the most pressing gaps in the community's current understanding, and provide a long-term perspective. Hope to see you there in person or virtually. https://t.co/Qa2eNkVsnZ @DigEconLab
We furthered AI research by reproducing CRUX #1 for Windows using @getnenai's infrastructure without needing to buy a Windows machine-- checkout our blog post https://t.co/Mjz1D5myF9
Benchmarks are saturated more quickly than ever. How should frontier AI evaluations evolve? In a new paper, we argue that the AI community is already converging on an answer: Open-world evaluations. They are long, messy, real-world tasks that would be impractical for benchmarks.
We benchmarked GPT-5.5 on document understanding ๐๐ We ran it through ParseBench, our comprehensive OCR benchmark over enterprise documents. We evaluated metrics across various dimensions: visual grounding, tables, charts, and more. We evaluated GPT-5.5 on mid thinking and zero-thinking modes. When compared against GPT-5.4 (0 thinking) and Opus 4.7 (adaptive thinking): ๐ GPT-5.5 wins on tables ๐ GPT-5.5 wins on visual grounding ๐ GPT-5.5 0-thinking does worse on charts than GPT-5.4 0-thinking ๐ Higher thinking does worse than lower thinking of content faithfulness, semantic formatting ๐ Opus 4.7 wins overall on content faithfulness and semantic formatting ๐ธ GPT-5.5 is expensive: 13c per page at mid-thinking modes and 5.93 at 0-thinking! This is 5x the cost of any competitive OCR solution. Conclusion: GPT-5.5 is one of the better frontier models out there in terms of pure accuracy, but def not pound for pound w.r.t price.
Loan processors spend 40โ60% of their time reconciling income across tax returns, pay stubs, W-2s, and bank statements. We built an end-to-end pipeline that automates it with LlamaParse + the Claude Agent SDK: ๐ Schema-driven extraction across 4 doc types with confidence scores + citations ๐ Cross-document validation with Claude โ catches W-2/pay-stub gaps, unexplained Zelle/Venmo deposits, employer name mismatches ๐ Self-contained HTML report with a COMPLETE / REVIEW / FLAG decision Full code + walkthrough: https://t.co/ozm4VwWmj3
Let's talk document formatting. Bold. Italics. Superscripts. Strikethroughs. The visual cues humans rely on every time we read a doc, and ones existing OCR benchmarks completely ignore. ๐ฑ"$199" struck through next to "$149" isn't decoration. It's the meaning. ๐ฑA superscript tells your agent "3" is a citation, not part of the number. Flatten that and your agent is reading a different doc than you are. Two weeks ago we released ParseBench, the first document OCR benchmark for AI agents. One of five metrics: the Semantic Formatting Score. Read more๐ https://t.co/2sq5ncGiel
Parsing documents with AI agents just got a lot more seamless๐ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ Once connected, you'll be able to: ๐ Parse documents into clean markdown ๐ Classify files against your own categories โ๏ธ Split long documents into labelled sections โฌ๏ธ Upload files via URL or a browser-based upload flow Building a production MCP server surfaced some non-obvious challenges: getting auth to align with an existing platform identity system using @WorkOS, working around MCP's lack of built-in file upload support, and making deployments, rate limiting and observability feel native with @vercel and @AxiomFM. We wrote up all of it, from the OAuth flow, to the token-based upload design, to the tradeoffs we hit along the way๐ ๐ Read the full blog: https://t.co/2E3qIVeUYI ๐ฉโ๐ป GitHub repository: https://t.co/ru1Evj7Zsr
We shipped a LlamaParse MCP server to let you parse, classify, split, and generally operate over your hardest documents with your favorite AI agent ๐๐ค Check out the MCP server: https://t.co/eOeDgzOCI5 This is both a useful feature and a great piece of engineering that we want to share with the community. ๐ก MCP does not have built-in file upload support, so we needed to implement a URL-based upload endpoint and couple it with parse operations ๐กWe built in an integration with @WorkOS OAuth. ๐กWe built in observability and rate-limiting Huge shoutout to @itsclelia for shipping this! Come check out our blog writeup: https://t.co/XsCZlKnzez LlamaParse: https://t.co/TqP6OT5U5O
Parsing documents with AI agents just got a lot more seamless๐ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ Once connected, you'll be able to: ๐ Pars
Building scalable, distributed document processing pipelines isnโt easy. Thatโs why we teamed up with @render to build a system that: ๐ Leverages the LlamaParse platform to parse, classify, extract, and retrieve information from documents โ๏ธ Uses Render Workflows to distribute tasks across nodes and accelerate background processing โก Deploys a lightweight server and database on Render, giving you an instant interface to interact with your pipeline ๐ฉโ๐ป Explore the repo to see it in action: https://t.co/eiJqklNVhj ๐ And check out the step-by-step breakdown by @ojusave and @itsclelia: https://t.co/gePmuL7YtV
Thank you AI Dev Day '26 @DeepLearningAI @jerryjliu0 shares why SOTA LLMs can build an app but can't read a PDF ๐คฏ https://t.co/L2WMAUVaJ4

Parsing PDFs is hard This past week I gave a few talks (at both AI Dev '26 by @DeepLearningAI and @Capgemini ) on why this is still such an open problem, and itโs even more important as agents become the consumers of documents, and need the OCR tools to read them properly. The fundamental issue is that PDFs are designed for print and display purposes, not to give back a linearized, semantically meaningful string of text. Text and tables are represented as a bunch of chars and lines, without any guaranteed order. This is what the community is solving with VLM-based approaches, including our own efforts around LlamaParse and ParseBench. If youโre interested in learning more about the problem, check out the blog post I wrote on this a while ago! https://t.co/740ZiAFyOk

๐ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

Iโm excited to announce that @llama_index is on the @CBInsights AI 100 list for 2026 ๐ฅ Weโre on a mission to parse all of the worldโs PDFs, and make them accessible to both humans and AI agents. List: https://t.co/8pKDaPnHI8 If you havenโt done so already, our website design is awesome, check it out: https://t.co/YiIfjVlzb6
๐ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powered by the LlamaParse TypeScript SDK ๐ฑ Three steps, thatโs it: ๐ Add your API key (securely stored on-device) ๐ธ Snap a photo of anything with text ๐ Parse it and, in under a minute, get clean, copyable text No hassle, no manual typing. ๐ Try it now: https://t.co/mEeraV4zap ๐ฆ Get started with LlamaParse: https://t.co/Pmc8usyjTwโ