Your curated collection of saved posts and media
@AhmedShahnab Some country files are being screwed by overseas territories https://t.co/8RRmMpROkO
Iโm looking for the first full-time developer for https://t.co/nprGVKUWBn! Youโd work with me on making new web projects and games for the site. Preferably in nyc If anyone is interested or has leads dm me or email hi@neal.fun! https://t.co/0bOvzn7UVc
๐จNew preprint! We find evidence of LLMs enabling people to file lawsuits without lawyers (filing "pro se") at historically unprecedented rates in federal courts.๐ 1/n https://t.co/JCj8oq5Jym
This meditation app is was invented, designed and coded, and then submitted to the app store by an AI model (it made a few mistakes along the way). In a new research paper, @sayashk and others say having AI take on this kind of messy open world task could offer a better way to measure progress. Very interesting! (paper https://t.co/ki0ymYzGkT) (app https://t.co/jbWeNcGMma)

AI evaluation is becoming its own compute bottleneck. We often talk about the cost of training frontier models, but the cost of evaluating them is starting to matter just as much, especially for agents, scientific ML systems, and training-in-the-loop benchmarks. In our new Evaluating Evaluations post, we look at how evals are crossing a threshold where cost changes who can participate. The Holistic Agent Leaderboard spent about $40K on 21,730 agent rollouts across 9 models and 9 benchmarks. A single GAIA run on a frontier model can cost $2,829 before caching. And once you care about reliability, repeated runs can multiply these costs many times over. This creates a real accountability problem. If only large labs can afford statistically credible evals, independent researchers, auditors, journalists, and public-interest organizations are left with partial visibility into frontier systems. The core issue is that benchmark design is changing. Static benchmarks could often be compressed aggressively while preserving rankings. Agent benchmarks are noisier and scaffold-sensitive. Training-in-the-loop benchmarks are expensive by construction. As evals move closer to real work, they also become harder to make cheap. Some takeaways: โ Leaderboards should report cost alongside accuracy. โ Reliability should not be treated as optional. โ We need reusable eval artifacts! Shared documentation formats, such as Every Eval Ever, can help the field stop paying repeatedly for the same measurements. Read the full post: https://t.co/sArlZMkytF Thanks for the insights @LChoshen , Yifan Mai, and @cgeorgiaw๐ค

Imagine a 19-year-old scrolling TikTok. She watches a creator list five "signs you have undiagnosed anxiety." She recognizes three in herself. By the end of the week, she's describing herself as anxious to her friends. A month later, she's avoiding situations she used to handle fine. What went wrong? In a new paper by my PhD student Dasha Sandra, titled "Why mental health awareness can harm: Converging explanations for a societal problem", we argue that well-meaning mental health awareness can backfire, and we identify how. Four separate literatures (concept creep, nocebo effects, prevalence inflation, and illness self-labeling) have been circling the same problem from different angles. We show they converge on three mechanisms: 1.Awareness lowers the threshold for what counts as a disorder. 2. It trains people to scan their inner lives for symptoms and reinterpret normal distress as pathology. 3. Once someone adopts an illness identity, they behave in ways that confirm and deepen it. The evidence is wide. Learning that loneliness is harmful makes solitude feel worse. Learning that stress is harmful worsens well-being and performance. Awareness videos about fake conditions like "wind turbine syndrome" produce real headaches. Trigger warnings raise anticipatory anxiety without reducing distress. This does not mean awareness should stop. It means awareness can have unintended consequences, including manufacturing the suffering it tries to prevent. Inoculating people against these mechanisms works, and we already have evidence it does. Link to paper: https://t.co/ucoGyhEuAj
1800: If Thomas Jefferson is elected "Murder, robbery, rape, adultery, and incest will all be openly taught and practiced." When it comes to politics we have a bad habit of romanticizing the past and imagining that today's politics are worse and coarser. To make this visceral, I built a little app that shows what the 1800 election would have felt like if X had been around. Scrolling through it really does give you a sense that vicious, indecorous politics long pre-dates present day. Check it out here: https://t.co/eo2vFOf6TF
Experts have three views on the future of work, each credible but sharply opposed. Whoโs right? In a new paper for @CarnegieEndow & @CEIPTechProgram, I lay out the best arguments made by the alarmed, patient, and excited groups. ๐งตOn the most important points and what policymakers can do today
A recent viral paper claims to reverse-engineer the parameter counts of frontier models: GPT-5.5 = 9.7T, Opus 4.7 = 4.0T, o1 = 3.5T, etc. @ben_sturgeon and I investigated and found serious issues in the paper; fixing them gives GPT-5.5 as ~1.5T (90% CI: 256B-8.3T). https://t.co/MbWQyVlmsE
Some people are being way too alarmist about Mythos. 80,000 Hours called Mythos "an AI that can break into almost any computer on Earth". Zvi Mowshowitz said, "If given to anyone with a credit card, Claude Mythos would give attackers a cornucopia of zero-day exploits for essentially all the software on Earth". But these descriptions are unfounded. @natalia__coelho shows in her latest blog post that the cyber capabilities of Mythos are nearly tied with GPT-5.5, across practically every public benchmark we have available. This includes both narrow and broad cyber evaluations. It is not way ahead of trend. The same holds for general capability evaluations. Except for a somewhat impressive score on SWE-bench Pro, Claude Mythos is nearly on-trend, at most a few months ahead. In other words, it's barely better than models that millions of people already have access to. It's an impressive model, but I'm very skeptical that Mythos is going to take down our digital infrastructure or cause a cyber catastrophe.

Excited to give this talk at the Stanford Digital Economy Lab on May 18! I will do three things: discuss my group's recent research, identify the most pressing gaps in the community's current understanding, and provide a long-term perspective. Hope to see you there in person or virtually. https://t.co/Qa2eNkVsnZ @DigEconLab
We furthered AI research by reproducing CRUX #1 for Windows using @getnenai's infrastructure without needing to buy a Windows machine-- checkout our blog post https://t.co/Mjz1D5myF9
Benchmarks are saturated more quickly than ever. How should frontier AI evaluations evolve? In a new paper, we argue that the AI community is already converging on an answer: Open-world evaluations. They are long, messy, real-world tasks that would be impractical for benchmarks.
We benchmarked GPT-5.5 on document understanding ๐๐ We ran it through ParseBench, our comprehensive OCR benchmark over enterprise documents. We evaluated metrics across various dimensions: visual grounding, tables, charts, and more. We evaluated GPT-5.5 on mid thinking and zero-thinking modes. When compared against GPT-5.4 (0 thinking) and Opus 4.7 (adaptive thinking): ๐ GPT-5.5 wins on tables ๐ GPT-5.5 wins on visual grounding ๐ GPT-5.5 0-thinking does worse on charts than GPT-5.4 0-thinking ๐ Higher thinking does worse than lower thinking of content faithfulness, semantic formatting ๐ Opus 4.7 wins overall on content faithfulness and semantic formatting ๐ธ GPT-5.5 is expensive: 13c per page at mid-thinking modes and 5.93 at 0-thinking! This is 5x the cost of any competitive OCR solution. Conclusion: GPT-5.5 is one of the better frontier models out there in terms of pure accuracy, but def not pound for pound w.r.t price.
Loan processors spend 40โ60% of their time reconciling income across tax returns, pay stubs, W-2s, and bank statements. We built an end-to-end pipeline that automates it with LlamaParse + the Claude Agent SDK: ๐ Schema-driven extraction across 4 doc types with confidence scores + citations ๐ Cross-document validation with Claude โ catches W-2/pay-stub gaps, unexplained Zelle/Venmo deposits, employer name mismatches ๐ Self-contained HTML report with a COMPLETE / REVIEW / FLAG decision Full code + walkthrough: https://t.co/ozm4VwWmj3
Let's talk document formatting. Bold. Italics. Superscripts. Strikethroughs. The visual cues humans rely on every time we read a doc, and ones existing OCR benchmarks completely ignore. ๐ฑ"$199" struck through next to "$149" isn't decoration. It's the meaning. ๐ฑA superscript tells your agent "3" is a citation, not part of the number. Flatten that and your agent is reading a different doc than you are. Two weeks ago we released ParseBench, the first document OCR benchmark for AI agents. One of five metrics: the Semantic Formatting Score. Read more๐ https://t.co/2sq5ncGiel
Parsing documents with AI agents just got a lot more seamless๐ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ Once connected, you'll be able to: ๐ Parse documents into clean markdown ๐ Classify files against your own categories โ๏ธ Split long documents into labelled sections โฌ๏ธ Upload files via URL or a browser-based upload flow Building a production MCP server surfaced some non-obvious challenges: getting auth to align with an existing platform identity system using @WorkOS, working around MCP's lack of built-in file upload support, and making deployments, rate limiting and observability feel native with @vercel and @AxiomFM. We wrote up all of it, from the OAuth flow, to the token-based upload design, to the tradeoffs we hit along the way๐ ๐ Read the full blog: https://t.co/2E3qIVeUYI ๐ฉโ๐ป GitHub repository: https://t.co/ru1Evj7Zsr
We shipped a LlamaParse MCP server to let you parse, classify, split, and generally operate over your hardest documents with your favorite AI agent ๐๐ค Check out the MCP server: https://t.co/eOeDgzOCI5 This is both a useful feature and a great piece of engineering that we want to share with the community. ๐ก MCP does not have built-in file upload support, so we needed to implement a URL-based upload endpoint and couple it with parse operations ๐กWe built in an integration with @WorkOS OAuth. ๐กWe built in observability and rate-limiting Huge shoutout to @itsclelia for shipping this! Come check out our blog writeup: https://t.co/XsCZlKnzez LlamaParse: https://t.co/TqP6OT5U5O
Parsing documents with AI agents just got a lot more seamless๐ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ Once connected, you'll be able to: ๐ Pars
Building scalable, distributed document processing pipelines isnโt easy. Thatโs why we teamed up with @render to build a system that: ๐ Leverages the LlamaParse platform to parse, classify, extract, and retrieve information from documents โ๏ธ Uses Render Workflows to distribute tasks across nodes and accelerate background processing โก Deploys a lightweight server and database on Render, giving you an instant interface to interact with your pipeline ๐ฉโ๐ป Explore the repo to see it in action: https://t.co/eiJqklNVhj ๐ And check out the step-by-step breakdown by @ojusave and @itsclelia: https://t.co/gePmuL7YtV
Thank you AI Dev Day '26 @DeepLearningAI @jerryjliu0 shares why SOTA LLMs can build an app but can't read a PDF ๐คฏ https://t.co/L2WMAUVaJ4

Parsing PDFs is hard This past week I gave a few talks (at both AI Dev '26 by @DeepLearningAI and @Capgemini ) on why this is still such an open problem, and itโs even more important as agents become the consumers of documents, and need the OCR tools to read them properly. The fundamental issue is that PDFs are designed for print and display purposes, not to give back a linearized, semantically meaningful string of text. Text and tables are represented as a bunch of chars and lines, without any guaranteed order. This is what the community is solving with VLM-based approaches, including our own efforts around LlamaParse and ParseBench. If youโre interested in learning more about the problem, check out the blog post I wrote on this a while ago! https://t.co/740ZiAFyOk

๐ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

Iโm excited to announce that @llama_index is on the @CBInsights AI 100 list for 2026 ๐ฅ Weโre on a mission to parse all of the worldโs PDFs, and make them accessible to both humans and AI agents. List: https://t.co/8pKDaPnHI8 If you havenโt done so already, our website design is awesome, check it out: https://t.co/YiIfjVlzb6
๐ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powered by the LlamaParse TypeScript SDK ๐ฑ Three steps, thatโs it: ๐ Add your API key (securely stored on-device) ๐ธ Snap a photo of anything with text ๐ Parse it and, in under a minute, get clean, copyable text No hassle, no manual typing. ๐ Try it now: https://t.co/mEeraV4zap ๐ฆ Get started with LlamaParse: https://t.co/Pmc8usyjTwโ
I โค๏ธ NYC We're hosting two in-person events next Wednesday: 1๏ธโฃ FinParse workshop: Build AI agents to extract and act over the most complex financial documents 2๏ธโฃ AI Happy Hour with Tabs: drinks, conversations, and L'Industrie Pizza ๐ Workshop: https://t.co/RQXRxxAwFk Happy hour: https://t.co/8mlnOYpAgg Come check it out!
LlamaIndex NYC takeover, 5/13 ๐ฝ Our CEO Jerry Liu is in town. Two events, open to every NYC builder: ๐ ๏ธ FinParse Workshop โ laptops out, hands-on with @jerryjliu0 โ https://t.co/MeaWrBU2ca ๐ AI Engineers on Tap โ happy hour w/ @tabs โ https://t.co/8FQTwFLFk1
Last week I gave a talk at AI Dev โ26 by @DeepLearningAI on โAI canโt read PDFs, how do we fix itโ . Iโm sharing the slides publicly if others are interested in doing a deep dive into document understanding. AI agents are going to automate huge amounts of knowledge work, but knowledge work depends on data, a lot of that data is in documents/PDFs, and existing OCR tools suck. PDFs are a format that is inherently hard to read, and I dive into specific reasons why (tables, layouts), and why frontier VLMs and benchmarks are still insufficient. Even as agents get better and more general, they need the right tools to read and act over PDFs. They both need this at the data ingest layer, as well as tools they can call on the fly. Check out the slides: https://t.co/IyaTlQMPlW Weโre building high-quality AI document processing, both with LlamaParse, along with OSS efforts like LiteParse and ParseBench. If you have a ton of PDFs that youโre hoping to unlock with AI, come talk to us! https://t.co/TsPLwqZ8yg

A few weeks ago @simonw got Claude to port LiteParse to the browser. Today, we are launching that work as a complete guide in our docs! https://t.co/sqmrVxx3YU The guide itself relies on some fun hacks with vite and mocking. We expect this process to improve with future releases, so stay tuned!
Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet ๐๐ฎ๐ป๐ฑ๐ฏ๐ผ๐ ๐ฒ๐ฑ-๐น๐ถ๐, a Rust ๐ฆ CLI agent that combines: - LiteParse, our lightning-fast local parser for PDFs, images, Office files, and more - A secure sandbox powered by @microsandbox - Full filesystem mounting, so your agent can safely interact with local files inside the sandbox Mount your local workspace, give the agent shell access, and let it do its magic ๐ช ๐ฉโ๐ป GitHub: https://t.co/3NpXcFw58p ๐ Learn more about LiteParse: https://t.co/XoAIrbPtE4โ
GDP.pdf was accepted to the CVPR 2026 Workshop on Multimodal Reasoning can frontier models handle the three-letter document type that runs the world? we partnered with hundreds of expert surgers - ER physicians, construction engineers, corporate litigators - to find out. every one scored under 15%. paper, leaderboard, and dataset below. paper: https://t.co/y4hmPXbpx4 dataset: https://t.co/rElrfOntcU leaderboard: https://t.co/9CMY6JVhDL blog: https://t.co/0Wj97DBr44
As I mentioned before, I am now sharing an example from GPT-5.5 Pro, also featured by OpenAI, that really left me stunned by what it is capable of in biomedical science. (full report on the website I created with Codex, link in the thread). To push GPT-5.5 Pro hard, I uploaded a real data set of immune subset (T cells) gene-expression spreadsheet: 62 sorted T cell samples, 27,906 gene columns, and millions of underlying data points across different T cell subsets. Importantly, this public dataset also had paired structure making it possible to separate true cell-state biology from donor-to-donor variation. I asked GPT-5.5 Pro not merely to summarize the spreadsheet, but to analyze it deeply: What can we learn from this dataset? What are the mechanistic insights? What are the most important biological questions that emerge? What follow-up experiments should we do next? It thought for about 100 minutes and produced a roughly 40-page report! What amazed me was not just the length or even the initial analysis, since previous models are also capable of doing this. What amazed me was the quality of the reasoning and insights it provided! The report recognized that this was not just a table of genes, but two overlapping experimental designs. It identified the major biological axis, which in plain language was that the cells were not just โdifferent categories.โ They formed a coherent differentiation landscape, moving from future potential toward immediate function. It also understood the caveats. It did not overclaim from bulk gene-expression data. It clearly explained that bulk transcriptomics cannot distinguish whether every cell in a sorted population has shifted or whether a smaller subpopulation is dominating the signal. It recommended the right next steps experiments, and integration with donor metadata. This is what made the report feel so special to me. It was not just doing statistics. It was reasoning like an expert systems immunologist. It saw the structure of the experiment, interpreted the patterns, built a mechanistic model, identified limitations, proposed causal hypotheses, and laid out a translational roadmap. Other advanced models have been able to generate excellent biomedical reports before, including previous GPT-5 models. So I don't want to claim this is an entirely new type of capability. But this one felt different in an important way. It had more scientific elegance, more restraint, more biological intuition, and more of the nuanced judgment that usually comes only from years of hands-on experience in the field. It felt like this AI model had crossed another threshold. This is the kind of analysis that could easily take a research team months to perform, refine, interpret, and write up. Even then, many teams might not produce something this integrated, this mechanistically coherent, and this useful as a launchpad for future experiments. I know a 40-page T-cell gene-expression analysis may not be exciting to everyone. To illustrate how good it is, also had Codex built a web site with it anyone can explore, link below. ๐ Those interested can go deeper into the report. I also wanted this example on the record because, because to me, it is evidence that we are entering a new stage in AI-assisted biomedical science. The important point is no longer that AI can "analyze data and write a report.โ The important point is that AI can now help transform complex biological data into mechanistic understanding, experimental priorities, and testable hypotheses at a speed and depth that would have been almost unimaginable a short time ago. For biomedical science, this is a very big deal! Of course, this may vary across domains, and every analysis still needs expert review, validation, and experimental follow-up. But in my own field, with data I understand deeply, this felt like another inflection point. I feel strongly that we have crossed another milestone threshold in the age of AI, with the release of GPT-5.5.
.@pmarca says AI is the biggest technological revolution of his life: "This is the biggest technological revolution of my life. This is clearly bigger than the internet. The comps on this are things like the microprocessor, the steam engine, and electricity." "The neural network as an idea continued to be explored in academia for the last 80 years. And essentially it didn't work, it was decade after decade of excessive optimism, followed by disappointment." "Then basically we all saw what happened with the ChatGPT moment. All of a sudden it crystallized, and it was like, 'Oh my God, it turns out it works.' We're sort of three years into effectively an 80-year revolution."
What if you could test a drug on a human organ and simultaneously know what would have happened without it? That's what we built. Proud to share our work in @NatureBiotech: digital twins of human lungs ๐ซ๐ฅ (paper: https://t.co/Y8wzxNwWfd) We created digital twins of ex vivo human lungs: multimodal AI models trained on 951 human lungs from the world's largest EVLP dataset at @UHN. Physics-informed ML across physiology, biochemistry, transcriptomics, proteomics, metabolomics, and imaging, all forecasting together. The key insight: the physical lung receives the treatment. The twin is the untreated control. Paired causal inference on the same organ. No separate cohort. No intersubject noise. Result: we detected drug efficacy with 6 lungs. Traditional methods would need 18. This is what precision preclinical evaluation looks like. From Virtual Cells โ Virtual Organs โ Virtual Patients. One step toward virtual organs replacing animal testing. Huge congratulations to Elly Zhou (a phd student I co-supervise) for leading this work with exceptional rigor, and to Andrew Sage and @SKeshavjee for building the foundation that made it possible.

Our paper โABD: DefaultโException Abduction in Finite First-Order Worldsโ was accepted to KR 2026, in the KR Meets Machine Learning and Explanation track. arXiv: https://t.co/HYb5icwsel code/data: https://t.co/ZXXqzouuni A very nice plain-language summary: https://t.co/Zl5lHu4EGH Motivation: many reasoning benchmarks ask whether a model gets an answer right, but give only a weak view of *why* it succeeds or fails. In ABD, the model has to synthesize an explicit first-order formula explaining the exceptions to a default rule. A typical task is not: โclassify this object.โ It is: given a background first-order theory, and several finite relational worlds, output one formula ฮฑ(x) that defines when an object is abnormal (i.e., does not satisfy the background theory), while keeping the exception set sparse. That gives us three things that are hard to get from many natural-language benchmarks: 1. exact semantic checking; 2. an explicit hypothesis produced by the model; 3. a cost measure, so we can distinguish โvalid but bloatedโ explanations from compact ones. The benchmark has three observation regimes: โข Full / closed-world: the observed facts are complete. โข Partial / existential completion: unknowns may be completed favorably. โข Skeptical / universal completion: the explanation must work under all completions. So the same high-level task can be tested under increasingly demanding assumptions about incomplete information. A central design choice is that each instance contains multiple prompt worlds. The model must output one shared ฮฑ(x) that works across all of them jointly. This is meant to test rule synthesis, not per-world labeling. We also evaluate on matched holdout worlds. This lets us separate two failure modes: formulas that stop being valid on fresh worlds, and formulas that remain valid but become too expensive i.e., call too many objects "abnormal". Empirically, evaluating ten frontier LLMs on 600 instances, the best models often achieve high validity, but parsimony remains a real issue. Many hypotheses repair the theory but use unnecessarily broad exception definitions. Holdout evaluation also shows distinct generalization failures across the three observation regimes. ABD aims to test one specific type of reasoning, formulating simple rules (first-order formulas) that explain exceptions to a given background theory. Finite first-order worlds give us a clean testbed for doing so, by being explicitly solvable and feasible for generating problems automatically. I hope it is useful both for evaluating language models and for comparing them with more symbolic or solver-aided synthesis methods.
