Your curated collection of saved posts and media
// The Consistency Illusion // Multi-agent debate can make agents agree on the final answer while their underlying reasoning stays misaligned. This work finds that consensus on the output hides disagreement on the path that produced it, and you only ever see the output. A lot of pipelines treat debate or self-consistency as a correctness signal. This work shows that agreement can be an illusion, papering over reasoning that never actually lined up. If you trust convergence as a proxy for being right, you may be measuring the wrong thing. Paper: https://t.co/fd2edva1qu Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c
// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and remain frozen or mostly unchanged. The harness, like the skills, needs to evolve with new models. What if the scaffold rewrites itself? This new work treats the harness, the prompts, tools, and control flow around the model as a learnable artifact that improves from its own runs rather than staying a fixed wrapper you hand-maintain. The scaffolding becomes the part that compounds, run after run. If you run long-horizon agents, a self-modifying harness turns scaffold upkeep from manual work into something the system earns on its own. Paper: https://t.co/byh1MP99xU Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
This is awesome! I am spending a lot of time on diffusion LLMs these days, so this is perfect timing. I feel like there are so many underexplored research questions around text diffusion. Weight available in HF. https://t.co/BpZM7Vxwvm
DiffusionGemma is our new experimental open model with up to 4x faster output on dedicated GPUs. Instead of predicting word-by-word, it generates entire blocks of text simultaneously. This lets the model self-correct and format complex markdown in real time.
It looks like great weather for tomorrow at Red Rocks for It Burns Joe Fitness free bootcamp 9am come try it! Life changing fun. https://t.co/4e0QmhaZY0
Awesome weather for Sunday It Burns Joe Fitness FREE bootcamp at Red Rocks 9am Come check it out! Also every Wednesday night 7:30pm indoor bootcamp at Motion Fitness https://t.co/UcQw7wDr31 https://t.co/NlSkuq1RD5
@danwilliamsphil @ATabarrok Exactly right. https://t.co/cDbkjaYtwx
This by @snewmanpv is spot on, and exactly what we called the "false summit" phenomenon in AI as Normal Technology β as we climb the mountain of AGI, what we thought was the peak is repeatedly revealed to be a false summit. This is what leads to the accusation that skeptics keep
Very excited to share that our paper "Towards a Science of AI Agent Reliability" was accepted at ICML 2026! See you in Seoul! π We just released our camera ready version with three important updates (details below). We also recorded a short video on the paper's contributions. Main changes (full discussion at https://t.co/1a5r1jNFF4): 1οΈβ£We have added the latest set of frontier models to our evaluation (GPT 5.5, Gemini 3.1 Pro and 3.5 Flash, and Claude Opus 4.7) and find that they are not meaningfully more reliable than previously released models. Agent reliability is still far from being solved. 2οΈβ£We have updated the definition and measurement of our outcome consistency metric, which contained a typo in the pre-print we initially released. This caused us to under-estimate outcome consistency in our initial set of results. We have updated the paper and our codebase to the corrected metric. Despite this change, our new results show that outcome consistency is still surprisingly low across many reported models. 3οΈβ£We discovered multiple issues in our HAL Generalist Agent scaffold that we used for our experiments on GAIA. Notably, we discovered multiple instances of answer leakage and agents cheating on our evaluation. This caused us to slightly over-estimate both accuracy and reliability. At the same time, we noticed that the scaffold was overly constrained in terms of permissible software library imports. This caused us to slightly under-estimate both accuracy and reliability. We have done a rigorous audit of the scaffold and have fixed those issues. Overall, we saw that our resulting accuracy and reliability numbers are not meaningfully impacted by this change when compared to our original numbers. πOur paper: https://t.co/HAKHzASrOZ πOur dashboard: https://t.co/apbtxtsdvz π₯Short video: https://t.co/uqIourw6C6 Joint work w/ @sayashk, @PKirgis, @khl53182440, @SaitejaUtpala, and @random_walker.
Massive output uptick due to agentic AI. Complete flat adoption. https://t.co/s6ubPsy0SL
For all the attention AI gets, data shows that the parties are still not adopting it as a major issue. I've been playing with @derekwillis 's awesome data on party fundraising emails. As my plot here shows, AI is just starting to tick up as a Democratic talking point, but it's still very, very modest. The Republicans' 2020-era freakout about social media remains much, much larger than the current AI freakout in terms of email focus...and that never amounted to much in terms of policy. At the same time, policy proposals around AI are getting bolder, like Sanders and Trump floating national ownership---yet surprisingly, the party rank and file are actually moving pretty slowly to adopt AI as an issue, maybe because it's not very salient with the American public. Will be very curious how this changes in the coming months. Seems likely Democrats may start messaging around it more...but maybe not, if we don't start to see meaningful employment effects.
Present Trump on Air Force One taking to the press: Reporter: 'Sir, on AI companies, potentially taking these equity stakes, have you Spoken to Sam Altman or any of the-' President Trump: 'No, there's a concept out there, there's so much money, and it's so big, that there are
The popular conversation around AI in America looks nothing like the narratives the elites are driving. For our new research, we analyzed 25,000 TikTok and YouTube videos about AI---and watched thousands of them ourselves---to understand how Americans are encountering AI in their everyday lives. Despite an elite conversation focused largely on backlash, AI videos embracing AI outnumber videos about resisting AI 3 to 1. These "adopter" videos don't focus on the things elites talk about: they talk about funny memes and effects AI can help make and ways you can use AI to help you with your job search. There is a significant and organized social media community focused on resisting AI, but surprisingly, it's not mainly about job loss, data centers, or existential risk. Instead, it's about creative theft and the erosion of human-made art. This has all the hallmarks of a genuine movement---with organized efforts to support human artists, to report AI-generated content, and to oppose the technology in the real world. All in all, when we look past the efforts of the labs and the media to impose a top-down narrative around job loss and existential risk, we find everyday Americans having a far different and in many ways more "normal" conversation (@random_walker)---one in which AI offers immediate and personal opportunities and challenges all at the same time. Check out the full research piece, which is loaded with interesting real example videos, here: https://t.co/AbFTqM4g7e
Nice to see "churnalism" get automated. An agent cranked this out based on @sayashk's tweet. I hope journalists will leave behind this low-value stuff to AI and focus on the hard parts β digging up non-public info; verification and provenance; supplying unique analysis. https://t.co/KEVYIrLZSX
Financial analysts spend ~70% of their time pulling numbers out of PDFs. We built a demo agent that ingests SEC filings and answers questions with exact citations highlighted on the original PDF page. About 600 lines of Next.js. No vector DB. Just LiteParse. https://t.co/dmV641aZi1
New in LlamaParse: Latency Metrics is now live. For every Parse, Extract, and Classify job, you can now get a full latency breakdown. All broken down by tier. β± Queue time β‘Processing time π Total latency There's also a new Metrics tab with a latency scatter plot and job volume histogram if you want to dig into patterns across your usage over time. Head to your Parse History tab to check it out. π¦ https://t.co/zqUqveJRgD

How do you know your document parser is ready for production? π€Existing benchmarks miss what AI agents actually need. That's the gap ParseBench, the first doc OCR benchmark for AI agents, fills. We'll unveil all the magic behind it in a live webinar π https://t.co/odSaGMAlkz https://t.co/hjZGEol4df
LlamaParse now parses HEIC files natively π . HEIC is Apple's default image format, so it shows up all over enterprise file systems. Photos of whiteboards, scanned docs, receipts snapped on an iPhone. You no longer need to convert to JPEG first. Point LlamaParse at the .heic file and it parses. Go ahead, parse that messy whiteboard.
Automate a loan underwriting pipeline in just a few lines of codeβ¨οΈ A typical loan file is a stack of pay stubs and brokerage statements, every one formatted differently, every number re-typed by hand. Here's a pipeline that does it automatically with LlamaParse: PDFs to clean markdown, fields into Pydantic models, then cross-document analysis that produces an underwriting summary with discrepancy flags. Full post and repo: https://t.co/QnuLitgmEn
LiteParse v2.0 is out now, and it is blazing fast + runs everywhere! We rewrote everything from scratch in Rust, and now: - up to 100x faster parsing - install natively in Rust, JS/TS, and Python - a custom WASM package enables browser and edge runtime usage pip install liteparse npm i οΌ llamaindex/liteparse npm i οΌ llamaindex/liteparse-wasm cargo install liteparse Blog: https://t.co/zWnhGNrgeb Repo: https://t.co/UJy6KQ1Dyi

Is grep π³π¦π’πππΊ all your AI agent needs for search? For a small codebase or a docs folder, the answer might be yes, but in most enterprise environments, agents face millions of PDFs, spreadsheets, and scanned documents. Lexical search alone can't read those formats, doesn't scale, and misses synonyms entirely. In our latest post, we break down: β Where grep shines (and why it's not going away) β Why RAG and semantic search are necessary at enterprise scale β How to layer lexical + semantic search for the best of both worlds The answer isn't grep vs. RAG, it is knowing when to reach for each and how to combine them. ποΈ Read the full breakdown: https://t.co/S758X1l3E5
Opus 4.8 dropped today. ParseBench results are out. β Slight gains: tables, semantic formatting, layout β οΈ Slight regressions: charts, content faithfulness π° Slight price/page increase Lots of alpha left in teaching LLMs to read docs like humans do. LlamaParse remains the best doc-ingestion API for AI agents.
When we say βLiteParse runs everywhere,β we mean it. Our WASM package is lightweight, minimal, and built for browser and edge runtimes, which makes it a perfect fit for @cloudflare Workers. Using WebAssembly, you can spin up a parser that runs directly on the Worker, takes PDF bytes as input, and returns extracted text plus page count (all in under 25 lines of code!)π π©βπ» Try it out now: https://t.co/zDYL0TCYQS ποΈ Get started with LiteParse: https://t.co/9zv8WOkbpS
Last week we revamped Liteparse to be the fastest PDF parser out there β‘οΈ An underrated part of liteparse is it doesn't just give you text. It gives you bounding boxes that a coding agent can use to paint exact audit trails back to the source document. For instance, check out the deep research skill we compiled in liteparse_samples: https://t.co/eoICBmvEnq Come check out liteparse: https://t.co/JNER0mVcB8 We are hard at work making liteparse even better (e.g. Markdown support). Please feel free to open up issues, PRs, and let us know your feature requests π
We've created the world's fastest PDF parser β‘οΈ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf, pypdf, markitdown, pdftotext, opendataloader, pymupdf4llm) Introducing LiteParse v2 - we rewrote the entire library into Rust and adapte
We're presenting ParseBench at CVPR 2026 today. π¦ Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise table is harder than it looks). The first doc-parsing benchmark built for AI agents: 2,000+ human-verified pages 167K+ test rules 5 dimensions: tables, charts, faithfulness, formatting, grounding Fully open source. π Talk TODAY, June 4, 9β10 AM at CVPR. Come say hi π π€ https://t.co/skla84GVTc π» https://t.co/h7SpuTWYVn π https://t.co/VnKcb48oJl
We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. β It contains 2k pages of real-world enterprise documents β It has comprehensive evaluation metrics around tables, charts, visual grounding, semantic formatting, and content faithfulness The core goal is measuring whether models can semantically interpret a document in the right way, without having models overfit to our precise benchmark. Parsing 100% of PDFs to 100% accuracy is the final boss for document OCR. In general, the latest frontier models have been tuned for coding, math, and scientific reasoning as opposed to precise visual understanding; hope more benchmarks that these will encourage overall progress towards solving this problem! Poster is below. If you want to learn more come check out our site or 30-page ArXiv paper: ParseBench: https://t.co/PWczfhp0OX ArXiv: https://t.co/2dEJIaBBkr
We're presenting ParseBench at CVPR 2026 today. π¦ Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise table is harder than it looks). The first doc-parsing benchmark built for AI ag
Most AI pipelines are only as good as the data we provide them with, and that usually means PDFs or other unstructured documents. Contracts, invoices, reports... All have special layout, language, and context mixed together, and getting reliable structured data out of them is still one of the hardest unsolved problems in enterprise AI. Parse-Flow is an open-source project we built to tackle this head-on. It puts four document processing primitives at the center of a visual workflow designer: π Parse β clean markdown and text from raw documents ποΈ Classify β assign documents to user-defined categories βοΈ Split β segment documents into typed chunks πͺ Extract β pull structured JSON against a schema You drag steps onto a canvas, drop in a document, and watch events stream back as the pipeline runs. Under the hood it's powered by a LlamaAgents workflow that walks your flow one step at a time, making every transition observable and every failure a first-class value. ποΈ Full write-up on the architecture here: https://t.co/my2ZT6wGnX π©βπ» Source code: https://t.co/vtuo7vXt9i
The Agent Open: AI's Pickleball Tournament π Come put your code and backhand to the test and embrace the full Open experience. Custom built out courts. Stadium seating. Exhibition matches by AI leaders. Fresh agent merch. Every infra startup you love, all in one place. Brought to you by in collaboration with: @braintrust , @browserbase , @cursor_ai , @modal , @p0 , @turbopuffer Where were you during the first Agent Open? Come make history. SF Edition. π https://t.co/B2rkZC9gMJ
Parsing a document accurately is one thing. Proving where every value came from is another. When a compliance team reviews an AI extraction, or an auditor needs to sign off on a figure pulled from a financial filing, "it came from this document" isn't enough. They need to see exactly where. The specific cell in the table, the exact line on the page, the precise word the agent used. Most parsers can get you to a paragraph or a table block. That's where the trail ends. Today we're shipping Granular Bounding Boxes in LlamaParse β word, line, and cell level coordinates for every value in your document. The result is a complete, verifiable trail from every extracted value back to its exact source in the document. Built for audit workflows, compliance review, and any pipeline where verification isn't optional. Read the full announcement β https://t.co/Me1QFEka8n
Day 0 Anthropic Fable 5 in ParseBench: We tested the model's advancements when it comes to document understanding. The model clearly peaks when it comes to adherence to the original text: π Content faithfulness: 90.02% vs 86.19% (Gemini 3 Flash) and 86.81% (GPT-5.5) π’ Semantic formatting: 72.62% vs 58.35% and 60.12%, a 12+ point lead These are two of the most important metrics for SOTA document understanding: does the output preserve what the document actually says, and does it preserve formatting that carries meaning? But ... it's not a sweep there continues to be a lot of alpha in unlocking document understanding for frontier models. Full results below π
Claude Fable 5 thinks document parsing is beneath it It is absolutely crushing on all reasoning-intensive/long horizon benchmarks: SWE-Bench Pro, FrontierCode, GDPval, Runescape, etc. But for document understanding tasks, it is roughly equivalent with Gemini 3 Flash in performance, at roughly 10-15x the token cost. We benchmarked the model on ParseBench and compared it against all other frontier models. It is definitely up there compared to other frontier models, but falls far short of specialized OCR providers. What we found interesting is that Fable 5 is self-aware about this. When we ask the model what tasks it enjoys the last, it actively said that it dislikes tasks "where the request is fully specified and the answer is fully known" - implying part of it being bad is due to laziness and lack of willingness to actually solve the task at hand. For a full list of results across different frontier models, check out ParseBench! https://t.co/PWczfhp0OX
Day 0 Anthropic Fable 5 in ParseBench: We tested the model's advancements when it comes to document understanding. The model clearly peaks when it comes to adherence to the original text: π Content faithfulness: 90.02% vs 86.19% (Gemini 3 Flash) and 86.81% (GPT-5.5) π’ Semantic f
LiteParse, our open-source/Rust-based doc parser, runs so quickly that Claude Fable 5 doesn't think it's real π₯ It is the fastest document parsing solution on the planet and a great choice for your AI document workloads. Check it out: https://t.co/VfA6yJwzcZ https://t.co/DUjhiF2CTq
LiteParse runs so fast that Claude Fable 5 doesn't think its real https://t.co/6CExW6owiw

As frontier models (e.g. Fable 5) continue to push the task horizon of knowledge work automation, it becomes ever more important for humans to be able to audit decisions back to the source context. It is extremely easy for agents to cite an entire document or document page, but much harder for them to trace back to the exact numbers/words/figures within a page. Today we've launched granular bounding boxes within LlamaParse, which allows you to obtain visual citations of every single word in the document. This allows human users to audit exact words and figures - not just general document regions or entire pages! Come check it out: https://t.co/TqP6OT5U5O
Parsing a document accurately is one thing. Proving where every value came from is another. When a compliance team reviews an AI extraction, or an auditor needs to sign off on a figure pulled from a financial filing, "it came from this document" isn't enough. They need to see
Congrats to the Microsoft AI team on MAI-Thinking-1! There's an art and science to post-training - navigating tradeoffs, grounding the model in what's meaningful to users, crafting a set of tastes and views. It shows in the model. Proud of the work we did together and to see them pushing the frontier. https://t.co/K9NxLO4jsP
Introducing ComplexConstraints β a new IF benchmark designed to test whether models can handle the kinds of constraints that show up in real work: 1. Conditional constraints (fire only when specific conditions are met) 2. Planning constraints (many requirements satisfied simultaneously) 3. Multistep constraints (each step feeds the next) 4. Implicit constraints (a competent colleague would just know). Models currently score between 0% and 40%. Here's an example: imagine you're a film producer drafting next week's shooting schedule. Two actresses are only available Tuesday and Thursday. The exterior scenes need daylight, but the forecast calls for rain on Wednesday. Tom's stunt double doesn't arrive until Friday, and Chris and Sydney don't get along. There are twenty-six interdependent constraints, and missing any one is a failure. What's interesting is also what the data teaches. We trained a Qwen-4B model on 1K ComplexConstraints companion examples. It reached parity with a model 60x its size, and the gains transferred to other IF benchmarks like MultiChallenge and AdvancedIF. Single-turn data even generalized to multi-turn behaviors, because tracking simultaneous requirements without dropping the lower-priority ones is the same skill multi-turn IF tests. Read more! Blog post: https://t.co/Su8UQH5EZd Leaderboard: https://t.co/WdMCiKov7Z