Your curated collection of saved posts and media

Showing 32 posts ยท last 14 days ยท by score
A
arnicas
@arnicas
๐Ÿ“…
May 04, 2026
82d ago
๐Ÿ†”29902794

@AhmedShahnab Some country files are being screwed by overseas territories https://t.co/8RRmMpROkO

Media 1
๐Ÿ–ผ๏ธ Media
N
nealagarwal
@nealagarwal
๐Ÿ“…
May 05, 2026
81d ago
๐Ÿ†”15486023

Iโ€™m looking for the first full-time developer for https://t.co/nprGVKUWBn! Youโ€™d work with me on making new web projects and games for the site. Preferably in nyc If anyone is interested or has leads dm me or email hi@neal.fun! https://t.co/0bOvzn7UVc

Media 1
๐Ÿ–ผ๏ธ Media
A
avshah99
@avshah99
๐Ÿ“…
Apr 22, 2026
94d ago
๐Ÿ†”42376698

๐ŸšจNew preprint! We find evidence of LLMs enabling people to file lawsuits without lawyers (filing "pro se") at historically unprecedented rates in federal courts.๐Ÿ‘‡ 1/n https://t.co/JCj8oq5Jym

Media 1
๐Ÿ–ผ๏ธ Media
W
willknight
@willknight
๐Ÿ“…
Apr 27, 2026
89d ago
๐Ÿ†”90929441

This meditation app is was invented, designed and coded, and then submitted to the app store by an AI model (it made a few mistakes along the way). In a new research paper, @sayashk and others say having AI take on this kind of messy open world task could offer a better way to measure progress. Very interesting! (paper https://t.co/ki0ymYzGkT) (app https://t.co/jbWeNcGMma)

Media 1Media 2
๐Ÿ–ผ๏ธ Media
E
evijit
@evijit
๐Ÿ“…
Apr 29, 2026
87d ago
๐Ÿ†”14577475

AI evaluation is becoming its own compute bottleneck. We often talk about the cost of training frontier models, but the cost of evaluating them is starting to matter just as much, especially for agents, scientific ML systems, and training-in-the-loop benchmarks. In our new Evaluating Evaluations post, we look at how evals are crossing a threshold where cost changes who can participate. The Holistic Agent Leaderboard spent about $40K on 21,730 agent rollouts across 9 models and 9 benchmarks. A single GAIA run on a frontier model can cost $2,829 before caching. And once you care about reliability, repeated runs can multiply these costs many times over. This creates a real accountability problem. If only large labs can afford statistically credible evals, independent researchers, auditors, journalists, and public-interest organizations are left with partial visibility into frontier systems. The core issue is that benchmark design is changing. Static benchmarks could often be compressed aggressively while preserving rankings. Agent benchmarks are noisier and scaffold-sensitive. Training-in-the-loop benchmarks are expensive by construction. As evals move closer to real work, they also become harder to make cheap. Some takeaways: โ†’ Leaderboards should report cost alongside accuracy. โ†’ Reliability should not be treated as optional. โ†’ We need reusable eval artifacts! Shared documentation formats, such as Every Eval Ever, can help the field stop paying repeatedly for the same measurements. Read the full post: https://t.co/sArlZMkytF Thanks for the insights @LChoshen , Yifan Mai, and @cgeorgiaw๐Ÿค—

Media 1Media 2
+3 more
๐Ÿ–ผ๏ธ Media
M
minzlicht
@minzlicht
๐Ÿ“…
Apr 29, 2026
87d ago
๐Ÿ†”53389276

Imagine a 19-year-old scrolling TikTok. She watches a creator list five "signs you have undiagnosed anxiety." She recognizes three in herself. By the end of the week, she's describing herself as anxious to her friends. A month later, she's avoiding situations she used to handle fine. What went wrong? In a new paper by my PhD student Dasha Sandra, titled "Why mental health awareness can harm: Converging explanations for a societal problem", we argue that well-meaning mental health awareness can backfire, and we identify how. Four separate literatures (concept creep, nocebo effects, prevalence inflation, and illness self-labeling) have been circling the same problem from different angles. We show they converge on three mechanisms: 1.Awareness lowers the threshold for what counts as a disorder. 2. It trains people to scan their inner lives for symptoms and reinterpret normal distress as pathology. 3. Once someone adopts an illness identity, they behave in ways that confirm and deepen it. The evidence is wide. Learning that loneliness is harmful makes solitude feel worse. Learning that stress is harmful worsens well-being and performance. Awareness videos about fake conditions like "wind turbine syndrome" produce real headaches. Trigger warnings raise anticipatory anxiety without reducing distress. This does not mean awareness should stop. It means awareness can have unintended consequences, including manufacturing the suffering it tries to prevent. Inoculating people against these mechanisms works, and we already have evidence it does. Link to paper: https://t.co/ucoGyhEuAj

Media 1
๐Ÿ–ผ๏ธ Media
A
ahall_research
@ahall_research
๐Ÿ“…
Apr 29, 2026
87d ago
๐Ÿ†”60885641

1800: If Thomas Jefferson is elected "Murder, robbery, rape, adultery, and incest will all be openly taught and practiced." When it comes to politics we have a bad habit of romanticizing the past and imagining that today's politics are worse and coarser. To make this visceral, I built a little app that shows what the 1800 election would have felt like if X had been around. Scrolling through it really does give you a sense that vicious, indecorous politics long pre-dates present day. Check it out here: https://t.co/eo2vFOf6TF

Media 2
๐Ÿ–ผ๏ธ Media
T
TawilTeddy
@TawilTeddy
๐Ÿ“…
May 01, 2026
85d ago
๐Ÿ†”11847556

Experts have three views on the future of work, each credible but sharply opposed. Whoโ€™s right? In a new paper for @CarnegieEndow & @CEIPTechProgram, I lay out the best arguments made by the alarmed, patient, and excited groups. ๐ŸงตOn the most important points and what policymakers can do today

Media 1
๐Ÿ–ผ๏ธ Media
J
justanotherlaw
@justanotherlaw
๐Ÿ“…
May 02, 2026
84d ago
๐Ÿ†”82155726

A recent viral paper claims to reverse-engineer the parameter counts of frontier models: GPT-5.5 = 9.7T, Opus 4.7 = 4.0T, o1 = 3.5T, etc. @ben_sturgeon and I investigated and found serious issues in the paper; fixing them gives GPT-5.5 as ~1.5T (90% CI: 256B-8.3T). https://t.co/MbWQyVlmsE

Media 1
๐Ÿ–ผ๏ธ Media
M
MatthewJBar
@MatthewJBar
๐Ÿ“…
May 05, 2026
81d ago
๐Ÿ†”68879935

Some people are being way too alarmist about Mythos. 80,000 Hours called Mythos "an AI that can break into almost any computer on Earth". Zvi Mowshowitz said, "If given to anyone with a credit card, Claude Mythos would give attackers a cornucopia of zero-day exploits for essentially all the software on Earth". But these descriptions are unfounded. @natalia__coelho shows in her latest blog post that the cyber capabilities of Mythos are nearly tied with GPT-5.5, across practically every public benchmark we have available. This includes both narrow and broad cyber evaluations. It is not way ahead of trend. The same holds for general capability evaluations. Except for a somewhat impressive score on SWE-bench Pro, Claude Mythos is nearly on-trend, at most a few months ahead. In other words, it's barely better than models that millions of people already have access to. It's an impressive model, but I'm very skeptical that Mythos is going to take down our digital infrastructure or cause a cyber catastrophe.

Media 1Media 2
+2 more
๐Ÿ–ผ๏ธ Media
R
random_walker
@random_walker
๐Ÿ“…
May 06, 2026
80d ago
๐Ÿ†”26218933

Excited to give this talk at the Stanford Digital Economy Lab on May 18! I will do three things: discuss my group's recent research, identify the most pressing gaps in the community's current understanding, and provide a long-term perspective. Hope to see you there in person or virtually. https://t.co/Qa2eNkVsnZ @DigEconLab

Media 1
๐Ÿ–ผ๏ธ Media
D
dongyangzi
@dongyangzi
๐Ÿ“…
May 05, 2026
80d ago
๐Ÿ†”26700234

We furthered AI research by reproducing CRUX #1 for Windows using @getnenai's infrastructure without needing to buy a Windows machine-- checkout our blog post https://t.co/Mjz1D5myF9

@sayashk โ€ข Thu Apr 16 17:49

Benchmarks are saturated more quickly than ever. How should frontier AI evaluations evolve? In a new paper, we argue that the AI community is already converging on an answer: Open-world evaluations. They are long, messy, real-world tasks that would be impractical for benchmarks.

Media 1
๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
Apr 24, 2026
91d ago
๐Ÿ†”37656389

We benchmarked GPT-5.5 on document understanding ๐Ÿ“„๐Ÿ“Š We ran it through ParseBench, our comprehensive OCR benchmark over enterprise documents. We evaluated metrics across various dimensions: visual grounding, tables, charts, and more. We evaluated GPT-5.5 on mid thinking and zero-thinking modes. When compared against GPT-5.4 (0 thinking) and Opus 4.7 (adaptive thinking): ๐Ÿ“ˆ GPT-5.5 wins on tables ๐Ÿ“ˆ GPT-5.5 wins on visual grounding ๐Ÿ“‰ GPT-5.5 0-thinking does worse on charts than GPT-5.4 0-thinking ๐Ÿ“‰ Higher thinking does worse than lower thinking of content faithfulness, semantic formatting ๐Ÿ“‰ Opus 4.7 wins overall on content faithfulness and semantic formatting ๐Ÿ’ธ GPT-5.5 is expensive: 13c per page at mid-thinking modes and 5.93 at 0-thinking! This is 5x the cost of any competitive OCR solution. Conclusion: GPT-5.5 is one of the better frontier models out there in terms of pure accuracy, but def not pound for pound w.r.t price.

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 27, 2026
89d ago
๐Ÿ†”21791003

Loan processors spend 40โ€“60% of their time reconciling income across tax returns, pay stubs, W-2s, and bank statements. We built an end-to-end pipeline that automates it with LlamaParse + the Claude Agent SDK: ๐Ÿ“„ Schema-driven extraction across 4 doc types with confidence scores + citations ๐Ÿ” Cross-document validation with Claude โ€” catches W-2/pay-stub gaps, unexplained Zelle/Venmo deposits, employer name mismatches ๐Ÿ“Š Self-contained HTML report with a COMPLETE / REVIEW / FLAG decision Full code + walkthrough: https://t.co/ozm4VwWmj3

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 28, 2026
88d ago
๐Ÿ†”16946011

Let's talk document formatting. Bold. Italics. Superscripts. Strikethroughs. The visual cues humans rely on every time we read a doc, and ones existing OCR benchmarks completely ignore. ๐Ÿ˜ฑ"$199" struck through next to "$149" isn't decoration. It's the meaning. ๐Ÿ˜ฑA superscript tells your agent "3" is a citation, not part of the number. Flatten that and your agent is reading a different doc than you are. Two weeks ago we released ParseBench, the first document OCR benchmark for AI agents. One of five metrics: the Semantic Formatting Score. Read more๐Ÿ‘‡ https://t.co/2sq5ncGiel

๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 29, 2026
87d ago
๐Ÿ†”90606809

Parsing documents with AI agents just got a lot more seamless๐Ÿš€ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ŸŒ Once connected, you'll be able to: ๐Ÿ“ Parse documents into clean markdown ๐Ÿ” Classify files against your own categories โœ‚๏ธ Split long documents into labelled sections โฌ†๏ธ Upload files via URL or a browser-based upload flow Building a production MCP server surfaced some non-obvious challenges: getting auth to align with an existing platform identity system using @WorkOS, working around MCP's lack of built-in file upload support, and making deployments, rate limiting and observability feel native with @vercel and @AxiomFM. We wrote up all of it, from the OAuth flow, to the token-based upload design, to the tradeoffs we hit along the way๐Ÿ“ ๐Ÿ“š Read the full blog: https://t.co/2E3qIVeUYI ๐Ÿ‘ฉโ€๐Ÿ’ป GitHub repository: https://t.co/ru1Evj7Zsr

๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
Apr 30, 2026
86d ago
๐Ÿ†”38085689

We shipped a LlamaParse MCP server to let you parse, classify, split, and generally operate over your hardest documents with your favorite AI agent ๐Ÿ“„๐Ÿค– Check out the MCP server: https://t.co/eOeDgzOCI5 This is both a useful feature and a great piece of engineering that we want to share with the community. ๐Ÿ’ก MCP does not have built-in file upload support, so we needed to implement a URL-based upload endpoint and couple it with parse operations ๐Ÿ’กWe built in an integration with @WorkOS OAuth. ๐Ÿ’กWe built in observability and rate-limiting Huge shoutout to @itsclelia for shipping this! Come check out our blog writeup: https://t.co/XsCZlKnzez LlamaParse: https://t.co/TqP6OT5U5O

@llama_index โ€ข Wed Apr 29 16:00

Parsing documents with AI agents just got a lot more seamless๐Ÿš€ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ŸŒ Once connected, you'll be able to: ๐Ÿ“ Pars

Media 2
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 30, 2026
86d ago
๐Ÿ†”79436469

Building scalable, distributed document processing pipelines isnโ€™t easy. Thatโ€™s why we teamed up with @render to build a system that: ๐Ÿ“ Leverages the LlamaParse platform to parse, classify, extract, and retrieve information from documents โš™๏ธ Uses Render Workflows to distribute tasks across nodes and accelerate background processing โšก Deploys a lightweight server and database on Render, giving you an instant interface to interact with your pipeline ๐Ÿ‘ฉโ€๐Ÿ’ป Explore the repo to see it in action: https://t.co/eiJqklNVhj ๐Ÿ“š And check out the step-by-step breakdown by @ojusave and @itsclelia: https://t.co/gePmuL7YtV

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 30, 2026
86d ago
๐Ÿ†”01071648

Thank you AI Dev Day '26 @DeepLearningAI @jerryjliu0 shares why SOTA LLMs can build an app but can't read a PDF ๐Ÿคฏ https://t.co/L2WMAUVaJ4

Media 1Media 2
+6 more
๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
May 03, 2026
83d ago
๐Ÿ†”42086427

Parsing PDFs is hard This past week I gave a few talks (at both AI Dev '26 by @DeepLearningAI and @Capgemini ) on why this is still such an open problem, and itโ€™s even more important as agents become the consumers of documents, and need the OCR tools to read them properly. The fundamental issue is that PDFs are designed for print and display purposes, not to give back a linearized, semantically meaningful string of text. Text and tables are represented as a bunch of chars and lines, without any guaranteed order. This is what the community is solving with VLM-based approaches, including our own efforts around LlamaParse and ParseBench. If youโ€™re interested in learning more about the problem, check out the blog post I wrote on this a while ago! https://t.co/740ZiAFyOk

Media 1Media 2
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
May 05, 2026
81d ago
๐Ÿ†”99831547

๐ŸŽ‰ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

Media 1Media 2
๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
May 05, 2026
81d ago
๐Ÿ†”37779563

Iโ€™m excited to announce that @llama_index is on the @CBInsights AI 100 list for 2026 ๐Ÿ”ฅ Weโ€™re on a mission to parse all of the worldโ€™s PDFs, and make them accessible to both humans and AI agents. List: https://t.co/8pKDaPnHI8 If you havenโ€™t done so already, our website design is awesome, check it out: https://t.co/YiIfjVlzb6

@llama_index โ€ข Tue May 05 15:06

๐ŸŽ‰ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

Media 1Media 2
+1 more
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
May 06, 2026
80d ago
๐Ÿ†”54837865

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powered by the LlamaParse TypeScript SDK ๐Ÿ“ฑ Three steps, thatโ€™s it: ๐Ÿ”‘ Add your API key (securely stored on-device) ๐Ÿ“ธ Snap a photo of anything with text ๐Ÿ“„ Parse it and, in under a minute, get clean, copyable text No hassle, no manual typing. ๐Ÿš€ Try it now: https://t.co/mEeraV4zap ๐Ÿฆ™ Get started with LlamaParse: https://t.co/Pmc8usyjTwโ€“

๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
May 06, 2026
80d ago
๐Ÿ†”23058683

I โค๏ธ NYC We're hosting two in-person events next Wednesday: 1๏ธโƒฃ FinParse workshop: Build AI agents to extract and act over the most complex financial documents 2๏ธโƒฃ AI Happy Hour with Tabs: drinks, conversations, and L'Industrie Pizza ๐Ÿ• Workshop: https://t.co/RQXRxxAwFk Happy hour: https://t.co/8mlnOYpAgg Come check it out!

@llama_index โ€ข Wed May 06 13:00

LlamaIndex NYC takeover, 5/13 ๐Ÿ—ฝ Our CEO Jerry Liu is in town. Two events, open to every NYC builder: ๐Ÿ› ๏ธ FinParse Workshop โ€” laptops out, hands-on with @jerryjliu0 โ†’ https://t.co/MeaWrBU2ca ๐Ÿ• AI Engineers on Tap โ€” happy hour w/ @tabs โ†’ https://t.co/8FQTwFLFk1

Media 1
๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
May 06, 2026
80d ago
๐Ÿ†”41731268

Last week I gave a talk at AI Dev โ€™26 by @DeepLearningAI on โ€œAI canโ€™t read PDFs, how do we fix itโ€ . Iโ€™m sharing the slides publicly if others are interested in doing a deep dive into document understanding. AI agents are going to automate huge amounts of knowledge work, but knowledge work depends on data, a lot of that data is in documents/PDFs, and existing OCR tools suck. PDFs are a format that is inherently hard to read, and I dive into specific reasons why (tables, layouts), and why frontier VLMs and benchmarks are still insufficient. Even as agents get better and more general, they need the right tools to read and act over PDFs. They both need this at the data ingest layer, as well as tools they can call on the fly. Check out the slides: https://t.co/IyaTlQMPlW Weโ€™re building high-quality AI document processing, both with LlamaParse, along with OSS efforts like LiteParse and ParseBench. If you have a ton of PDFs that youโ€™re hoping to unlock with AI, come talk to us! https://t.co/TsPLwqZ8yg

Media 1Media 2
+3 more
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
May 07, 2026
79d ago
๐Ÿ†”72472890

A few weeks ago @simonw got Claude to port LiteParse to the browser. Today, we are launching that work as a complete guide in our docs! https://t.co/sqmrVxx3YU The guide itself relies on some fun hacks with vite and mocking. We expect this process to improve with future releases, so stay tuned!

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
May 11, 2026
75d ago
๐Ÿ†”62374980

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet ๐˜€๐—ฎ๐—ป๐—ฑ๐—ฏ๐—ผ๐˜…๐—ฒ๐—ฑ-๐—น๐—ถ๐˜, a Rust ๐Ÿฆ€ CLI agent that combines: - LiteParse, our lightning-fast local parser for PDFs, images, Office files, and more - A secure sandbox powered by @microsandbox - Full filesystem mounting, so your agent can safely interact with local files inside the sandbox Mount your local workspace, give the agent shell access, and let it do its magic ๐Ÿช„ ๐Ÿ‘ฉโ€๐Ÿ’ป GitHub: https://t.co/3NpXcFw58p ๐Ÿ“š Learn more about LiteParse: https://t.co/XoAIrbPtE4โ€“

Media 1
๐Ÿ–ผ๏ธ Media
H
HelloSurgeAI
@HelloSurgeAI
๐Ÿ“…
Apr 29, 2026
87d ago
๐Ÿ†”94059834

GDP.pdf was accepted to the CVPR 2026 Workshop on Multimodal Reasoning can frontier models handle the three-letter document type that runs the world? we partnered with hundreds of expert surgers - ER physicians, construction engineers, corporate litigators - to find out. every one scored under 15%. paper, leaderboard, and dataset below. paper: https://t.co/y4hmPXbpx4 dataset: https://t.co/rElrfOntcU leaderboard: https://t.co/9CMY6JVhDL blog: https://t.co/0Wj97DBr44

Media 1
๐Ÿ–ผ๏ธ Media
D
DeryaTR_
@DeryaTR_
๐Ÿ“…
May 01, 2026
85d ago
๐Ÿ†”22220542

As I mentioned before, I am now sharing an example from GPT-5.5 Pro, also featured by OpenAI, that really left me stunned by what it is capable of in biomedical science. (full report on the website I created with Codex, link in the thread). To push GPT-5.5 Pro hard, I uploaded a real data set of immune subset (T cells) gene-expression spreadsheet: 62 sorted T cell samples, 27,906 gene columns, and millions of underlying data points across different T cell subsets. Importantly, this public dataset also had paired structure making it possible to separate true cell-state biology from donor-to-donor variation. I asked GPT-5.5 Pro not merely to summarize the spreadsheet, but to analyze it deeply: What can we learn from this dataset? What are the mechanistic insights? What are the most important biological questions that emerge? What follow-up experiments should we do next? It thought for about 100 minutes and produced a roughly 40-page report! What amazed me was not just the length or even the initial analysis, since previous models are also capable of doing this. What amazed me was the quality of the reasoning and insights it provided! The report recognized that this was not just a table of genes, but two overlapping experimental designs. It identified the major biological axis, which in plain language was that the cells were not just โ€œdifferent categories.โ€ They formed a coherent differentiation landscape, moving from future potential toward immediate function. It also understood the caveats. It did not overclaim from bulk gene-expression data. It clearly explained that bulk transcriptomics cannot distinguish whether every cell in a sorted population has shifted or whether a smaller subpopulation is dominating the signal. It recommended the right next steps experiments, and integration with donor metadata. This is what made the report feel so special to me. It was not just doing statistics. It was reasoning like an expert systems immunologist. It saw the structure of the experiment, interpreted the patterns, built a mechanistic model, identified limitations, proposed causal hypotheses, and laid out a translational roadmap. Other advanced models have been able to generate excellent biomedical reports before, including previous GPT-5 models. So I don't want to claim this is an entirely new type of capability. But this one felt different in an important way. It had more scientific elegance, more restraint, more biological intuition, and more of the nuanced judgment that usually comes only from years of hands-on experience in the field. It felt like this AI model had crossed another threshold. This is the kind of analysis that could easily take a research team months to perform, refine, interpret, and write up. Even then, many teams might not produce something this integrated, this mechanistically coherent, and this useful as a launchpad for future experiments. I know a 40-page T-cell gene-expression analysis may not be exciting to everyone. To illustrate how good it is, also had Codex built a web site with it anyone can explore, link below. ๐Ÿ˜Š Those interested can go deeper into the report. I also wanted this example on the record because, because to me, it is evidence that we are entering a new stage in AI-assisted biomedical science. The important point is no longer that AI can "analyze data and write a report.โ€ The important point is that AI can now help transform complex biological data into mechanistic understanding, experimental priorities, and testable hypotheses at a speed and depth that would have been almost unimaginable a short time ago. For biomedical science, this is a very big deal! Of course, this may vary across domains, and every analysis still needs expert review, validation, and experimental follow-up. But in my own field, with data I understand deeply, this felt like another inflection point. I feel strongly that we have crossed another milestone threshold in the age of AI, with the release of GPT-5.5.

Media 1
๐Ÿ–ผ๏ธ Media
A
a16z
@a16z
๐Ÿ“…
Apr 30, 2026
86d ago
๐Ÿ†”89065914

.@pmarca says AI is the biggest technological revolution of his life: "This is the biggest technological revolution of my life. This is clearly bigger than the internet. The comps on this are things like the microprocessor, the steam engine, and electricity." "The neural network as an idea continued to be explored in academia for the last 80 years. And essentially it didn't work, it was decade after decade of excessive optimism, followed by disappointment." "Then basically we all saw what happened with the ChatGPT moment. All of a sudden it crystallized, and it was like, 'Oh my God, it turns out it works.' We're sort of three years into effectively an 80-year revolution."

๐Ÿ–ผ๏ธ Media
B
BoWang87
@BoWang87
๐Ÿ“…
May 04, 2026
82d ago
๐Ÿ†”36636119

What if you could test a drug on a human organ and simultaneously know what would have happened without it? That's what we built. Proud to share our work in @NatureBiotech: digital twins of human lungs ๐Ÿซ๐Ÿ”ฅ (paper: https://t.co/Y8wzxNwWfd) We created digital twins of ex vivo human lungs: multimodal AI models trained on 951 human lungs from the world's largest EVLP dataset at @UHN. Physics-informed ML across physiology, biochemistry, transcriptomics, proteomics, metabolomics, and imaging, all forecasting together. The key insight: the physical lung receives the treatment. The twin is the untreated control. Paired causal inference on the same organ. No separate cohort. No intersubject noise. Result: we detected drug efficacy with 6 lungs. Traditional methods would need 18. This is what precision preclinical evaluation looks like. From Virtual Cells โ†’ Virtual Organs โ†’ Virtual Patients. One step toward virtual organs replacing animal testing. Huge congratulations to Elly Zhou (a phd student I co-supervise) for leading this work with exceptional rigor, and to Andrew Sage and @SKeshavjee for building the foundation that made it possible.

Media 1Media 2
๐Ÿ–ผ๏ธ Media
S
s_batzoglou
@s_batzoglou
๐Ÿ“…
May 09, 2026
77d ago
๐Ÿ†”83623399

Our paper โ€œABD: Defaultโ€“Exception Abduction in Finite First-Order Worldsโ€ was accepted to KR 2026, in the KR Meets Machine Learning and Explanation track. arXiv: https://t.co/HYb5icwsel code/data: https://t.co/ZXXqzouuni A very nice plain-language summary: https://t.co/Zl5lHu4EGH Motivation: many reasoning benchmarks ask whether a model gets an answer right, but give only a weak view of *why* it succeeds or fails. In ABD, the model has to synthesize an explicit first-order formula explaining the exceptions to a default rule. A typical task is not: โ€œclassify this object.โ€ It is: given a background first-order theory, and several finite relational worlds, output one formula ฮฑ(x) that defines when an object is abnormal (i.e., does not satisfy the background theory), while keeping the exception set sparse. That gives us three things that are hard to get from many natural-language benchmarks: 1. exact semantic checking; 2. an explicit hypothesis produced by the model; 3. a cost measure, so we can distinguish โ€œvalid but bloatedโ€ explanations from compact ones. The benchmark has three observation regimes: โ€ข Full / closed-world: the observed facts are complete. โ€ข Partial / existential completion: unknowns may be completed favorably. โ€ข Skeptical / universal completion: the explanation must work under all completions. So the same high-level task can be tested under increasingly demanding assumptions about incomplete information. A central design choice is that each instance contains multiple prompt worlds. The model must output one shared ฮฑ(x) that works across all of them jointly. This is meant to test rule synthesis, not per-world labeling. We also evaluate on matched holdout worlds. This lets us separate two failure modes: formulas that stop being valid on fresh worlds, and formulas that remain valid but become too expensive i.e., call too many objects "abnormal". Empirically, evaluating ten frontier LLMs on 600 instances, the best models often achieve high validity, but parsimony remains a real issue. Many hypotheses repair the theory but use unnecessarily broad exception definitions. Holdout evaluation also shows distinct generalization failures across the three observation regimes. ABD aims to test one specific type of reasoning, formulating simple rules (first-order formulas) that explain exceptions to a given background theory. Finite first-order worlds give us a clean testbed for doing so, by being explicitly solvable and feasible for generating problems automatically. I hope it is useful both for evaluating language models and for comparing them with more symbolic or solver-aided synthesis methods.

Media 1Media 2
๐Ÿ–ผ๏ธ Media