Your curated collection of saved posts and media

Showing 32 posts ยท last 7 days ยท newest first
D
DennisonBertram
@DennisonBertram
๐Ÿ“…
Apr 26, 2026
91d ago
๐Ÿ†”75539816

Yeah okay, Lego bros, brodettes and brotheys are cooked with this one. GPT 2 Image can create full Lego sets! With actual Bricklink IDs so you can order the parts and build it. Whole new business opportunity here for the taking. https://t.co/4d4aJeEGTU

Media 1Media 2
+1 more
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
May 07, 2026
80d ago
๐Ÿ†”91879996

Side quest: trying to write custom firmware to a DVD drive and get low-level hardware control. This has been a long-term AI bench of mine. Not quite there yet, but 5.5 made a ton of progress - it feels like we're quite close to the fun + scary world where any sufficiently complex device can probably be reverse-engineered and hacked with a bit of effort! Rough write-up of the current state: https://t.co/yIbRLMDstA

Media 1
๐Ÿ–ผ๏ธ Media
T
takex5g
@takex5g
๐Ÿ“…
May 07, 2026
80d ago
๐Ÿ†”61751750

ๆš‡ใ™ใŽใฆDVDใฎใ‚ขใƒฌใ‚’ๅ†็พใ™ใ‚‹ใƒžใ‚ทใƒณไฝœใฃใฆใŸใ‚‰้€ฃไผ‘็ต‚ใ‚ใฃใŸ.. ่ง’ใซๅฝ“ใŸใ‚‹ใจ็ˆ†่ฃ‚ใƒฌใ‚คใƒณใƒœใƒผใซๅ…‰ใ‚‹ https://t.co/QlTobdhEFj

๐Ÿ–ผ๏ธ Media
๐Ÿ”johnowhitaker retweeted
T
ใ‚†ใ†ใ‚‚ใ‚„
@takex5g
๐Ÿ“…
May 07, 2026
80d ago
๐Ÿ†”61751750

ๆš‡ใ™ใŽใฆDVDใฎใ‚ขใƒฌใ‚’ๅ†็พใ™ใ‚‹ใƒžใ‚ทใƒณไฝœใฃใฆใŸใ‚‰้€ฃไผ‘็ต‚ใ‚ใฃใŸ.. ่ง’ใซๅฝ“ใŸใ‚‹ใจ็ˆ†่ฃ‚ใƒฌใ‚คใƒณใƒœใƒผใซๅ…‰ใ‚‹ https://t.co/QlTobdhEFj

โค๏ธ108,473
likes
๐Ÿ”11,037
retweets
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
May 08, 2026
79d ago
๐Ÿ†”98852999

Microbial traffic jam. (Ciliates in biofilm, seawater sample from the Oregon coast, polarized light microscopy) https://t.co/ylDP9RyTeF

๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
May 08, 2026
79d ago
๐Ÿ†”43962283

More tests with duckweed transformation and regeneration https://t.co/8gkCsuQozp

Media 1Media 2
+1 more
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
May 08, 2026
79d ago
๐Ÿ†”82388512

Mildly obsessed with this bacteria growing on an old plate (cc @ContamClub ) https://t.co/lW2XDVSHJ6

Media 1Media 2
+2 more
๐Ÿ–ผ๏ธ Media
S
sam_paech
@sam_paech
๐Ÿ“…
Apr 26, 2026
92d ago
๐Ÿ†”03947444

New EQ-Bench results! Opus 4.7: clean sweep, still the king. Deepseek 4: very strong, near frontier on EQ-Bench & longform writing. Kimi k2.6: Strong in shortform but seems to suffer degradation in longform writing. GPT-5.5: performs ~identically to GPT-5.4. https://t.co/slwFfDoBVj

Media 1Media 2
+1 more
๐Ÿ–ผ๏ธ Media
A
arnicas
@arnicas
๐Ÿ“…
May 03, 2026
84d ago
๐Ÿ†”06394707

Newsletter update of last 2 weeks of AI/game/web/narrative links of interest... more splats and 360ies, web games, narrative and creativity, vis of Aaron Reed's Subcutanean procgen novel, and more! 1/2 https://t.co/vaHQM9gO48

Media 1
๐Ÿ–ผ๏ธ Media
A
arnicas
@arnicas
๐Ÿ“…
May 04, 2026
83d ago
๐Ÿ†”29902794

@AhmedShahnab Some country files are being screwed by overseas territories https://t.co/8RRmMpROkO

Media 1
๐Ÿ–ผ๏ธ Media
N
nealagarwal
@nealagarwal
๐Ÿ“…
May 05, 2026
82d ago
๐Ÿ†”15486023

Iโ€™m looking for the first full-time developer for https://t.co/nprGVKUWBn! Youโ€™d work with me on making new web projects and games for the site. Preferably in nyc If anyone is interested or has leads dm me or email hi@neal.fun! https://t.co/0bOvzn7UVc

Media 1
๐Ÿ–ผ๏ธ Media
A
avshah99
@avshah99
๐Ÿ“…
Apr 22, 2026
95d ago
๐Ÿ†”42376698

๐ŸšจNew preprint! We find evidence of LLMs enabling people to file lawsuits without lawyers (filing "pro se") at historically unprecedented rates in federal courts.๐Ÿ‘‡ 1/n https://t.co/JCj8oq5Jym

Media 1
๐Ÿ–ผ๏ธ Media
W
willknight
@willknight
๐Ÿ“…
Apr 27, 2026
90d ago
๐Ÿ†”90929441

This meditation app is was invented, designed and coded, and then submitted to the app store by an AI model (it made a few mistakes along the way). In a new research paper, @sayashk and others say having AI take on this kind of messy open world task could offer a better way to measure progress. Very interesting! (paper https://t.co/ki0ymYzGkT) (app https://t.co/jbWeNcGMma)

Media 1Media 2
๐Ÿ–ผ๏ธ Media
E
evijit
@evijit
๐Ÿ“…
Apr 29, 2026
88d ago
๐Ÿ†”14577475

AI evaluation is becoming its own compute bottleneck. We often talk about the cost of training frontier models, but the cost of evaluating them is starting to matter just as much, especially for agents, scientific ML systems, and training-in-the-loop benchmarks. In our new Evaluating Evaluations post, we look at how evals are crossing a threshold where cost changes who can participate. The Holistic Agent Leaderboard spent about $40K on 21,730 agent rollouts across 9 models and 9 benchmarks. A single GAIA run on a frontier model can cost $2,829 before caching. And once you care about reliability, repeated runs can multiply these costs many times over. This creates a real accountability problem. If only large labs can afford statistically credible evals, independent researchers, auditors, journalists, and public-interest organizations are left with partial visibility into frontier systems. The core issue is that benchmark design is changing. Static benchmarks could often be compressed aggressively while preserving rankings. Agent benchmarks are noisier and scaffold-sensitive. Training-in-the-loop benchmarks are expensive by construction. As evals move closer to real work, they also become harder to make cheap. Some takeaways: โ†’ Leaderboards should report cost alongside accuracy. โ†’ Reliability should not be treated as optional. โ†’ We need reusable eval artifacts! Shared documentation formats, such as Every Eval Ever, can help the field stop paying repeatedly for the same measurements. Read the full post: https://t.co/sArlZMkytF Thanks for the insights @LChoshen , Yifan Mai, and @cgeorgiaw๐Ÿค—

Media 1Media 2
+3 more
๐Ÿ–ผ๏ธ Media
M
minzlicht
@minzlicht
๐Ÿ“…
Apr 29, 2026
88d ago
๐Ÿ†”53389276

Imagine a 19-year-old scrolling TikTok. She watches a creator list five "signs you have undiagnosed anxiety." She recognizes three in herself. By the end of the week, she's describing herself as anxious to her friends. A month later, she's avoiding situations she used to handle fine. What went wrong? In a new paper by my PhD student Dasha Sandra, titled "Why mental health awareness can harm: Converging explanations for a societal problem", we argue that well-meaning mental health awareness can backfire, and we identify how. Four separate literatures (concept creep, nocebo effects, prevalence inflation, and illness self-labeling) have been circling the same problem from different angles. We show they converge on three mechanisms: 1.Awareness lowers the threshold for what counts as a disorder. 2. It trains people to scan their inner lives for symptoms and reinterpret normal distress as pathology. 3. Once someone adopts an illness identity, they behave in ways that confirm and deepen it. The evidence is wide. Learning that loneliness is harmful makes solitude feel worse. Learning that stress is harmful worsens well-being and performance. Awareness videos about fake conditions like "wind turbine syndrome" produce real headaches. Trigger warnings raise anticipatory anxiety without reducing distress. This does not mean awareness should stop. It means awareness can have unintended consequences, including manufacturing the suffering it tries to prevent. Inoculating people against these mechanisms works, and we already have evidence it does. Link to paper: https://t.co/ucoGyhEuAj

Media 1
๐Ÿ–ผ๏ธ Media
A
ahall_research
@ahall_research
๐Ÿ“…
Apr 29, 2026
88d ago
๐Ÿ†”60885641

1800: If Thomas Jefferson is elected "Murder, robbery, rape, adultery, and incest will all be openly taught and practiced." When it comes to politics we have a bad habit of romanticizing the past and imagining that today's politics are worse and coarser. To make this visceral, I built a little app that shows what the 1800 election would have felt like if X had been around. Scrolling through it really does give you a sense that vicious, indecorous politics long pre-dates present day. Check it out here: https://t.co/eo2vFOf6TF

Media 2
๐Ÿ–ผ๏ธ Media
T
TawilTeddy
@TawilTeddy
๐Ÿ“…
May 01, 2026
86d ago
๐Ÿ†”11847556

Experts have three views on the future of work, each credible but sharply opposed. Whoโ€™s right? In a new paper for @CarnegieEndow & @CEIPTechProgram, I lay out the best arguments made by the alarmed, patient, and excited groups. ๐ŸงตOn the most important points and what policymakers can do today

Media 1
๐Ÿ–ผ๏ธ Media
J
justanotherlaw
@justanotherlaw
๐Ÿ“…
May 02, 2026
86d ago
๐Ÿ†”82155726

A recent viral paper claims to reverse-engineer the parameter counts of frontier models: GPT-5.5 = 9.7T, Opus 4.7 = 4.0T, o1 = 3.5T, etc. @ben_sturgeon and I investigated and found serious issues in the paper; fixing them gives GPT-5.5 as ~1.5T (90% CI: 256B-8.3T). https://t.co/MbWQyVlmsE

Media 1
๐Ÿ–ผ๏ธ Media
M
MatthewJBar
@MatthewJBar
๐Ÿ“…
May 05, 2026
82d ago
๐Ÿ†”68879935

Some people are being way too alarmist about Mythos. 80,000 Hours called Mythos "an AI that can break into almost any computer on Earth". Zvi Mowshowitz said, "If given to anyone with a credit card, Claude Mythos would give attackers a cornucopia of zero-day exploits for essentially all the software on Earth". But these descriptions are unfounded. @natalia__coelho shows in her latest blog post that the cyber capabilities of Mythos are nearly tied with GPT-5.5, across practically every public benchmark we have available. This includes both narrow and broad cyber evaluations. It is not way ahead of trend. The same holds for general capability evaluations. Except for a somewhat impressive score on SWE-bench Pro, Claude Mythos is nearly on-trend, at most a few months ahead. In other words, it's barely better than models that millions of people already have access to. It's an impressive model, but I'm very skeptical that Mythos is going to take down our digital infrastructure or cause a cyber catastrophe.

Media 1Media 2
+2 more
๐Ÿ–ผ๏ธ Media
R
random_walker
@random_walker
๐Ÿ“…
May 06, 2026
81d ago
๐Ÿ†”26218933

Excited to give this talk at the Stanford Digital Economy Lab on May 18! I will do three things: discuss my group's recent research, identify the most pressing gaps in the community's current understanding, and provide a long-term perspective. Hope to see you there in person or virtually. https://t.co/Qa2eNkVsnZ @DigEconLab

Media 1
๐Ÿ–ผ๏ธ Media
D
dongyangzi
@dongyangzi
๐Ÿ“…
May 05, 2026
82d ago
๐Ÿ†”26700234

We furthered AI research by reproducing CRUX #1 for Windows using @getnenai's infrastructure without needing to buy a Windows machine-- checkout our blog post https://t.co/Mjz1D5myF9

@sayashk โ€ข Thu Apr 16 17:49

Benchmarks are saturated more quickly than ever. How should frontier AI evaluations evolve? In a new paper, we argue that the AI community is already converging on an answer: Open-world evaluations. They are long, messy, real-world tasks that would be impractical for benchmarks.

Media 1
๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
Apr 24, 2026
93d ago
๐Ÿ†”37656389

We benchmarked GPT-5.5 on document understanding ๐Ÿ“„๐Ÿ“Š We ran it through ParseBench, our comprehensive OCR benchmark over enterprise documents. We evaluated metrics across various dimensions: visual grounding, tables, charts, and more. We evaluated GPT-5.5 on mid thinking and zero-thinking modes. When compared against GPT-5.4 (0 thinking) and Opus 4.7 (adaptive thinking): ๐Ÿ“ˆ GPT-5.5 wins on tables ๐Ÿ“ˆ GPT-5.5 wins on visual grounding ๐Ÿ“‰ GPT-5.5 0-thinking does worse on charts than GPT-5.4 0-thinking ๐Ÿ“‰ Higher thinking does worse than lower thinking of content faithfulness, semantic formatting ๐Ÿ“‰ Opus 4.7 wins overall on content faithfulness and semantic formatting ๐Ÿ’ธ GPT-5.5 is expensive: 13c per page at mid-thinking modes and 5.93 at 0-thinking! This is 5x the cost of any competitive OCR solution. Conclusion: GPT-5.5 is one of the better frontier models out there in terms of pure accuracy, but def not pound for pound w.r.t price.

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 27, 2026
90d ago
๐Ÿ†”21791003

Loan processors spend 40โ€“60% of their time reconciling income across tax returns, pay stubs, W-2s, and bank statements. We built an end-to-end pipeline that automates it with LlamaParse + the Claude Agent SDK: ๐Ÿ“„ Schema-driven extraction across 4 doc types with confidence scores + citations ๐Ÿ” Cross-document validation with Claude โ€” catches W-2/pay-stub gaps, unexplained Zelle/Venmo deposits, employer name mismatches ๐Ÿ“Š Self-contained HTML report with a COMPLETE / REVIEW / FLAG decision Full code + walkthrough: https://t.co/ozm4VwWmj3

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 28, 2026
89d ago
๐Ÿ†”16946011

Let's talk document formatting. Bold. Italics. Superscripts. Strikethroughs. The visual cues humans rely on every time we read a doc, and ones existing OCR benchmarks completely ignore. ๐Ÿ˜ฑ"$199" struck through next to "$149" isn't decoration. It's the meaning. ๐Ÿ˜ฑA superscript tells your agent "3" is a citation, not part of the number. Flatten that and your agent is reading a different doc than you are. Two weeks ago we released ParseBench, the first document OCR benchmark for AI agents. One of five metrics: the Semantic Formatting Score. Read more๐Ÿ‘‡ https://t.co/2sq5ncGiel

๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 29, 2026
88d ago
๐Ÿ†”90606809

Parsing documents with AI agents just got a lot more seamless๐Ÿš€ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ŸŒ Once connected, you'll be able to: ๐Ÿ“ Parse documents into clean markdown ๐Ÿ” Classify files against your own categories โœ‚๏ธ Split long documents into labelled sections โฌ†๏ธ Upload files via URL or a browser-based upload flow Building a production MCP server surfaced some non-obvious challenges: getting auth to align with an existing platform identity system using @WorkOS, working around MCP's lack of built-in file upload support, and making deployments, rate limiting and observability feel native with @vercel and @AxiomFM. We wrote up all of it, from the OAuth flow, to the token-based upload design, to the tradeoffs we hit along the way๐Ÿ“ ๐Ÿ“š Read the full blog: https://t.co/2E3qIVeUYI ๐Ÿ‘ฉโ€๐Ÿ’ป GitHub repository: https://t.co/ru1Evj7Zsr

๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
Apr 30, 2026
88d ago
๐Ÿ†”38085689

We shipped a LlamaParse MCP server to let you parse, classify, split, and generally operate over your hardest documents with your favorite AI agent ๐Ÿ“„๐Ÿค– Check out the MCP server: https://t.co/eOeDgzOCI5 This is both a useful feature and a great piece of engineering that we want to share with the community. ๐Ÿ’ก MCP does not have built-in file upload support, so we needed to implement a URL-based upload endpoint and couple it with parse operations ๐Ÿ’กWe built in an integration with @WorkOS OAuth. ๐Ÿ’กWe built in observability and rate-limiting Huge shoutout to @itsclelia for shipping this! Come check out our blog writeup: https://t.co/XsCZlKnzez LlamaParse: https://t.co/TqP6OT5U5O

@llama_index โ€ข Wed Apr 29 16:00

Parsing documents with AI agents just got a lot more seamless๐Ÿš€ We've rebuilt the LlamaParse MCP server to handle your document processing workflows, and you can connect it today to any MCP-compatible client at https://t.co/NF40qtKnQc ๐ŸŒ Once connected, you'll be able to: ๐Ÿ“ Pars

Media 2
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 30, 2026
87d ago
๐Ÿ†”79436469

Building scalable, distributed document processing pipelines isnโ€™t easy. Thatโ€™s why we teamed up with @render to build a system that: ๐Ÿ“ Leverages the LlamaParse platform to parse, classify, extract, and retrieve information from documents โš™๏ธ Uses Render Workflows to distribute tasks across nodes and accelerate background processing โšก Deploys a lightweight server and database on Render, giving you an instant interface to interact with your pipeline ๐Ÿ‘ฉโ€๐Ÿ’ป Explore the repo to see it in action: https://t.co/eiJqklNVhj ๐Ÿ“š And check out the step-by-step breakdown by @ojusave and @itsclelia: https://t.co/gePmuL7YtV

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 30, 2026
87d ago
๐Ÿ†”01071648

Thank you AI Dev Day '26 @DeepLearningAI @jerryjliu0 shares why SOTA LLMs can build an app but can't read a PDF ๐Ÿคฏ https://t.co/L2WMAUVaJ4

Media 1Media 2
+6 more
๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
May 03, 2026
84d ago
๐Ÿ†”42086427

Parsing PDFs is hard This past week I gave a few talks (at both AI Dev '26 by @DeepLearningAI and @Capgemini ) on why this is still such an open problem, and itโ€™s even more important as agents become the consumers of documents, and need the OCR tools to read them properly. The fundamental issue is that PDFs are designed for print and display purposes, not to give back a linearized, semantically meaningful string of text. Text and tables are represented as a bunch of chars and lines, without any guaranteed order. This is what the community is solving with VLM-based approaches, including our own efforts around LlamaParse and ParseBench. If youโ€™re interested in learning more about the problem, check out the blog post I wrote on this a while ago! https://t.co/740ZiAFyOk

Media 1Media 2
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
May 05, 2026
82d ago
๐Ÿ†”99831547

๐ŸŽ‰ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

Media 1Media 2
๐Ÿ–ผ๏ธ Media
J
jerryjliu0
@jerryjliu0
๐Ÿ“…
May 05, 2026
82d ago
๐Ÿ†”37779563

Iโ€™m excited to announce that @llama_index is on the @CBInsights AI 100 list for 2026 ๐Ÿ”ฅ Weโ€™re on a mission to parse all of the worldโ€™s PDFs, and make them accessible to both humans and AI agents. List: https://t.co/8pKDaPnHI8 If you havenโ€™t done so already, our website design is awesome, check it out: https://t.co/YiIfjVlzb6

@llama_index โ€ข Tue May 05 15:06

๐ŸŽ‰ @CBinsights AI 100 2026 is out and LlamaIndex made the list. We're proud to provide the leading document understanding API for AI agents. Congrats to all honorees in the AI Infrastructure category. Full list here: https://t.co/2VWBboucSN https://t.co/IBc4dQqRno

Media 1Media 2
+1 more
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
May 06, 2026
81d ago
๐Ÿ†”54837865

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powered by the LlamaParse TypeScript SDK ๐Ÿ“ฑ Three steps, thatโ€™s it: ๐Ÿ”‘ Add your API key (securely stored on-device) ๐Ÿ“ธ Snap a photo of anything with text ๐Ÿ“„ Parse it and, in under a minute, get clean, copyable text No hassle, no manual typing. ๐Ÿš€ Try it now: https://t.co/mEeraV4zap ๐Ÿฆ™ Get started with LlamaParse: https://t.co/Pmc8usyjTwโ€“

๐Ÿ–ผ๏ธ Media