Your curated collection of saved posts and media

Showing 9 posts · last 14 days · by score
➕ Add New Post
F
François Chollet
@fchollet
📅
Aug 21, 2026
11d ago
🆔37645398
0.40

This is very nice work from NVIDIA. Like all high-performing approaches on ARC-AGI-3, it uses deep learning-guided on-the-fly synthesis of symbolic world models, i.e. navigating the world by generating programs to represent what you know. To be clear, like with several other recent claims, scoring 100% on the public demonstration set is not the same as "scoring 100% on the ARC-AGI-3 benchmark". It would be like saying you beat a videogame because you cleared the tutorial level.

@NVIDIAAI • Fri Aug 21 13:05

NVIDIA AVO continuously inspects, plans, implements, and evaluates, using memory, tools, and execution feedback to build on what it learns along the way. This allows the system to sustain progress across long-running tasks rather than starting over with each model context. Read

❤️452
likes
🔁35
retweets
A
Aravind Srinivas
@AravSrinivas
📅
Aug 25, 2026
7d ago
🆔71598820
0.40

In a compute and power-constrained world, a good chunk of agentic inference needs to move to local hardware. A drastic version of that is a fully local agent runtime, where the model (orchestrator and subagents) and the harness run locally. Portable Computer from Perplexity is this. Launching today for @nvidia DGX Spark.

@perplexity_ai • Tue Aug 25 15:10

Today we’re launching Portable Computer on @NVIDIA DGX Spark. Portable Computer is a fully local version of Perplexity Computer, where the entire runtime: orchestrator LLM, subagent LLM, agent harness all run on your local hardware. No cloud dependency. https://t.co/plVWz5PaAw

❤️352
likes
🔁21
retweets
Q
Xiaoyin Qu
@quxiaoyin
📅
Aug 24, 2026
8d ago
🆔09553036

I ran the same task on Claude Code and DeepSeek's new agent harness. One cost $150. The other cost $2. Today we're launching https://t.co/twx6etZb3X (@agentsky_dev), the "OpenRouter for Agents" — one API → Claude Code, Codex, DeepSeek, Kimi, OpenCode, and every major agent in the cloud. And Agent Playground on top: race them on your own task, with your real tools (GitHub, Gmail, more), side by side in a browser: time, cost, tokens burnt. Guess which one was $2.

Media 2
❤️344
likes
🔁99
retweets
🖼️ Media
J
Joe Hansen
@joehansen
📅
Aug 21, 2026
10d ago
🆔95211147

Grok 4.6 just took #1 on CursorBench. Not only the highest score. It did it at a fraction of the cost of the models sitting right behind it. That combination is the real signal. Top-tier results are one thing. Top-tier results that stay cheap enough to run for long agentic coding sessions are something else. I see this as the practical edge that matters for real work. Benchmarks are useful. Sustained performance at low cost is what actually gets used.

@elonmusk • Fri Aug 21 16:36

Grok 4.6 on extra high thinking mode now achieves #1 score on CursorBench!

Media 1
❤️293
likes
🔁58
retweets
🖼️ Media
E
Steven Brunton
@eigensteve
📅
Aug 20, 2026
11d ago
🆔52811460

HydroGym: A Reinforcement Learning Platform for Fluid Dynamics Now published in Nature!! https://t.co/o1zTn2DIHg GitHub: https://t.co/9MLBzJAMiJ Amazing collaboration with Christian Lagemann, S Mokbel, M Gondrum, M Rüttgers, Y Wang, P Suárez, L Paehler, D A Bezgin, A B Buhendwa, J L Callaham, S Ahnert, N Zolman, X Shao, J-Ch Loiseau, N A. Adams, M Meinke, W Schröder, K Lagemann, E Lagemann, R Vinuesa & S L Brunton

Media 1Media 2
❤️226
likes
🔁37
retweets
🖼️ Media
🔁AK retweeted
E
Einsia
@EinsiaAI
📅
Aug 21, 2026
11d ago
🆔01771909
0.34

1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are trained —not just tuning hyperparameters, but improving the training algorithm itself. We tested this directly with AI4AI-Bench: 10 real research repositories spanning 10 distinct algorithm families. Full breakdown 👇 GitHub: [https://t.co/s0f0NY7PdQ] Paper Link: [https://t.co/x0qY8wnlwB] Einsia Website:[https://t.co/Rewt4FJwl8] 📊 The results: The average score is just 0.166. Even the best-performing model, Opus 5, reaches only 0.288. The median exploration cost per task rises from $1.69 to $34.60. #AI4AI #RecursiveSelfImprovement #AIResearch #AI4AI_Bench

❤️205
likes
🔁71
retweets
N
Niels Rogge
@NielsRogge
📅
Aug 31, 2026
1d ago
🆔06339969

For folks wondering what Sliding Window Attention is, there's a method for it on Papers with Code Sliding Window Attention (SWA): A local attention pattern that restricts each token to attending only within a fixed-size neighborhood instead of the full sequence. This reduces attention and KV-cache memory for long-context models, while periodic global-attention layers can preserve broader context. Find it here: https://t.co/K1MhZVasL8

@jm_alexia • Mon Aug 31 13:19

Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: https://t.co/h8DIc223Su

Media 1Media 2
❤️196
likes
🔁25
retweets
🖼️ Media
I
Luke The Dev
@iamlukethedev
📅
Aug 27, 2026
4d ago
🆔23456637

Hermes HUD mode just became my live stock analyst I opened a chart and asked Hermes to analyze what it was seeing Instead of only explaining it in text, HUD drew directly over my screen: • Resistance • Support • Trend direction • Current price context So now Hermes can look at the same chart I’m looking at and annotate the analysis in real time. This is exactly what I wanted HUD mode to become. Not another chat window. An AI layer on top of whatever I’m already doing 👀 https://t.co/uAbJm94AQ3

@iamlukethedev • Wed Aug 26 23:36

Hermes HUD mode can now translate movies live while you watch No pausing No copying subtitles No switching apps HUD watches what is on your screen and translates it in real time as the movie plays This is exactly why I love building on HUD mode Hermes does not pull you away f

Media 2
❤️189
likes
🔁13
retweets
🖼️ Media
🔁LlamaIndex 🦙 retweeted
J
Jerry Liu
@jerryjliu0
📅
Aug 23, 2026
9d ago
🆔22077885
0.32

The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2️⃣ A “just-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the “out of the box” doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within @llama_index to help any agent do two-pass document processing with higher accuracy and lower cost. 1️⃣ We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types 2️⃣ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a “zoom-in” pass. Come check it out! LiteParse: https://t.co/JNER0mVcB8 LlamaParse: https://t.co/XYZmx5TFz8 All the relevant docs, including MCP, are here: https://t.co/qc9Q5NT3Jr

❤️174
likes
🔁25
retweets
← PreviousPage 5 of 149Next →