Your curated collection of saved posts and media
I made portals with some open cv and projection mapping :) https://t.co/TbQ88YktpD
Gemini 3.7 Flash from @Google on ARC-AGI (Verified): - ARC-AGI-2: 84.6%, $0.25/task - ARC-AGI-1: 95.5%, $0.12/task Gemini 3.7 Flash stands out for its low cost and high scores on ARC-AGI-1 and ARC-AGI-2 relative to other frontier models. https://t.co/7cY43PW7Db
Today we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models. RL training is easy to start and hard to debug. Miles helps you ensure your run is correct, use hardware efficiently, and keep RL running at scale. Over the past 9 months, 72 contributors have landed 1,326 commits, 85 GPU E2E CI tests, battle-testing Miles on frontier open models like Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, MiniMax H3, etc. Miles powers frontier-model development and production RL workloads at @humansand, @periodiclabs, @modal, @DecagonAI, @Eigent_AI, @nebiusai, @IBM and more, on both @NVIDIAAI and @AIatAMD hardware. Here is what we built, and why teams picked Milesπ§΅
also got access to it, it's still in training and i found the wandb this is crazy, here is the training loss π€― https://t.co/olrWjgdunA https://t.co/Ax2SnyaJaV
Mind blown from a new model I just got access to. I think this will be one of the most (the most?) significant drops this year. Excited. And sorry to be annoyingly vague. Just excited.

trained my own with trl + openenv blog with open artifacts (code, models, dataset, rl env...) soon! https://t.co/HTxA9HLRv9
You can just RL a coding model to paint with javascript btw https://t.co/4x5B81kjUh
Meet Portable Computer, Perplexity's new local-first agent stack on NVIDIA DGX Spark. When running locally, Portable Computer offers one-click local inference setup and an optimized agentic experience for DGX Spark. Learn more and get started today: https://t.co/0UElxhCnkY https://t.co/W07ZafKYdC
I got the MicroDuck π¦ back flipping clean! All trained on my MacBook Pro. Going to open source my repo soon. https://t.co/X8wVcrQxWJ
Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on @huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task. https://t.co/w6OCWyWRi5
This is very nice work from NVIDIA. Like all high-performing approaches on ARC-AGI-3, it uses deep learning-guided on-the-fly synthesis of symbolic world models, i.e. navigating the world by generating programs to represent what you know. To be clear, like with several other recent claims, scoring 100% on the public demonstration set is not the same as "scoring 100% on the ARC-AGI-3 benchmark". It would be like saying you beat a videogame because you cleared the tutorial level.
NVIDIA AVO continuously inspects, plans, implements, and evaluates, using memory, tools, and execution feedback to build on what it learns along the way. This allows the system to sustain progress across long-running tasks rather than starting over with each model context. Read
In a compute and power-constrained world, a good chunk of agentic inference needs to move to local hardware. A drastic version of that is a fully local agent runtime, where the model (orchestrator and subagents) and the harness run locally. Portable Computer from Perplexity is this. Launching today for @nvidia DGX Spark.
Today weβre launching Portable Computer on @NVIDIA DGX Spark. Portable Computer is a fully local version of Perplexity Computer, where the entire runtime: orchestrator LLM, subagent LLM, agent harness all run on your local hardware. No cloud dependency. https://t.co/plVWz5PaAw
I ran the same task on Claude Code and DeepSeek's new agent harness. One cost $150. The other cost $2. Today we're launching https://t.co/twx6etZb3X (@agentsky_dev), the "OpenRouter for Agents" β one API β Claude Code, Codex, DeepSeek, Kimi, OpenCode, and every major agent in the cloud. And Agent Playground on top: race them on your own task, with your real tools (GitHub, Gmail, more), side by side in a browser: time, cost, tokens burnt. Guess which one was $2.
Grok 4.6 just took #1 on CursorBench. Not only the highest score. It did it at a fraction of the cost of the models sitting right behind it. That combination is the real signal. Top-tier results are one thing. Top-tier results that stay cheap enough to run for long agentic coding sessions are something else. I see this as the practical edge that matters for real work. Benchmarks are useful. Sustained performance at low cost is what actually gets used.
Grok 4.6 on extra high thinking mode now achieves #1 score on CursorBench!
Today, we're introducing AC2, the Applied Compute Agent Cloud, to enable every team to train, serve, and improve their own frontier models. https://t.co/pvqq2t8HHY
1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are trained βnot just tuning hyperparameters, but improving the training algorithm itself. We tested this directly with AI4AI-Bench: 10 real research repositories spanning 10 distinct algorithm families. Full breakdown π GitHub: [https://t.co/s0f0NY7PdQ] Paper Link: [https://t.co/x0qY8wnlwB] Einsia Website:[https://t.co/Rewt4FJwl8] π The results: The average score is just 0.166. Even the best-performing model, Opus 5, reaches only 0.288. The median exploration cost per task rises from $1.69 to $34.60. #AI4AI #RecursiveSelfImprovement #AIResearch #AI4AI_Bench
For folks wondering what Sliding Window Attention is, there's a method for it on Papers with Code Sliding Window Attention (SWA): A local attention pattern that restricts each token to attending only within a fixed-size neighborhood instead of the full sequence. This reduces attention and KV-cache memory for long-context models, while periodic global-attention layers can preserve broader context. Find it here: https://t.co/K1MhZVasL8
Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: https://t.co/h8DIc223Su

Hermes HUD mode just became my live stock analyst I opened a chart and asked Hermes to analyze what it was seeing Instead of only explaining it in text, HUD drew directly over my screen: β’ Resistance β’ Support β’ Trend direction β’ Current price context So now Hermes can look at the same chart Iβm looking at and annotate the analysis in real time. This is exactly what I wanted HUD mode to become. Not another chat window. An AI layer on top of whatever Iβm already doing π https://t.co/uAbJm94AQ3
Hermes HUD mode can now translate movies live while you watch No pausing No copying subtitles No switching apps HUD watches what is on your screen and translates it in real time as the movie plays This is exactly why I love building on HUD mode Hermes does not pull you away f
The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1οΈβ£ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1kβs of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2οΈβ£ A βjust-in-timeβ VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that itβs slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context thatβs needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the βout of the boxβ doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool wouldβve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within @llama_index to help any agent do two-pass document processing with higher accuracy and lower cost. 1οΈβ£ We have liteparse for the first pass - a free/OSS parser written in Rust thatβs faster/more accurate than other OSS parsers, and supports 50+ document types 2οΈβ£ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a βzoom-inβ pass. Come check it out! LiteParse: https://t.co/JNER0mVcB8 LlamaParse: https://t.co/XYZmx5TFz8 All the relevant docs, including MCP, are here: https://t.co/qc9Q5NT3Jr
we've been iterating a lot on how DocWriter internally represents a user's writing style! the naive approach is to dump all a user's prior writing in context and pray the AI figures out how to sound like them. this doesn't really work. and the user can't just manually specify their style either, because nobody knows what their meaningful writing preferences (for an AI agent) actually are. so we've built a multi-agent pipeline that runs lexical, grammatical, and discourse analyses on the user's writing, extracts specific patterns with evidence, and then surfaces them in a UI where the user can confirm or reject each pattern by comparing two passages side by side TBD on how well it performs in our next round of user studies π it feels like we are really trying to do anything possible to avoid fine-tuning, but perhaps we need to bite the bullet eventually
Connecting AI to hardware requires days or weeks of bespoke integration, with no standard way for agents to operate equipment safely. MHS cuts integration to hours or minutes, provides an interface that makes devices discoverable, and enables agents to operate them safely.
We need to normalize measuring and judging models against a standardized test harness "Oh but model X performs best in their own proprietary harness" I could not care less. When I take exams, I go to the standardized classroom, get the standardized pencil and exam sheet, and have to solve it under 2 hours This system arose because we have a LOT of people to test Guess what? We now have a LOT of models, and they are multiplying by the day "Oh but model X performs substantially better in ARC-AGI-3 with a custom harness" I don't care... Then imbue model X with enough knowledge so that it can reconstruct that harness on the spot The main harness could be mini-swe-agent, terminus 2, vanilla pi or something along those lines It needs to be simple, and stay roughly the same over time There is already too much complexity in the benchmarking space right now, and I feel like not enough people are putting their feet down to cut away some part of it
π§΅Scaling Laws for T2I Diffusion We do a scaling laws ladder for t2i diffusion models all the way from 60M to 2B, spanning 3 orders of magnitude in training flops https://t.co/GtxmABJAcq
Grok 4.6 just took the #1 spot on CursorBench 3.2.....and the efficiency is insane Here's the cost comparison: β’ Grok 4.6 Extra High β 70.8% | $2.81/task β’ Fable 5 Max β 70.5% | $17.32/task β’ Opus 5 Max β 70.0% | $8.23/task β’ GPT-5.6 Sol Max β 67.2% | $5.69/task Grok achieved the highest score while costing roughly 6X less than Fable 5 Max and nearly 3X less than Opus 5 Max per task Thatβs what makes Grok so powerful for agents Top-tier intelligence is great.....but top-tier intelligence that can keep working across long coding tasks without burning ridiculous amounts of compute is even better Grokβs agentic coding efficiency is insane
π New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread π§΅ https://t.co/zRWY3j3ivJ
Great new paper from AWS on agent handoff tax. If you build agents today, you need to understand the so-called handoff tax. (bookmark it) Escalating to a stronger model mid-run is usually the resort when a cheap agent stalls. New work from AWS AI Labs measures how much that switch actually costs. Coding agents run for dozens of model calls, so teams escalate when a weak model struggles and downshift once the hard reasoning is done. Every switch forces the receiving model to continue a trajectory another model wrote. Across pairs of Claude and GPT models, full-trajectory escalation recovers less than half the quality gap between the weak and strong model while adding a substantial cost premium. The authors call that penalty the handoff tax. Downshifting lands at a much better cost-quality point. Cutting the weak model's trajectory information improves escalation quality, while removing the strong model's trajectory hurts downshift quality. Paper: https://t.co/59ozKgugB8 Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
Banger paper from Stanford on efficient test-time scaling. If you run agents that think for a long time, this one is worth your time. (bookmark it) Long reasoning keeps the entire trace in memory through full attention. This means that the hardest problems, the ones that need the most thinking, are also the ones that cost the most to run. The authors measured what the middle of a reasoning trace is actually worth. Intermediate tokens steadily lose importance as the model keeps going. Their new approach, Prefix Sliding, drops those tokens. It keeps the prefix, which holds the instructions and the available tools, plus a window of the last few thousand tokens. Everything in between gets discarded during generation. Total memory stays capped no matter how long the model reasons. Without any training, this runs existing models 3x faster while matching full-attention performance, and it enables RL rollouts past 100,000 tokens. Paper: https://t.co/HzwSZ7fCdh Chat with Paper: https://t.co/OfjtVjamIC
a new open-weight model family where the model card ships with Hermes Agent setup out of the box. point your Hermes Agent at a local Ornith 1.5 server, 9B to 397B! benchmark coming soon :)
new paper: Prefix Sliding for efficient test-time scaling vanilla full attention OOMs on long tasks & compaction loses important details -- prefix sliding is a simple & fast alternative that can outperform both πhttps://t.co/fUw7yJAN5D https://t.co/PgcyrsfIZb
"ZipSplat: Fewer Gaussians, Better Splats" TL;DR: feed-forward 3DGS model that decouples Gaussian placement from pixels, reconstructing unposed scenes in under a second with ~6Γ fewer Gaussians while achieving state-of-the-art quality. https://t.co/T6YmrQqlHy
GPT-5.4 xhigh scored 53 on Artificial Analysis in March. By August, Qwen3.8-Flash-Next scores 56 and GLM-5.3-Flash 57 with only 6B / 18B active params per token. Yesterdayβs frontier is todayβs Flash tier.
I'm implementing a tiny transformer on tinyshakespear. The 2d matrices went to muon and 1d to adam. The model was still learning and generating some real words. But turns out my muon implementation was bugged and was no-oping. So 99.2% of my model was frozen at init. Adam still managed to tweak those biases into having the model still output some real words. Oh and I forgot the positional embeddings too. It's kinda crazy how you can have the shittiest implementation and a neural network still manages to learn
Karpathy's recipe for training neural networks is still relevant today. this lesson in particular is one we've been feeling very viscerally recently... neural network training can sometimes be very resilient and you may not realize there's an error for a very long time... https
I guess this is as good a way as any to announce that I'm now the Head of Evals at @every https://t.co/gPWYBoNrBY
