Your curated collection of saved posts and media

Showing 32 posts Β· last 14 days Β· by score
B
Brendan Johnston
@builtwbrendan
πŸ“…
Aug 26, 2026
5d ago
πŸ†”37790930

I made portals with some open cv and projection mapping :) https://t.co/TbQ88YktpD

❀️1,576
likes
πŸ”95
retweets
πŸ–ΌοΈ Media
πŸ”Ivan Leo retweeted
A
ARC Prize
@arcprize
πŸ“…
Aug 20, 2026
11d ago
πŸ†”50539327
⭐0.34

Gemini 3.7 Flash from @Google on ARC-AGI (Verified): - ARC-AGI-2: 84.6%, $0.25/task - ARC-AGI-1: 95.5%, $0.12/task Gemini 3.7 Flash stands out for its low cost and high scores on ARC-AGI-1 and ARC-AGI-2 relative to other frontier models. https://t.co/7cY43PW7Db

❀️1,348
likes
πŸ”96
retweets
R
RadixArk
@radixark
πŸ“…
Aug 18, 2026
14d ago
πŸ†”39384068

Today we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models. RL training is easy to start and hard to debug. Miles helps you ensure your run is correct, use hardware efficiently, and keep RL running at scale. Over the past 9 months, 72 contributors have landed 1,326 commits, 85 GPU E2E CI tests, battle-testing Miles on frontier open models like Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, MiniMax H3, etc. Miles powers frontier-model development and production RL workloads at @humansand, @periodiclabs, @modal, @DecagonAI, @Eigent_AI, @nebiusai, @IBM and more, on both @NVIDIAAI and @AIatAMD hardware. Here is what we built, and why teams picked Miles🧡

❀️578
likes
πŸ”89
retweets
πŸ–ΌοΈ Media
E
elie
@eliebakouch
πŸ“…
Aug 24, 2026
8d ago
πŸ†”58569854

also got access to it, it's still in training and i found the wandb this is crazy, here is the training loss 🀯 https://t.co/olrWjgdunA https://t.co/Ax2SnyaJaV

@martin_casado β€’ Sun Aug 23 22:17

Mind blown from a new model I just got access to. I think this will be one of the most (the most?) significant drops this year. Excited. And sorry to be annoyingly vague. Just excited.

Media 1Media 2
❀️573
likes
πŸ”25
retweets
πŸ–ΌοΈ Media
S
Sergio Paniego
@SergioPaniego
πŸ“…
Aug 30, 2026
1d ago
πŸ†”30894045

trained my own with trl + openenv blog with open artifacts (code, models, dataset, rl env...) soon! https://t.co/HTxA9HLRv9

@kickingkeys β€’ Sun Aug 23 16:59

You can just RL a coding model to paint with javascript btw https://t.co/4x5B81kjUh

❀️547
likes
πŸ”29
retweets
πŸ–ΌοΈ Media
πŸ”Aravind Srinivas retweeted
N
NVIDIA
@nvidia
πŸ“…
Aug 25, 2026
7d ago
πŸ†”86126575
⭐0.32

Meet Portable Computer, Perplexity's new local-first agent stack on NVIDIA DGX Spark. When running locally, Portable Computer offers one-click local inference setup and an optimized agentic experience for DGX Spark. Learn more and get started today: https://t.co/0UElxhCnkY https://t.co/W07ZafKYdC

❀️515
likes
πŸ”64
retweets
J
jonathanhawkins
@jonathanhawkins
πŸ“…
Aug 31, 2026
21h ago
πŸ†”90192385

I got the MicroDuck πŸ¦† back flipping clean! All trained on my MacBook Pro. Going to open source my repo soon. https://t.co/X8wVcrQxWJ

❀️514
likes
πŸ”31
retweets
πŸ–ΌοΈ Media
πŸ”Robert Scoble retweeted
K
kwindla
@kwindla
πŸ“…
Aug 27, 2026
5d ago
πŸ†”47339026
⭐0.38

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on @huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

❀️512
likes
πŸ”51
retweets
T
tobi lutke
@tobi
πŸ“…
Sep 01, 2026
1h ago
πŸ†”55191249

Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task. https://t.co/w6OCWyWRi5

Media 1
❀️474
likes
πŸ”34
retweets
πŸ–ΌοΈ Media
F
FranΓ§ois Chollet
@fchollet
πŸ“…
Aug 21, 2026
11d ago
πŸ†”37645398
⭐0.40

This is very nice work from NVIDIA. Like all high-performing approaches on ARC-AGI-3, it uses deep learning-guided on-the-fly synthesis of symbolic world models, i.e. navigating the world by generating programs to represent what you know. To be clear, like with several other recent claims, scoring 100% on the public demonstration set is not the same as "scoring 100% on the ARC-AGI-3 benchmark". It would be like saying you beat a videogame because you cleared the tutorial level.

@NVIDIAAI β€’ Fri Aug 21 13:05

NVIDIA AVO continuously inspects, plans, implements, and evaluates, using memory, tools, and execution feedback to build on what it learns along the way. This allows the system to sustain progress across long-running tasks rather than starting over with each model context. Read

❀️452
likes
πŸ”35
retweets
A
Aravind Srinivas
@AravSrinivas
πŸ“…
Aug 25, 2026
7d ago
πŸ†”71598820
⭐0.40

In a compute and power-constrained world, a good chunk of agentic inference needs to move to local hardware. A drastic version of that is a fully local agent runtime, where the model (orchestrator and subagents) and the harness run locally. Portable Computer from Perplexity is this. Launching today for @nvidia DGX Spark.

@perplexity_ai β€’ Tue Aug 25 15:10

Today we’re launching Portable Computer on @NVIDIA DGX Spark. Portable Computer is a fully local version of Perplexity Computer, where the entire runtime: orchestrator LLM, subagent LLM, agent harness all run on your local hardware. No cloud dependency. https://t.co/plVWz5PaAw

❀️352
likes
πŸ”21
retweets
Q
Xiaoyin Qu
@quxiaoyin
πŸ“…
Aug 24, 2026
8d ago
πŸ†”09553036

I ran the same task on Claude Code and DeepSeek's new agent harness. One cost $150. The other cost $2. Today we're launching https://t.co/twx6etZb3X (@agentsky_dev), the "OpenRouter for Agents" β€” one API β†’ Claude Code, Codex, DeepSeek, Kimi, OpenCode, and every major agent in the cloud. And Agent Playground on top: race them on your own task, with your real tools (GitHub, Gmail, more), side by side in a browser: time, cost, tokens burnt. Guess which one was $2.

Media 2
❀️344
likes
πŸ”99
retweets
πŸ–ΌοΈ Media
J
Joe Hansen
@joehansen
πŸ“…
Aug 21, 2026
10d ago
πŸ†”95211147

Grok 4.6 just took #1 on CursorBench. Not only the highest score. It did it at a fraction of the cost of the models sitting right behind it. That combination is the real signal. Top-tier results are one thing. Top-tier results that stay cheap enough to run for long agentic coding sessions are something else. I see this as the practical edge that matters for real work. Benchmarks are useful. Sustained performance at low cost is what actually gets used.

@elonmusk β€’ Fri Aug 21 16:36

Grok 4.6 on extra high thinking mode now achieves #1 score on CursorBench!

Media 1
❀️293
likes
πŸ”58
retweets
πŸ–ΌοΈ Media
A
Applied Compute
@appliedcompute
πŸ“…
Aug 25, 2026
6d ago
πŸ†”57319367

Today, we're introducing AC2, the Applied Compute Agent Cloud, to enable every team to train, serve, and improve their own frontier models. https://t.co/pvqq2t8HHY

❀️255
likes
πŸ”29
retweets
πŸ–ΌοΈ Media
πŸ”AK retweeted
E
Einsia
@EinsiaAI
πŸ“…
Aug 21, 2026
10d ago
πŸ†”01771909
⭐0.34

1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are trained β€”not just tuning hyperparameters, but improving the training algorithm itself. We tested this directly with AI4AI-Bench: 10 real research repositories spanning 10 distinct algorithm families. Full breakdown πŸ‘‡ GitHub: [https://t.co/s0f0NY7PdQ] Paper Link: [https://t.co/x0qY8wnlwB] Einsia Website:[https://t.co/Rewt4FJwl8] πŸ“Š The results: The average score is just 0.166. Even the best-performing model, Opus 5, reaches only 0.288. The median exploration cost per task rises from $1.69 to $34.60. #AI4AI #RecursiveSelfImprovement #AIResearch #AI4AI_Bench

❀️205
likes
πŸ”71
retweets
N
Niels Rogge
@NielsRogge
πŸ“…
Aug 31, 2026
1d ago
πŸ†”06339969

For folks wondering what Sliding Window Attention is, there's a method for it on Papers with Code Sliding Window Attention (SWA): A local attention pattern that restricts each token to attending only within a fixed-size neighborhood instead of the full sequence. This reduces attention and KV-cache memory for long-context models, while periodic global-attention layers can preserve broader context. Find it here: https://t.co/K1MhZVasL8

@jm_alexia β€’ Mon Aug 31 13:19

Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: https://t.co/h8DIc223Su

Media 1Media 2
❀️196
likes
πŸ”25
retweets
πŸ–ΌοΈ Media
I
Luke The Dev
@iamlukethedev
πŸ“…
Aug 27, 2026
4d ago
πŸ†”23456637

Hermes HUD mode just became my live stock analyst I opened a chart and asked Hermes to analyze what it was seeing Instead of only explaining it in text, HUD drew directly over my screen: β€’ Resistance β€’ Support β€’ Trend direction β€’ Current price context So now Hermes can look at the same chart I’m looking at and annotate the analysis in real time. This is exactly what I wanted HUD mode to become. Not another chat window. An AI layer on top of whatever I’m already doing πŸ‘€ https://t.co/uAbJm94AQ3

@iamlukethedev β€’ Wed Aug 26 23:36

Hermes HUD mode can now translate movies live while you watch No pausing No copying subtitles No switching apps HUD watches what is on your screen and translates it in real time as the movie plays This is exactly why I love building on HUD mode Hermes does not pull you away f

Media 2
❀️189
likes
πŸ”13
retweets
πŸ–ΌοΈ Media
J
Jerry Liu
@jerryjliu0
πŸ“…
Aug 23, 2026
9d ago
πŸ†”22077885

The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2️⃣ A β€œjust-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the β€œout of the box” doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within @llama_index to help any agent do two-pass document processing with higher accuracy and lower cost. 1️⃣ We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types 2️⃣ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a β€œzoom-in” pass. Come check it out! LiteParse: https://t.co/JNER0mVcB8 LlamaParse: https://t.co/XYZmx5TFz8 All the relevant docs, including MCP, are here: https://t.co/qc9Q5NT3Jr

Media 2
+1 more
❀️174
likes
πŸ”25
retweets
πŸ–ΌοΈ Media
S
Shreya Shankar
@sh_reya
πŸ“…
Aug 22, 2026
9d ago
πŸ†”51549617

we've been iterating a lot on how DocWriter internally represents a user's writing style! the naive approach is to dump all a user's prior writing in context and pray the AI figures out how to sound like them. this doesn't really work. and the user can't just manually specify their style either, because nobody knows what their meaningful writing preferences (for an AI agent) actually are. so we've built a multi-agent pipeline that runs lexical, grammatical, and discourse analyses on the user's writing, extracts specific patterns with evidence, and then surfaces them in a UI where the user can confirm or reject each pattern by comparing two passages side by side TBD on how well it performs in our next round of user studies πŸ˜† it feels like we are really trying to do anything possible to avoid fine-tuning, but perhaps we need to bite the bullet eventually

Media 1
❀️167
likes
πŸ”7
retweets
πŸ–ΌοΈ Media
A
Anthropic
@AnthropicAI
πŸ“…
Aug 27, 2026
4d ago
πŸ†”42134590
⭐0.36

Connecting AI to hardware requires days or weeks of bespoke integration, with no standard way for agents to operate equipment safely. MHS cuts integration to hours or minutes, provides an interface that makes devices discoverable, and enables agents to operate them safely.

❀️151
likes
πŸ”3
retweets
O
Onur Solmaz
@onusoz
πŸ“…
Aug 23, 2026
9d ago
πŸ†”51267162
⭐0.42

We need to normalize measuring and judging models against a standardized test harness "Oh but model X performs best in their own proprietary harness" I could not care less. When I take exams, I go to the standardized classroom, get the standardized pencil and exam sheet, and have to solve it under 2 hours This system arose because we have a LOT of people to test Guess what? We now have a LOT of models, and they are multiplying by the day "Oh but model X performs substantially better in ARC-AGI-3 with a custom harness" I don't care... Then imbue model X with enough knowledge so that it can reconstruct that harness on the spot The main harness could be mini-swe-agent, terminus 2, vanilla pi or something along those lines It needs to be simple, and stay roughly the same over time There is already too much complexity in the benchmarking space right now, and I feel like not enough people are putting their feet down to cut away some part of it

❀️149
likes
πŸ”3
retweets
πŸ”Jonathan Whitaker retweeted
S
sway
@SwayStar123
πŸ“…
Aug 19, 2026
13d ago
πŸ†”72447314
⭐0.34

🧡Scaling Laws for T2I Diffusion We do a scaling laws ladder for t2i diffusion models all the way from 60M to 2B, spanning 3 orders of magnitude in training flops https://t.co/GtxmABJAcq

❀️133
likes
πŸ”25
retweets
X
X Freeze
@XFreeze
πŸ“…
Aug 21, 2026
11d ago
πŸ†”85377458

Grok 4.6 just took the #1 spot on CursorBench 3.2.....and the efficiency is insane Here's the cost comparison: β€’ Grok 4.6 Extra High β€” 70.8% | $2.81/task β€’ Fable 5 Max β€” 70.5% | $17.32/task β€’ Opus 5 Max β€” 70.0% | $8.23/task β€’ GPT-5.6 Sol Max β€” 67.2% | $5.69/task Grok achieved the highest score while costing roughly 6X less than Fable 5 Max and nearly 3X less than Opus 5 Max per task That’s what makes Grok so powerful for agents Top-tier intelligence is great.....but top-tier intelligence that can keep working across long coding tasks without burning ridiculous amounts of compute is even better Grok’s agentic coding efficiency is insane

Media 1
❀️124
likes
πŸ”19
retweets
πŸ–ΌοΈ Media
T
tomaarsen
@tomaarsen
πŸ“…
Aug 26, 2026
6d ago
πŸ†”90713066

πŸ“ˆ New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread 🧡 https://t.co/zRWY3j3ivJ

Media 1
❀️115
likes
πŸ”14
retweets
πŸ–ΌοΈ Media
πŸ”elvis retweeted
O
elvis
@omarsar0
πŸ“…
Aug 26, 2026
6d ago
πŸ†”17953811
⭐0.34

Great new paper from AWS on agent handoff tax. If you build agents today, you need to understand the so-called handoff tax. (bookmark it) Escalating to a stronger model mid-run is usually the resort when a cheap agent stalls. New work from AWS AI Labs measures how much that switch actually costs. Coding agents run for dozens of model calls, so teams escalate when a weak model struggles and downshift once the hard reasoning is done. Every switch forces the receiving model to continue a trajectory another model wrote. Across pairs of Claude and GPT models, full-trajectory escalation recovers less than half the quality gap between the weak and strong model while adding a substantial cost premium. The authors call that penalty the handoff tax. Downshifting lands at a much better cost-quality point. Cutting the weak model's trajectory information improves escalation quality, while removing the strong model's trajectory hurts downshift quality. Paper: https://t.co/59ozKgugB8 Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX

❀️101
likes
πŸ”14
retweets
πŸ”elvis retweeted
O
elvis
@omarsar0
πŸ“…
Aug 30, 2026
1d ago
πŸ†”04099888
⭐0.34

Banger paper from Stanford on efficient test-time scaling. If you run agents that think for a long time, this one is worth your time. (bookmark it) Long reasoning keeps the entire trace in memory through full attention. This means that the hardest problems, the ones that need the most thinking, are also the ones that cost the most to run. The authors measured what the middle of a reasoning trace is actually worth. Intermediate tokens steadily lose importance as the model keeps going. Their new approach, Prefix Sliding, drops those tokens. It keeps the prefix, which holds the instructions and the available tools, plus a window of the last few thousand tokens. Everything in between gets discarded during generation. Total memory stays capped no matter how long the model reasons. Without any training, this runs existing models 3x faster while matching full-attention performance, and it enables RL rollouts past 100,000 tokens. Paper: https://t.co/HzwSZ7fCdh Chat with Paper: https://t.co/OfjtVjamIC

❀️91
likes
πŸ”20
retweets
πŸ”Teknium πŸͺ½ retweeted
W
witcheer
@witcheer
πŸ“…
Aug 19, 2026
13d ago
πŸ†”34643678
⭐0.34

a new open-weight model family where the model card ships with Hermes Agent setup out of the box. point your Hermes Agent at a local Ornith 1.5 server, 9B to 397B! benchmark coming soon :)

❀️83
likes
πŸ”5
retweets
M
Niklas Muennighoff
@Muennighoff
πŸ“…
Aug 27, 2026
5d ago
πŸ†”85692974

new paper: Prefix Sliding for efficient test-time scaling vanilla full attention OOMs on long tasks &amp; compaction loses important details -- prefix sliding is a simple &amp; fast alternative that can outperform both πŸ“œhttps://t.co/fUw7yJAN5D https://t.co/PgcyrsfIZb

Media 1
❀️82
likes
πŸ”8
retweets
πŸ–ΌοΈ Media
A
Alexandre Morgand
@Almorgand
πŸ“…
Aug 20, 2026
11d ago
πŸ†”99432196

"ZipSplat: Fewer Gaussians, Better Splats" TL;DR: feed-forward 3DGS model that decouples Gaussian placement from pixels, reconstructing unposed scenes in under a second with ~6Γ— fewer Gaussians while achieving state-of-the-art quality. https://t.co/T6YmrQqlHy

❀️71
likes
πŸ”2
retweets
πŸ–ΌοΈ Media
πŸ”Nader Khalil🍊 retweeted
S
sunny madra
@sundeep
πŸ“…
Aug 30, 2026
2d ago
πŸ†”48453189
⭐0.34

GPT-5.4 xhigh scored 53 on Artificial Analysis in March. By August, Qwen3.8-Flash-Next scores 56 and GLM-5.3-Flash 57 with only 6B / 18B active params per token. Yesterday’s frontier is today’s Flash tier.

❀️68
likes
πŸ”6
retweets
S
sway
@SwayStar123
πŸ“…
Aug 31, 2026
1d ago
πŸ†”61220983
⭐0.42

I'm implementing a tiny transformer on tinyshakespear. The 2d matrices went to muon and 1d to adam. The model was still learning and generating some real words. But turns out my muon implementation was bugged and was no-oping. So 99.2% of my model was frozen at init. Adam still managed to tweak those biases into having the model still output some real words. Oh and I forgot the positional embeddings too. It's kinda crazy how you can have the shittiest implementation and a neural network still manages to learn

@iScienceLuvr β€’ Sat Aug 29 08:53

Karpathy's recipe for training neural networks is still relevant today. this lesson in particular is one we've been feeling very viscerally recently... neural network training can sometimes be very resilient and you may not realize there's an error for a very long time... https

❀️63
likes
πŸ”4
retweets
πŸ”Hamel Husain retweeted
H
Mike Taylor
@hammer_mt
πŸ“…
Aug 28, 2026
4d ago
πŸ†”49516910

I guess this is as good a way as any to announce that I'm now the Head of Evals at @every https://t.co/gPWYBoNrBY

Media 1Media 2
❀️59
likes
πŸ”1
retweets
πŸ–ΌοΈ Media