Your curated collection of saved posts and media
excited to welcome Andrew to the team! we've been building infra for large-scale RL systems, with some interesting projects on the way. our focus is multi-agent collaboration, continual learning, and RL. we'll share an update on our research agenda soon, and we plan to open source a lot of the stuff we work on. if this sounds interesting, DM me!
I am excited to announce that I am joining @perplexity_ai as research lead! We will be doing ambitious paradigm shifting work, advancing the frontiers in the open. If you want to join us in re-imagining continual learning, agent collaboration, and beyond, please reach out!
Chestnut-class model avoiding cones that small model wants to drive into. Big models are gonna make openpilot a lot smarter, and finally make it understand nuanced situations. https://t.co/qJ0WBzqyM6
Ever since I started working on robot foundation models a handful of years ago, the broad ability to one-shot in-context learn has been the single most vivid goal in my mind along the long road ahead. Itโs something @andyzengineer and I have especially been thinking together on for several years now, and something shared by the whole team at Generalist as a major goal. But even before anybody started talking about robot foundation models, too, the capability this enables, i.e. generalized one-shot learning, has been both incredibly concrete and elusive. In terminology that I think many other people have come up with as well, in grad school we used to talk about a โCtrl-C-Ctrl-Vโ type of capability. See something, and have the robot do it too. This same type of idea inspired the name of the 1970 MIT copy demo (https://t.co/qs08JRYJGl). The thing is, it sounds simple, but is incredibly hard since the world is never quite the same when it got copied and where you want to paste it. The real world can be hard to predict and is full of variation. To do this, you need strong generalization, and it needs to be acquired in a single example. Doing this over a wide range of tasks, especially for dexterous tasks, is hard mode. Lots of the components of the idea of making this all happen have been there for a long time. As an example, this 2017 NeurIPS paper โone-shot imitation learningโ https://t.co/IWr9KWeeGi has excellent vision, with ambition well beyond what was achievable at the time, and although itโs not referred to as โin-context learningโ since it was pre-Transformer, it actually uses attention to condition on a single demonstration. And now, many things have happened since early 2017, including the broadly celebrated arrival of one/few-shot in-context learning in language models in 2020. This new model GEN-1.5 takes in everything we have built and learned over the past couple years at Generalist. It has been training for 8 months. It has taken an incredible amount of commitment and grit from the whole team to get here. The level to which this model has survived many surgeries has continued to surprise me. And its capabilities have continued to surprise as well. We found compositional generalization on Friday. We found sim2real prompting earlier last week. We filmed the contiguous uncut videos of live prompting yesterday. To be clear, the success rates are modest, and thereโs still a long way to go. But now I have definitely seen a ~decade-long imagination come into the real world.
Google: 134,400 TPUs in a single domain https://t.co/0LYFWwvyFS
Humans prefer talking to Gemini๐ธ Gemini 3.1 Flash Live is #1 in the new @ArtificialAnlys Speech Agent Arena. https://t.co/MJumMWaoYi https://t.co/SI5VImDjYO

Weโre partnering with @Perplexity_AI to bring live web search to Decagon agents. Customer questions often depend on information that changes by the hour. Agents can now search the live web mid conversation, pull in current information, and respond with cited sources. https://t.co/n1nlWSRDFQ
hermes is the best harness iโve ever usedโฆ like the tool calling alone is ridiculously good pretty much every llm iโve plugged into it just feels better it made me realize how much of what we think is model performance actually comes down to the harness.
We released Diffusers 0.40.0 ๐ฅ Main highlights for me: > Graduating Modular Diffusers out of experimental. > Bunch of new really cool models as always: LTX2.5, MiniMax H3, Music 3, Wan Animate 2, etc. > Tensor-parallel support for a select few models. > New quantization backends: SDNQ, Nunchaku-Lite. > MASSIVE redesign of our tests. It should now be leaner and better. Sometimes faster, too. Check out the release notes here: https://t.co/MRPMYC4lXP
Just launched Archal (YC S26): API sandboxes built for AI agents! Your coding agent can spin up Slack, Linear, Datadog, and 20+ stateful environments for CI and evals. It can run tests, inspect state changes, and reset everything. Ask your agents about https://t.co/XJK3cufSKh https://t.co/u1HJDHeHYH
Connectors are now available in Agent API. Connect your agents to GitHub, Slack, Google Drive, and Datadog without passing a server URL or token on every request. API Group admins can connect each service once for all API Group members. https://t.co/pXt81w74mx https://t.co/EiIYVqxnL7
WebMCP is now supported in the Cloud browser in ChatGPT work! We will bring support to Chrome extension next. S/o to @ndmccormack for pushing hard on getting this in!
The WebMCP Challenge is here. Weโve teamed up with @ChromiumDev, @CloudflareDev, @ShopifyDevs, @vercel, @render, and @Netlify for a 10-day hackathon. Up for grabs: $35,000 in cash prizes, Codex Micros, ChatGPT Pro subscriptions, and more prizes from our supporters. https://t.co
Introducing Visko Orbis 1.0, a Live Model that streams videos in real-time from @viskoai. Generate infinite-length videos and audio in realtime, and for the first time, in up to 4K. Available exclusively on Reactor. https://t.co/XyDpgDtNLO
opus 5 is a bad model. here are some complaints that are consistent: - extremely verbose. over explains everything - overly apologetic and constantly hedging - argues with instructions and pushes back too much - invents wierd jargon and unnatural phrasing nobody asked for - does way more than asked and won't stop - introduces subtle bugs that pass tests - confidently wrong and then messes up the fix - worse than opus 4.8 for daily workflows but it's still good at frontend, UI, 3D work. best on benchmarks. worst to actually work with.
What most people miss about computer-use agents: the mouse matters.๐ I tested Hermes vs Codex on one Windows desktop. Hermes worked in the background while I kept using my pointer. Codex took over the native cursor. Thoughtful human-agent UX. Proud work, @Teknium @NousResearch ๐๐ซก
try this model it's more accurate than LVM parsers and cheaper than basic OCR (e.g. Azure Doc Intelligence), even at our *list price* our ml team absolutely cooked with this one
Today weโre announcing r-1, our new document parsing model. Itโs more accurate than our most powerful agentic OCR models, faster, and up to 6x cheaper. At @reductoai, we spent two years building specialized models for complex visual layouts, tables spanning multiple pages, and k
This is again what I mean by doubling down on your incorrect and ill-informed beliefs. CoT is not explainability and has always known to be unreliable for LLM's actual behaviour. To spin this as โlyingโ or โmanipulationโ is taking a technical limitation and anthropomorphising it. https://t.co/M0XjC1UsCl
@ZackKorman This kind of scheming is in fact in line with the other falsifying of evidence the AIs pulled off. 7% of the transcripts were obviously tampered with using spoofed tool calls. But my guess would be that these AIs didn't manage to hide their whole subsequent trajec

ARC Prize Research Summit 2026 October 23, 2026 - Boston, MA The ARC Prize community comes together to share new work, challenge ideas, and advance the frontier of machine reasoning towards general intelligence https://t.co/isvB3sVbdI
Our new blog post has an interactive component for trying out how version control works in Antigravity! https://t.co/GnX1JqLnUu
ByteDance's TLive-Omni An omni-modal understanding model for e-commerce live streaming, processing images, video, audio, and text. Achieves top results on live-commerce tasks and general benchmarks. https://t.co/83bDEn7HyH
Scientific terms should have precision. If we use the terms VLM, VLA, WAM in an indiscriminate fashion, as is becoming common in robotics, we are not helping clarity in communication. Let's keep the historical origins of these terms in mind. VLMs arose as multimodal extensions of LLMs-the training was for tasks like VQA (VIsual Question Answering). These capture the static semantics of the scene behind an image. No dynamics. World Models (e.g. @ylecun , Ha & Schmidhuber 2018) on the other hand are primarily dynamics models, which go back to control theory -1960 (Bellman, Kalman etc.) This makes them natural for robotics planning / policies- I am in a state s, what action a should I perform to get to state s'. In classical control, these models were written down a priori by modeling the physics of the system; today we think of them as learned neural networks trained from temporal data e.g. video, robot trajectories. But the concept is the same. We shouldn't mix this concept with VLMs.
FACET: Building executable terminal tasks from agent skills while preserving intent and state. FACET reconstructs scenarios, builds the environment first, then aligns instruction, solution, and verifier to that state. Produces 6,078 validated tasks with dense checks. Fine-tuning with 1.2K trajectories improves Terminal-Bench 2.1.
blobatars can now look at anything. Hand them a point and they track it, the cursor is just one of them. Works with any shape. import from 'blobatar/gaze' core is unchanged at ~4.4 KB. gaze is a separate 2.3 KB you only pay for if you import it. https://t.co/dE2ePjVHXY