Your curated collection of saved posts and media
I evaluated Ox Alpha on the induction benchmark, accessing the model through OpenRouter. The performance of Ox Alpha on this benchmark is OK but not great. It ranks below Luna and DeepSeek v4 Pro and just above Gemini 3.7 Flash. I have seen other benchmarks posted on X where Ox Alpha rocks, so this could be very benchmark-specific. It was not easy to get evaluable results (at least on OpenRouter). Ox Alpha often returned "" in responses, or API errors. I made 551 API calls to be able to get a reasonable completion rate of 87/100. I wonder if this is an OpenRouter issue or something broader like the model failing to return answers when it hasn't solved the problem. This is very common with reasoning models, but more pronounced than average here. Some words on the induction benchmark: This is a challenging reasoning benchmark, described in ICML 2026. The models are given several small graphs in which some nodes are marked as targets. The task is to provide a first-order logical formula that picks precisely the target nodes in all graphs simultaneously. Correct: a formula that correctly picks precisely the marked nodes. Holdout correct: a formula that correctly picks precisely the marked nodes in held out problems. Formula complexity (in AST): the avg tree size of the correct formula. GitHub public repository: https://t.co/ZXXqzouuni Paper: https://t.co/gBelIZQEaa
@0xEronn Unfortunately thereโs no waves in transformers but a single vector (residual stream) and linear transformations (inference, single forward pass). The only place there are actual waves is in positional encoding.
weird that there's a "make a lot of money" button and nobody's pressing it (take your SaaS, make it headless, let agents use it, charge per interaction esp for enterprises)
Claude has become a language, so I built a translator. English <-> Claudish https://t.co/871CI2XdUz
Iโm thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc. Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.
Cursor has been a trusted partner of Anthropic since Sonnet 3.5. Weโll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! ๐ ๐น This experimental multimodal model matches DeepSeek-V4-Flash on text capabilitiesโincluding agents, reasoning, and world knowledge. ๐น On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
Introducing TimesFM-3, a state-of-the-art time series foundation model that enables accurate multivariate time series forecasting in a single forward pass, significantly outperforming other forecasting models across major benchmarks. More on the blog โhttps://t.co/uSlnIdUJ4Q https://t.co/NfbONpFYDz

๐ I also just uploaded a new version of the lecture notes. https://t.co/bpIvdgG8KO (it's still, and perhaps will always be work-in-progress, given the field is moving fast. Any mistakes are mine :) and any feedback is more than welcome :).
Instead of watching 2 hours of Netflix tonight, watch this Stanford lecture. Itโs the clearest end-to-end explanation Iโve seen of how ChatGPT and Claude are actually built. From tokenization and BPE all the way to the transformer architecture, training pipeline, and next-token
Please welcome to the world a beautiful new geometric object, to do with a problem iโve always loved. claude really contains multitudes:D Does S^6 admit a complex structure? Yup
mojo is now open source https://t.co/oHZf5o6LgR https://t.co/AW0hjxcJPM
many people asked me how to write CLAUDE.md or AGENTS.md, and i see lots of bad advice flying around so i took some time to write down a guide in https://t.co/v9rrkWKEFr tl;dr - handwrite your user level AGENTS.md - for project level ones, you don't write it. you train it like a neural net i also open sourced my private solution "backpass" at https://t.co/DkM9b4TcZ0 - it samples your past agent sessions for a repo, distill key learnings and losses, synthesize them, and produce a gradient descent step as a proposal that you can review and apply to improve your AGENTS.md and project level skills easiest way to run it is just "npx -y backpass" in your repo hope it helps! please share with whoever you think can benefit from it
Today we're releasing abliterated-model-large-v2. Based on GLM-5.3, which is #3 on Terminal-Bench 4.0 (behind only Opus 5 and Fable), with 2ร the cyber exploitation of 5.2. We abliterated and hosted it so it does the offensive cyber, red teaming, and agent testing work other models refuse to do. - US-hosted - FP8 - 1 million context window - Zero input/output prompt retention Live now. ๐งต
Might be genuinely worth explaining to the masses the difference between LLMs / diffusion models that everyone uses for ChatGPT essays and image slop and deep learning algorithms used to detect breast cancer to preemptively combat stupid takes like these https://t.co/REJd2dBUcW
Neil Movva (@neilmovva) started his career at Nvidia, working on GPUs and kernels, and has an unusually deep understanding of inference, from software to chips to power. We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are. What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow. Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible. We discuss: - Latency versus throughput - Why there are no bad chips, only bad pricing - The end of kernel engineering - Buying chips and power no one else wants - New chip architectures - Nvidia lore + his contrarian view of the company - Open source and the frontier labs I learned a ton. Enjoy! TIMESTAMPS 0:00 Intro 0:38 Building a โToken Factoryโ 4:21 The Future of Background Agents 13:09 Nvidia and the GPU Stack 23:27 Chips, Memory, and Transformers 36:14 The Future of AI Training Data 44:32 Chip Scarcity and Compute Arbitrage 52:44 Reinventing the AI Data Center 59:01 Power and the โScavenger Strategyโ 1:10:10 Open vs. Closed AI
MacใใญใผใซใซLLMใซ่ๅณใใใไบบใปใฉใHermes Agentใฏๅๅผทใใใปใใใใใ ChatGPTใClaudeใฏใใขใใซๅไฝใใใชใใ ใกใขใชใใใผใซใในใญใซใชใฉใฎใใใผใในใ่พผใฟ่พผใฟใฎใตใผใในใ ใญใผใซใซLLMใฏ้ใใ LlamaใQwenใ่ฝใจใใฆใใใใใซใใใฎใฏใ้ ญ่ณใใ ใใ ใใใซHermes Agentใ้ใญใใ LLMใ่ใใใHermesใๅใใ่ชฟในใใ่ฆใใใใใผใซใไฝฟใใใใกใคใซใ่งฆใใใณใใณใใๅฎ่กใใใ ใ่ใใ้ ญ่ณใใซใๅใใ่บซไฝใใไธใใใ ใใใAIใจใผใธใงใณใใ ใ ใใๆฏ่ผใใใชใใQwen vs Fableใใ ใใใใชใใ ใQwen + HermesใๅฏพใFable + Claudeใขใใชใใฎใใใซใใใผใใน่พผใฟใง่ใใใ ใใฃใใใ่ฆ็ดใ ใใชใLM Studioใงๅๅใ ใงใใใขใใซใใๅใใจใผใธใงใณใใใซใใใใชใใHermes Agentใฏๆฌฒใใใชใใ ใขใใซใฏใใใๅทฎใๆฟใใใ ๆฎใใฎใฏใ่ชๅใงๆใคใใผใในใ Hermes Agentใฏใ็กๆใงไฝฟใใใจใผใธใงใณใใใผใในใจใใฆใใชใๅผทใใจๆใใ
Introducing Unreal MCP in UEFN. With Unreal MCP, you can connect agentic coding tools like Claude Code, Codex, or Cursor directly to UEFN, opening up new ways to build your experiences in the editor, from writing Verse and configuring devices to working with Scene Graph. Learn more: https://t.co/43PK3uGLiV Hereโs a look at what you can do ๐งต๐
I have no idea what they feed @vmg but this is an unbelievable read. One of the best systems posts Iโve read. Ever. With coding being automated this is the sort of architectural clarity we need.
We're making Git hosting more reliable, performant, and scalable. This post traces 20 years of Git infrastructure and explains how that history led us to design and operate our Git storage, Origin, as if it were a database. https://t.co/UW7jHuItSX
On showing an early demo of Portable Computer on DGX Spark to Jensen, he was kind to gift us a DGX Station, a beast of a local computer that can serve even frontier models like GLM 5.3. Unmetered frontier intelligence running on your own local hardware coming soon! https://t.co/HA16eiABnq
AI did not suddenly discover a cancer vaccine. That is engagement bait. Moderna and Merck reported positive Phase 3 results for Intismeran plus Keytruda in 1,137 melanoma patients after surgery. This is a real and important scientific result. But the vaccine entered human trials in 2017, and Moderna was already using internal bioinformatics algorithms to select each patientโs tumor targets. AI helps analyze mutations and choose up to 34 neoantigens for the personalized treatment. It did not independently invent the vaccine or cure cancer. The full Phase 3 numbers have not been released. Overall survival is still unknown. The treatment is not approved, not a universal cancer vaccine and was tested with Keytruda, not alone. BioNTech, Roche, NEC, Transgene and Evaxion have used similar computational or machine learning approaches for years. This is a breakthrough in genomics, immunology, mRNA, manufacturing and clinical science. Rebranding all of that as โAI discovered a cancer vaccine todayโ is pure AI hype.
AI did not suddenly discover a cancer vaccine. That is engagement bait. Moderna and Merck reported positive Phase 3 results for Intismeran plus Keytruda in 1,137 melanoma patients after surgery. This is a real and important scientific result. But the vaccine entered human trials in 2017, and Moderna was already using internal bioinformatics algorithms to select each patientโs tumor targets. AI helps analyze mutations and choose up to 34 neoantigens for the personalized treatment. It did not independently invent the vaccine or cure cancer. The full Phase 3 numbers have not been released. Overall survival is still unknown. The treatment is not approved, not a universal cancer vaccine and was tested with Keytruda, not alone. BioNTech, Roche, NEC, Transgene and Evaxion have used similar computational or machine learning approaches for years. This is a breakthrough in genomics, immunology, mRNA, manufacturing and clinical science. Rebranding all of that as โAI discovered a cancer vaccine todayโ is pure AI hype.
With cancer vaccines now being discovered with AI ($MRNA), it seems that the US government might actually grow its way out of its budget deficit. It feels like we're in the early innings of the healthcare system becoming unburdened by many terminal illnesses.
Continuous self-improvement needs an ever-expanding supply of training environments (goals). SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. โ ๏ธ https://t.co/5QRs5LksiW
Train in sim. Run it for real. Teach it new tricks. ๐ฆ This is the sim2real loop that powers Microduck. RL stack open source here : https://t.co/mveM7y5fYa Buy it here : https://t.co/Rmk8F4vpKd Git : https://t.co/kcoCKdBi52 Join our community : https://t.co/TcX5uaf0yL https://t.co/OyOpzl9hnh
Wow this is huge!!! ๐ฏ Hermes agent can now browse as you, with your existing logins and cookies. I see so many use cases!
Hermes Agent can now seamlessly browse as you. Turn on real-profile browsing and your agent acts with your logins, from a managed copy of your existing Chrome profile. https://t.co/20IKsY6LZK
Tip for your @NousResearch Hermes Agent Bots: Create an "Orchestrator" Bot and then have it design its own team to optimally fulfill any/all user requests. Here's El Jefe (my orchestrator) completing a very solid roster including comprehensive soul.md entries for each Bot! https://t.co/4VKL5UxG25
El Jefe and Fixer working together to assemble the optimal Hermes Bot team is 10/10 entertainment ๐คฉ https://t.co/5VMTjklvT8
Vercel Connect is now generally available. Give your apps and agents secure access to @slackhq, @linear, @github & 100+ other services. โข Short-lived, scoped access tokens โข Token and trigger observability โข RBAC and audit trails https://t.co/3JUlBz3ddo
Stop using your agent logs just for debugging. Use them to train your own model. Today we're launching Agnost AI (YC S26)'s first model: agnost-*******-0.1 Trained for our first customer on their existing production traces & it beat the frontier: +22.9% task success โ90.2% latency โ94.5% cost If you ever thought of fine-tuning your own model, we'll do it for you, talk to us!
Damn. Hermes Agent + local Qwen3.8-27B is a beast. So far it hasn't failed a single coding task I've given it, and I keep upping the complexity to see if I can make it fail. So far, it has been near-perfect. As good as Opus.
OpenWorker -- an open source agent that doesn't just chat but completes tasks on your laptop -- just released a new version with many features for security workflows. After our initial release, many users found it especially useful for cybersecurity. Attackers are already using AI; OpenWorker is committed to giving defenders the same leverage. Running an agent requires both (i) A model and (ii) A harness (the software around the model). Because the OpenWorker harness is fully open source, security teams can audit it to make sure we haven't built any backdoors that exfiltrate your code and data to some company or even a foreign adversary. OpenWorker now comes with built-in cybersecurity agents for (i) Scanning your code for vulnerabilities. (ii) Scanning dependencies for supply chain injections. (iii) Checking your cloud security configuration for attack surfaces. This enables developers to do much more security work before deployment (part of what's called the "shift left" movement). You choose the model: you can run open weight models fully locally so sensitive code never leaves your machine. This helps with legitimate security work (like reproducing a known exploit to defend against it) that can trigger refusals in leading closed models. Or use your ChatGPT subscription, or stealth preview models like Ox Alpha, or any model via API key. Thanks also to all the open source contributors! Join work with @rohitcprasad so please follow him too to get more frequent updates. Try it out: https://t.co/QPZLudn7ug Code: https://t.co/NYCiTD6hSq
SIGReg for pretraining Video Foundation Models! Our LeVJEPA opens many doors... - stable recipe with a simple loss (sigreg + prediction) - no tubelet, frame aggregation, EMA, stop-gradient, .... - 20X more FLOP efficient than VJEPA1/2 pretraining - open source + reproducible https://t.co/rCwJRx2Dim
SIGReg for pretraining Video Foundation Models! Our LeVJEPA opens many doors... - stable recipe with a simple loss (sigreg + prediction) - no tubelet, frame aggregation, EMA, stop-gradient, .... - 20X more FLOP efficient than VJEPA1/2 pretraining - open source + reproducible https://t.co/rCwJRx2Dim
A new Pareto frontier in video pretraining. Excited to introduce LeVJEPA ๐ฅ: a stable, efficient end-to-end pretraining method that matches V-JEPA 2 at up to 20x less pretraining compute! No target encoder, no masked prediction, no stop-gradient or teacher-student schedule. On
Perplexity Search debuts on the Artificial Analysis Search Index, with all three context size variants taking top positions on the leaderboard The @perplexity_ai Search API comes with three context settings (low, medium, and high) that control how much extracted content each search result carries. We tested all three variants using our standardized methodology: the same model (GPT-5.6 Luna at medium reasoning), running inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web. Only the provider behind the search tool changes. Key results: โค Perplexity Search (medium) scores 80 on the Artificial Analysis Search Index, ahead of the previous leaders, Parallel (advanced) and Brave Search (LLM context), at 75. The high and low variants score 79 and 77 respectively. Its lead is concentrated in BrowseComp results, with AA-Omniscience and DeepSearchQA scoring comparably to other leading providers โค Efficient search payloads: smaller overall search results mean the model reads less per task, so Perplexity has the lowest model inference cost per task of providers weโve tested so far, ranging from $0.028 to $0.034 across the three variants vs $0.036 for the next lowest provider โค Total cost per task is ~$0.091 for the medium and high context variants, at mid-pack latency. For comparison, Parallel (advanced) costs $0.084 per task and Brave (LLM context) costs $0.13 per task