Your curated collection of saved posts and media
Our paper is top on HuggingFace Daily now: https://t.co/voDYYyL13m Also check out our Github Repo on Agent Team Harness: https://t.co/9Ujn8AoBHo
Apodex 1.1 Scaling Agentic Intelligence for Complex Work paper: https://t.co/bZUS55viAw https://t.co/UWeTBvUj4S

One dead giveaway is also that Zhipu is one of the only labs that serves models with pretty poor partial prefill tok/s while output tok/s are fast. This model is ~30% faster than GLM 5.3. If the infras is the same, likely ~300-500B or fewer params vs activated params.
A new Kimi model, likely K3.1, is now being tested on the Code @arena under the name "korrine" K3 was tested on the Arena as "kivine" prior to its launch If anyone's wondering, "Ox Alpha" on OpenRouter is the upcoming GLM 5.3 Flash from fellow Chinese lab Zhipu
I evaluated Ox Alpha on the induction benchmark, accessing the model through OpenRouter. The performance of Ox Alpha on this benchmark is OK but not great. It ranks below Luna and DeepSeek v4 Pro and just above Gemini 3.7 Flash. I have seen other benchmarks posted on X where Ox Alpha rocks, so this could be very benchmark-specific. It was not easy to get evaluable results (at least on OpenRouter). Ox Alpha often returned "" in responses, or API errors. I made 551 API calls to be able to get a reasonable completion rate of 87/100. I wonder if this is an OpenRouter issue or something broader like the model failing to return answers when it hasn't solved the problem. This is very common with reasoning models, but more pronounced than average here. Some words on the induction benchmark: This is a challenging reasoning benchmark, described in ICML 2026. The models are given several small graphs in which some nodes are marked as targets. The task is to provide a first-order logical formula that picks precisely the target nodes in all graphs simultaneously. Correct: a formula that correctly picks precisely the marked nodes. Holdout correct: a formula that correctly picks precisely the marked nodes in held out problems. Formula complexity (in AST): the avg tree size of the correct formula. GitHub public repository: https://t.co/ZXXqzouuni Paper: https://t.co/gBelIZQEaa
@0xEronn Unfortunately thereβs no waves in transformers but a single vector (residual stream) and linear transformations (inference, single forward pass). The only place there are actual waves is in positional encoding.
weird that there's a "make a lot of money" button and nobody's pressing it (take your SaaS, make it headless, let agents use it, charge per interaction esp for enterprises)
The wait is over! Meet Arduino VENTUNO Q, where AI takes action. π¬ Run local LLMs like Qwen 3, Gemma 4, and Qwen 3 VLM directly on the board π§ NPU + CPU + GPU +MCU: @Qualcomm Dragonwing IQ-8275 with up to 40 dense TOPS of AI performance and STM32H5 microcontroller for real-time control ποΈ 16 GB RAM + 64 GB eMMC + expandable storage πͺ Linux-powered, pre-loaded with Ubuntu OS + Zephyr RTOS π οΈ Build AI faster using Arduino App Lab, @huggingface, @EdgeImpulse, and Qualcomm AI Hub π¨ Move seamlessly from prototype to production through Works with Arduino Certification Program The first batch wonβt last. Pre-order your VENTUNO Q with a free power supply and USB-C cable included! Get ready to enter a new era of Physical AI: https://t.co/YpDhKRuX6y
Iβm thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc. Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.
Cursor has been a trusted partner of Anthropic since Sonnet 3.5. Weβll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.
Fable 5.1 is now live in Claude Code and the Claude Platform. It's priced the same as Fable 5, with 75% cheaper API cache reads. It gets a lot further into a long task before it needs your input, is better at telling you when it's stuck, and its writing style is more natural.
Weβre introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the worldβs most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! π πΉ This experimental multimodal model matches DeepSeek-V4-Flash on text capabilitiesβincluding agents, reasoning, and world knowledge. πΉ On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n