Your curated collection of saved posts and media

Showing 10 posts Β· last 14 days Β· by score
βž• Add New Post
S
Simon Shaolei Du
@SimonShaoleiDu
πŸ“…
Aug 25, 2026
7d ago
πŸ†”01443194

Our paper is top on HuggingFace Daily now: https://t.co/voDYYyL13m Also check out our Github Repo on Agent Team Harness: https://t.co/9Ujn8AoBHo

@_akhaliq β€’ Tue Aug 25 15:13

Apodex 1.1 Scaling Agentic Intelligence for Complex Work paper: https://t.co/bZUS55viAw https://t.co/UWeTBvUj4S

Media 1Media 2
❀️15
likes
πŸ”3
retweets
πŸ–ΌοΈ Media
T
Tim Dettmers
@Tim_Dettmers
πŸ“…
Aug 21, 2026
11d ago
πŸ†”84608066
⭐0.36

One dead giveaway is also that Zhipu is one of the only labs that serves models with pretty poor partial prefill tok/s while output tok/s are fast. This model is ~30% faster than GLM 5.3. If the infras is the same, likely ~300-500B or fewer params vs activated params.

@synthwavedd β€’ Fri Aug 21 14:54

A new Kimi model, likely K3.1, is now being tested on the Code @arena under the name "korrine" K3 was tested on the Arena as "kivine" prior to its launch If anyone's wondering, "Ox Alpha" on OpenRouter is the upcoming GLM 5.3 Flash from fellow Chinese lab Zhipu

❀️15
likes
πŸ”1
retweets
S
Serafim Batzoglou
@s_batzoglou
πŸ“…
Aug 22, 2026
10d ago
πŸ†”97406978

I evaluated Ox Alpha on the induction benchmark, accessing the model through OpenRouter. The performance of Ox Alpha on this benchmark is OK but not great. It ranks below Luna and DeepSeek v4 Pro and just above Gemini 3.7 Flash. I have seen other benchmarks posted on X where Ox Alpha rocks, so this could be very benchmark-specific. It was not easy to get evaluable results (at least on OpenRouter). Ox Alpha often returned "" in responses, or API errors. I made 551 API calls to be able to get a reasonable completion rate of 87/100. I wonder if this is an OpenRouter issue or something broader like the model failing to return answers when it hasn't solved the problem. This is very common with reasoning models, but more pronounced than average here. Some words on the induction benchmark: This is a challenging reasoning benchmark, described in ICML 2026. The models are given several small graphs in which some nodes are marked as targets. The task is to provide a first-order logical formula that picks precisely the target nodes in all graphs simultaneously. Correct: a formula that correctly picks precisely the marked nodes. Holdout correct: a formula that correctly picks precisely the marked nodes in held out problems. Formula complexity (in AST): the avg tree size of the correct formula. GitHub public repository: https://t.co/ZXXqzouuni Paper: https://t.co/gBelIZQEaa

Media 1
❀️1
likes
πŸ–ΌοΈ Media
G
Gerard Sans | Axiom πŸ‡¬πŸ‡§
@gerardsans
πŸ“…
Aug 23, 2026
9d ago
πŸ†”53443901
⭐0.40

@0xEronn Unfortunately there’s no waves in transformers but a single vector (residual stream) and linear transformations (inference, single forward pass). The only place there are actual waves is in positional encoding.

T
Thariq
@trq212
πŸ“…
Aug 18, 2026
14d ago
πŸ†”91479333
⭐0.32

weird that there's a "make a lot of money" button and nobody's pressing it (take your SaaS, make it headless, let agents use it, charge per interaction esp for enterprises)

❀️6,553
likes
πŸ”231
retweets
A
Arduino
@arduino
πŸ“…
Aug 25, 2026
7d ago
πŸ†”37205982

The wait is over! Meet Arduino VENTUNO Q, where AI takes action. πŸ’¬ Run local LLMs like Qwen 3, Gemma 4, and Qwen 3 VLM directly on the board 🧠 NPU + CPU + GPU +MCU: @Qualcomm Dragonwing IQ-8275 with up to 40 dense TOPS of AI performance and STM32H5 microcontroller for real-time control πŸ—„οΈ 16 GB RAM + 64 GB eMMC + expandable storage πŸͺ Linux-powered, pre-loaded with Ubuntu OS + Zephyr RTOS πŸ› οΈ Build AI faster using Arduino App Lab, @huggingface, @EdgeImpulse, and Qualcomm AI Hub πŸ’¨ Move seamlessly from prototype to production through Works with Arduino Certification Program The first batch won’t last. Pre-order your VENTUNO Q with a free power supply and USB-C cable included! Get ready to enter a new era of Physical AI: https://t.co/YpDhKRuX6y

❀️5,377
likes
πŸ”693
retweets
πŸ–ΌοΈ Media
T
tobi lutke
@tobi
πŸ“…
Aug 25, 2026
7d ago
πŸ†”38495186
⭐0.38

I’m thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc. Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.

❀️5,100
likes
πŸ”221
retweets
N
Tom Brown
@NotTomBrown
πŸ“…
Aug 29, 2026
3d ago
πŸ†”27280657
⭐0.36

Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.

❀️4,764
likes
πŸ”227
retweets
C
ClaudeDevs
@ClaudeDevs
πŸ“…
Sep 01, 2026
7h ago
πŸ†”34277228
⭐0.36

Fable 5.1 is now live in Claude Code and the Claude Platform. It's priced the same as Fable 5, with 75% cheaper API cache reads. It gets a lot further into a long task before it needs your input, is better at telling you when it's stuck, and its writing style is more natural.

@claudeai β€’ Tue Sep 01 18:03

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3

❀️3,334
likes
πŸ”213
retweets
D
DeepSeek
@deepseek_ai
πŸ“…
Aug 21, 2026
11d ago
πŸ†”74631962

DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! πŸš€ πŸ”Ή This experimental multimodal model matches DeepSeek-V4-Flash on text capabilitiesβ€”including agents, reasoning, and world knowledge. πŸ”Ή On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n

Media 1
❀️2,219
likes
πŸ”273
retweets
πŸ–ΌοΈ Media