Your curated collection of saved posts and media

Showing 10 posts Β· last 14 days Β· by score
βž• Add New Post
πŸ”Ivan Leo retweeted
K
Ker Lee Yap
@klyap_
πŸ“…
Aug 30, 2026
2d ago
πŸ†”22782040
⭐0.36

Weekend project -- inference engineering step-by-step with GPT! It made a really fun beginner tutorial with exercises covering: - how to run the simplest working LLM server - intuition for prefill and decode - performance benchmarking on my system Next lesson plan: writing my own generation loop and learning about KV caches!

❀️37
likes
πŸ”1
retweets
πŸ”LlamaIndex πŸ¦™ retweeted
P
Perplexity
@perplexity_ai
πŸ“…
Aug 25, 2026
7d ago
πŸ†”60149371
⭐0.34

Parsing runs entirely on-device, so sensitive documents never leave the machine. On ParseBench-100, Computer scores 65.1% vs 34.6% for Hermes and 13.9% for Pi, in least time with fewest tokens. https://t.co/GQ590BYEn5

❀️35
likes
πŸ”1
retweets
D
DAIR.AI
@dair_ai
πŸ“…
Aug 22, 2026
10d ago
πŸ†”22808532

Banger paper from Microsoft. It's on agent reliability in real business workflows. (bookmark it) Thinkingbox is a sandbox with isolated MCP-compatible tool sessions, plus a benchmark of 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank IT, and consulting support. Every attempt is graded on the backend state the agent leaves behind. Executable checks accept valid trajectories and reject wrong, missing, or extra effects, so collateral damage counts against you. The strongest model reaches 65.36% pass@1 and 25.25% pass^20. Many failed trials terminate cleanly with valid state-changing tool calls. Watching the response or the tool call tells you very little about whether the task actually completed. Paper: https://t.co/EkMhBafbvI Track more trending AI papers in our academy: https://t.co/LRnpZN7L4c

Media 1Media 2
❀️28
likes
πŸ”4
retweets
πŸ–ΌοΈ Media
J
Jim Bohnslav
@jbohnslav
πŸ“…
Aug 18, 2026
14d ago
πŸ†”87573234
⭐0.38

if you want to know what it's really like building robotics foundation models, read this post. whether training WAMs or VLAs, your main job is shoveling good data into GPUs as fast as possible.

@DynaRobotics β€’ Mon Aug 17 17:02

When people talk about robotics, they usually talk about models, data, or hardware. Few people talk about the infrastructure that lets you iterate on all three quickly. Today we're publishing how we trained Dyna-2 on over 1,000,000 hours of egocentric video, repeatably. At this s

❀️27
likes
πŸ”4
retweets
S
Seldon
@seldon_tech
πŸ“…
Aug 21, 2026
11d ago
πŸ†”63298785

How good are agents actually at CAD? Today we release CADBench, evaluating 10 Models on 105 realistic tasks in Fusion 360 Fable 5 and Gemini 3.7 Flash are #1 with only 24.6% pass rate https://t.co/kMQ8U6uzKD

❀️24
likes
πŸ”4
retweets
πŸ–ΌοΈ Media
E
Ethan Mollick
@emollick
πŸ“…
Aug 27, 2026
5d ago
πŸ†”85378666
⭐0.40

Neat paper suggesting that human augmentation and task automation are not necessarily related. Models that are really good at doing work are not always good at helping humans do work better. Given the pressure to make models good agents, this may undermine human-AI cowork.

@abhishekn β€’ Wed Aug 26 23:37

the results? the best models in automate mode are not always the best models in augment mode. for example - Opus and Sonnet are great on performing tasks on their own, but not in proving assistance, while GPT-5-Mini was great on both dimensions. Gemini models were stronger as

❀️24
likes
πŸ”2
retweets
D
Datapoint AI
@datapointai
πŸ“…
Aug 20, 2026
12d ago
πŸ†”62065534

today, we’re releasing the largest open-source human image preferences dataset, along with a $1 million data grant - 2M+ annotations by real people - 30 SOTA image models ranked - 10 categories (marketing, product design, anime etc) dataset + benchmark + grant details below: https://t.co/K3U4uj1934

❀️22
likes
πŸ”9
retweets
πŸ–ΌοΈ Media
C
CoreWeave
@CoreWeave
πŸ“…
Aug 24, 2026
8d ago
πŸ†”03896156

27.8B parameters shouldn't be doing this. Qwen3.8-27B from @Alibaba_Qwen scores 52 on the @ArtificialAnlys Intelligence Index. Live on CoreWeave Serverless Inference with vision, tools + reasoning. Frontier-class model. CoreWeave-class infrastructure. https://t.co/2fNwmqYjoE https://t.co/0veUFgxwS0

Media 2
❀️19
likes
πŸ”3
retweets
πŸ–ΌοΈ Media
H
DailyPapers
@HuggingPapers
πŸ“…
Aug 21, 2026
11d ago
πŸ†”39457342

SWE-bench Science A new benchmark for scientific software engineering: 119 tasks across 98 repositories and 20 domains. Even the best agent, Claude Code with Opus-5, achieves under 50% pass@1. https://t.co/LtdNeJdwEh

Media 1
❀️19
likes
πŸ”2
retweets
πŸ–ΌοΈ Media
N
Nathan
@nathanhabib1011
πŸ“…
Aug 19, 2026
13d ago
πŸ†”40453027

NEW BENCHMARK ON THE @huggingface HUB Your favorite LLM agent can write a script. Can it survive 300+ steps in a terminal without losing the plot? That's what Long-Horizon Terminal-Bench (LHTB) measures, 46 tasks, contamination-resistant, hidden verifiers that check real state, not vibes. πŸ† Leaderboard πŸ₯‡ @MiniMax_AI 's Minimax M3 πŸ₯ˆ @Kimi_Moonshot 's Kimi k2.7 Code πŸ₯‰ @Zai_org's GLM 5.2 Congrats to @zli12321 and team. Super excited to see how smaller local models perf on this benchmark @0xSero :)

Media 1
❀️19
likes
πŸ”6
retweets
πŸ–ΌοΈ Media