Your curated collection of saved posts and media

Showing 10 posts Β· last 14 days Β· by score
βž• Add New Post
πŸ”Aravind Srinivas retweeted
P
Perplexity
@perplexity_ai
πŸ“…
Aug 25, 2026
6d ago
πŸ†”21432824
⭐0.38

New research: Portable Computer is a local-first agent for private and cost-effective work. With an on-device 27B model, our harness scores 82.6% on real knowledge work, beating open-source harnesses Pi and Hermes. Our post-trained PPLX 27B reaches 85.4%. https://t.co/Rb6d47clCI

❀️62
likes
πŸ”10
retweets
A
Alec Fong
@alecqfong
πŸ“…
Aug 26, 2026
5d ago
πŸ†”59411574

GLM 5.3 Flash will likely be my teams new daily driver Untuned I'm getting on 2xDGX Station TP=2 881 tok/s C=64 232 tok/s C= 1 At long context realistically can have 4-8 simultaneous users. + it supports vision!

Media 1Media 2
❀️14
likes
πŸ”1
retweets
πŸ–ΌοΈ Media
I
Tanishq Mathew Abraham, Ph.D.
@iScienceLuvr
πŸ“…
Aug 27, 2026
5d ago
πŸ†”68022979

Prefix Sliding for efficient test-time scaling "we find most intermediate reasoning tokens lose importance as the model continues reasoning. This calls into question whether retaining them is worth the cost. Based on this insight, we propose Prefix Sliding, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens." "Without training, Prefix Sliding can make existing models 3x faster while maintaining performance. Training with Prefix Sliding using reinforcement learning can achieve better performance by enabling scaling to reasoning traces beyond a hundred thousand tokens." code: https://t.co/8lGwjKkavI link: https://t.co/Sd2wre3L9A

Media 1
❀️12
likes
πŸ”1
retweets
πŸ–ΌοΈ Media
G
Giles Thomas
@gpjt
πŸ“…
Aug 20, 2026
12d ago
πŸ†”50812688

By switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000! https://t.co/NmNQu53cmW

Media 1
❀️5
likes
πŸ–ΌοΈ Media
H
Surge AI
@HelloSurgeAI
πŸ“…
Aug 20, 2026
11d ago
πŸ†”18515459

Qwen 3.8 Max on the Surge Scorecard: Tuesday Work Index β€” 58.7 ComplexConstraints β€” 45.5% Chartography β€” 29.1% HANDBOOK.md β€” 16.5% Riemann-bench β€” 15.2% Hemingway-bench β€” 1006 Elo Antidote β€” 990 Elo https://t.co/3MzHkqo1m4

Media 1
❀️2
likes
πŸ–ΌοΈ Media
M
Michael Truell
@mntruell
πŸ“…
Aug 29, 2026
3d ago
πŸ†”06063557
⭐0.38

We’re sorry to see that OpenAI put out a note saying they plan to block Cursor users from accessing OpenAI models in three months. OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this. Cursor was one of the very first users of OpenAI, we’ve worked closely with their team for years, and we’ve trusted their platform to be neutral infrastructure for our business.

❀️18,074
likes
πŸ”760
retweets
A
Qwen
@Alibaba_Qwen
πŸ“…
Aug 26, 2026
6d ago
πŸ†”24515114

⚑Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: πŸ₯³ - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.πŸš€ We can't wait to see what you build with Qwen3.8-Flash!πŸ‘€πŸ‘‡ - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

Media 1Media 2
+3 more
❀️5,238
likes
πŸ”738
retweets
πŸ–ΌοΈ Media
D
Ben Davis
@davis7
πŸ“…
Aug 21, 2026
11d ago
πŸ†”31298095

I ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh... gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80% (was a near miss on the "x"s so actually over 80%) I am very confused https://t.co/NdDSTjoHLj

@ β€’

Media 1
❀️3,724
likes
πŸ”150
retweets
πŸ–ΌοΈ Media
πŸ”Weights & Biases retweeted
P
Percy Liang
@percyliang
πŸ“…
Aug 21, 2026
10d ago
πŸ†”34684997
⭐0.36

🚒 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.

❀️3,579
likes
πŸ”506
retweets
πŸ”Hugging Face retweeted
Z
Z.ai
@Zai_org
πŸ“…
Aug 28, 2026
4d ago
πŸ†”22455713
⭐0.36

GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize. Weights: https://t.co/v1IbWMXxg4 Tech blog: https://t.co/ekQkO83jCv https://t.co/f8XlJksKyf

❀️2,604
likes
πŸ”333
retweets