Your curated collection of saved posts and media
new paper: Prefix Sliding for efficient test-time scaling vanilla full attention OOMs on long tasks & compaction loses important details -- prefix sliding is a simple & fast alternative that can outperform both ๐https://t.co/fUw7yJAN5D https://t.co/PgcyrsfIZb
Atlas is an autoregressive diffusion model built from the ground up for the task of "next frame prediction". It is simultaneously a world class method for camera-controlled video generation, novel view synthesis, and sparse 3D reconstruction.
One image. A world to train in, shoot in & explore. ๐ ๐Meet HYPER3D #WorldGen: independent foreground meshes + #3DGS background, built for interaction. For robotics๐ฆพ, film๐ฌ, games๐ฎ, DCC๐ & XR. ๐ฅPowered by our Best Paper Award Project ใCASTใ. Let's Connect the Dotsโฌ๏ธ https://t.co/L2GxXtxVLb
"ZipSplat: Fewer Gaussians, Better Splats" TL;DR: feed-forward 3DGS model that decouples Gaussian placement from pixels, reconstructing unposed scenes in under a second with ~6ร fewer Gaussians while achieving state-of-the-art quality. https://t.co/T6YmrQqlHy
GPT-5.4 xhigh scored 53 on Artificial Analysis in March. By August, Qwen3.8-Flash-Next scores 56 and GLM-5.3-Flash 57 with only 6B / 18B active params per token. Yesterdayโs frontier is todayโs Flash tier.
Atlas also outputs explicit 3D from one to many input images, beating top open source reconstruction models. Passing more images gives Atlas more context: the more it sees, the less it imagines. https://t.co/5nB2eFC7UQ
I'm implementing a tiny transformer on tinyshakespear. The 2d matrices went to muon and 1d to adam. The model was still learning and generating some real words. But turns out my muon implementation was bugged and was no-oping. So 99.2% of my model was frozen at init. Adam still managed to tweak those biases into having the model still output some real words. Oh and I forgot the positional embeddings too. It's kinda crazy how you can have the shittiest implementation and a neural network still manages to learn
Karpathy's recipe for training neural networks is still relevant today. this lesson in particular is one we've been feeling very viscerally recently... neural network training can sometimes be very resilient and you may not realize there's an error for a very long time... https
I guess this is as good a way as any to announce that I'm now the Head of Evals at @every https://t.co/gPWYBoNrBY
Join The AI Book Club for a live conversation with @rasbt about his new book, Build a Reasoning Model (From Scratch)! ๐ Sep. 3, 10 AM CT ๐ Online We'll talk about how reasoning models work, how to build them, and take questions from the community. https://t.co/yev3BlkgjF
Code: https://t.co/2KHMVKk28t Model: https://t.co/Ea8AKBkv4S Paper: https://t.co/RoPmG7boHH Tech blog: https://t.co/aBTMmOL0WS Discord: https://t.co/7hr2hzUbPC
Introducing LightNav-0, our first general-purpose navigation brain. Open-sourced starting today. Trained entirely in simulation, so it scales. See scalable real2sim2real transfer across robots, tasks, and scenes. https://t.co/zyFqL8gvQL
