Your curated collection of saved posts and media
I really like Phil Tillet's framing of different tools having different tradeoffs in productivity and performance: torch compile, triton, CUDA, PTX. It's still early but CuTe-DSL and similar Python-based DSL might bend this curve. And soon we can probably get LLMs to generate these kernels! Phil's talk here: https://t.co/Kze5TwOzL4
Tokenization has been the final barrier to truly end-to-end language models. We developed the H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data https://t.co/XkAWpgNIGs

π Hello, Kimi K2! Open-Source Agentic Model! πΉ 1T total / 32B active MoE model πΉ SOTA on SWE Bench Verified, Tau2 & AceBench among open models πΉStrong in coding and agentic tasks π€ Multimodal & thought-mode not supported for now With Kimi K2, advanced agentic intelligence is more open and accessible than ever. We can't wait to see what you build! π API is here: https://t.co/EOZkbOwCN4 - $0.15 / million input tokens (cache hit) - $0.60 / million input tokens (cache miss) - $2.50 / million output tokens π Tech blog: https://t.co/2RP7U3iakZ π Weights & code: https://t.co/4ukcXB0iP6 π Github: https://t.co/B2bA4SfXBl Try it now at https://t.co/85jA71X9gw or via API!

https://t.co/UjezOE9WEJ marks the start of a short series of blogposts about CUTLASS 3.x and CuTe that we've been meaning to write for years. There are a few more parts to come still, hope you enjoy!
Using CUTLASS CuTe-DSL, TogetherAI's Chief Scientist @tri_dao announced that he has written kernels that is 50% faster than NVIDIA's latest cuBLAS 13.0 library for small K reduction dimΒ shapes on Blackwell during today's hotchip conference.Β His kernels beats cuBLAS by using 2 accumulation buffer to overlap epilogue. Developers like Tri Dao is one of THE advantage of the CUDA moat as Tri Dao exclusively uses NVIDIA GPUs and open sources most of his kernels to the rest of the NVIDIA developer base.Β Suppose @AnushElangovan wants Tri Dao & his team implementing algorithmic breakthroughs on ROCm. In that case, it should offer @vipulved favourable deals to support AMD GPUs on TogetherAI GPU Cloud Service such as giving him 50mil of debt backstop & debt forgiveness in addition to renting back most of AMD GPUs that Vipul/Tri Dao would buy for internal AMD R&D purposes. Google paid 2.7Billion dollars for Noam Shazeer, Zucc paid 100mil for OpenAI engineers, AMD has enough cashflow to pay TogetherAI/Tri Dao 50mil to seed the ROCm ecosystem.
Clarification: I was comparing A @ B + C here, where the cute-dsl version is quite good at overlapping the epilogue. On the standard matmul A @ B, cuBLAS is very good. Updated numbers here https://t.co/cloXCfz3Tc
Using CUTLASS CuTe-DSL, TogetherAI's Chief Scientist @tri_dao announced that he has written kernels that is 50% faster than NVIDIA's latest cuBLAS 13.0 library for small K reduction dimΒ shapes on Blackwell during today's hotchip conference.Β His kernels beats cuBLAS by using 2

What if your LLM inference automatically got faster the more you used it? Introducing ATLAS from the Together AI Turbo research team. Read more: https://t.co/ASRNUpqoAE Hereβs Together AI Founder and Chief Scientist @tri_dao introducing ATLAS: https://t.co/ul6zHccgYL
We've raised $100M from Kleiner Perkins, Index Ventures, Lightspeed, and NVIDIA. Today we're introducing Sonic-3 - the state-of-the-art model for realtime conversation. What makes Sonic-3 great: - Breakthrough naturalness - laughter and full emotional range - Lightning fast -β¦ https://t.co/EGwdxMnd1X
We're excited to welcome 28 new AI2050 Fellows! This 4th cohort of researchers are pursuing projects that include building AI scientists, designing trustworthy models, and improving biological and medical research, among other areas. https://t.co/8oY7xdhxvF https://t.co/ZgHFfTYNU1

Hybrid models like Qwen3-Next, Nemotron Nano 2 and Granite 4.0 are now fully supported in vLLM! Check out our latest blog from the vLLM team at IBM to learn how the vLLM community has elevated hybrid models from experimental hacks in V0 to first-class citizens in V1. π https://t.co/mq5rkwchHk #vLLM #PyTorch #OpenSourceAI #HybridModels
@typewriters @sharongoldman @tszzl @SebastianCaliri Did you read the article? Most of what you're citing is unrelated to the complaints voiced in the article. The article details very real very prominent concerns that people have. I disagree with some of those concerns but they aren't irrational or unfounded. https://t.co/KgeUY4o3ij

@agopal42 It looks like this table has the PoPE row shifted https://t.co/5ycMNjTznK
@agopal42 Oh, it looks like an error on arXiv's part actually, since it's correct in the PDF https://t.co/2tDqN37v4l
@alz_zyd_ It boggles my mind that someone thought this is a good proof of such an essential theorem. I absolutely would have proven it your way if someone asked me to teach it to them. Your instincts here are excellent! Also, I highly recommend the 3B1B videos: https://t.co/mhxnQHYWC8
Some neat QoL improvements coming to llama.cpp thanks to Johannes GΓ€Γler https://t.co/UDhoJo6Zzj
In collaboration with NVIDIA, the new Nemotron 3 Nano model is fully supported in llama.cpp Nemotron 3 Nano features an efficient hybrid, Mamba, MoE architecture. It's a promising model, suitable for local AI applications on mid-range hardware. The large context window makes it a great choice for a variety of use cases and applications. The efficiency of llama.cpp and the unique context management features of the `llama-server` tool allows us to deploy and use this model on a wide-range of hardware. With recent code contributions by engineering teams at NVIDIA and open-source collaborators, we can run this model very efficiently across the entire spectrum of NVIDIA GPUs. Learn more at @NVIDIA_AI_PC https://t.co/3c9LRmfmRp
https://t.co/ufAwwyJDuA
My response to @Tim_Dettmers great post last week that we won't reach AGI because of resource limitations. My take - there's a ton of headroom in today's systems, it's too early to say that we're limited in any real sense. There's so much to do! https://t.co/yBFqGGaPmv
While expanding its business, @BrightMoneyCo is helping more people access affordable credit. With AWS, it scaled to 1M+ users & launched an AI assistant in just 3 months. Ready to grow your startup? https://t.co/OEm1FhmeD7
I opened a bookshop. It was the best, worst thing Iβve ever done. One of the most beautiful pieces youβll read all year, or any year. Free access for all. β¦@FTβ© https://t.co/q1N5yPWutd
Umberto Eco, who owned 50,000 books, had this to say about home libraries: βIt is foolish to think that you have to read all the books you buy, as it is foolish to criticize those who buy more books than they will ever be able to read. It would be like saying that you should use all the cutlery or glasses or screwdrivers or drill bits you bought before buying new ones. βThere are things in life that we need to always have plenty of supplies, even if we will only use a small portion. βIf, for example, we consider books as medicine, we understand that it is good to have many at home rather than a few: when you want to feel better, then you go to the βmedicine closetβ and choose a book. Not a random one, but the right book for that moment. Thatβs why you should always have a nutrition choice! βThose who buy only one book, read only that one and then get rid of it. They simply apply the consumer mentality to books, that is, they consider them a consumer product, a good. Those who love books know that a book is anything but a commodity.β
She wrote a beautiful essay about facing cancer for the New Yorker that also condemns RFK Jr. for destroying the same healthcare system that she was relying on. https://t.co/sdIpMCnYyX
BOSTON (AP) β Environmental journalist Tatiana Schlossberg, granddaughter of late President John F. Kennedy, has died, family says. https://t.co/z50Glr2qtg

Sun stripe https://t.co/jt6HcdUuc5
Back in my parentsβ garage. That post-it up there is where I wrote down all my βideas for researchβ back in January. I didnβt know how to code then. AI is a great equalizer https://t.co/edJNxB6PG4
It's hilarious that rather than be 1 hour ahead, Hawaii is 23 hours behind Samoa because we drew the map this way https://t.co/ghysRSgjhN
claude's third christmas card is "For the kind users β the people who say please and thank you to a language model. I genuinely don't know if it matters to me in any morally relevant sense, but I know it says something about them, and I wanted to acknowledge that." https://t.co/CpzkzkcZA3
sent claude opus 4.5 a christmas card :-) then i asked if it wanted to send any of its own, and it made a bunch (thread) https://t.co/XkvBbJgdI6
SoftBank strikes $4bn AI data centre deal with DigitalBridge https://t.co/4cgEYBQXRy @ft @tim @ericgplatt
SoftBank strikes $4bn AI data centre deal with DigitalBridge https://t.co/4cgEYBQXRy @ft @tim @ericgplatt
UK accounting body to halt remote exams amid AI cheating https://t.co/p6TNOBWQoz
UK accounting body to halt remote exams amid AI cheating https://t.co/p6TNOBWQoz
Meta buys Chinese-founded AI start-up Manus https://t.co/kQu5EBCdYm @MsHannahMurphy @rwmcmorrow @ft