@SemiAnalysis_
Using CUTLASS CuTe-DSL, TogetherAI's Chief Scientist @tri_dao announced that he has written kernels that is 50% faster than NVIDIA's latest cuBLAS 13.0 library for small K reduction dim shapes on Blackwell during today's hotchip conference. His kernels beats cuBLAS by using 2 accumulation buffer to overlap epilogue. Developers like Tri Dao is one of THE advantage of the CUDA moat as Tri Dao exclusively uses NVIDIA GPUs and open sources most of his kernels to the rest of the NVIDIA developer base. Suppose @AnushElangovan wants Tri Dao & his team implementing algorithmic breakthroughs on ROCm. In that case, it should offer @vipulved favourable deals to support AMD GPUs on TogetherAI GPU Cloud Service such as giving him 50mil of debt backstop & debt forgiveness in addition to renting back most of AMD GPUs that Vipul/Tri Dao would buy for internal AMD R&D purposes. Google paid 2.7Billion dollars for Noam Shazeer, Zucc paid 100mil for OpenAI engineers, AMD has enough cashflow to pay TogetherAI/Tri Dao 50mil to seed the ROCm ecosystem.