@tri_dao
By changing the recurrence from vector outer-product to matrix multiplication, we can increase the compute needed. During memory-bound decode, we get this compute and the stronger model for "free" (caveat is that training time increases since it's already quite optimized). 6/10 https://t.co/ffStdvl8AG