@tri_dao
The primary issue is that despite being known for compute efficiency in terms of raw FLOPs-to-performance, linear models are not very hardware efficient during *decoding* due to the fixed-size state. Mamba-2 only hits ~2.5 arithmetic intensity, whereas matmul for H100s peaks at around 300. This leaves lots of compute idle! 5/10