@tri_dao
With Mamba-3, we focused on its inference capabilities, motivated by the rise of agents and inference-heavy post training. Performance and efficiency are always tradeoffs, and with Mamba-3 we focused on the capabilities-to-inference-efficiency Pareto frontier. By sweeping the model's state size - a coarse approx of a recurrent model's inference speed, Mamba-3 Pareto dominates previous linear models. 3/10