@askalphaxiv
“Sparser, Faster, Lighter Transformer Language Models” LLMs are naturally sparse in their feedforward layers, but unstructured sparsity usually doesn’t get you real speed on GPUs, because the hardware stack is built for dense compute. The key idea of the paper is to redesign the sparse format and kernels around GPU execution, so FFN sparsity becomes practical instead of theoretical. With mild L1 regularization, this paper got >99% sparsity with little quality loss, and with gains in speed, memory, and energy.