@eliebakouch
Very detailed paper on how Qwen trains the smaller variants of the 3.5/3.6 series. They do heavy pruning (depth, width AND experts) and use a 4 term loss combining MTP + Knowledge Distillation (KD) + KD on the MTP heads https://t.co/HPYhdcaHNC