@HuggingPapers
ReMix: Reinforcement routing for mixtures of LoRAs A new approach to prevent routing weight collapse in Mixture-of-LoRAs models using non-learnable routing weights and the RLOO gradient estimator, ensuring all active LoRAs contribute equally to boost expressive power. https://t.co/eXRVT6AW2d