@QGallouedec
TRL v1.4 is out! two things I'm excited about: β chunked NLL loss for SFT. Way less VRAM, same loss, often faster. Qwen3-14B @ 16k seq: 58.9 β 38.9 GB. β first-class @OpenReward integration. One line wires up an env into GRPO. Plus: more chat templates, MFU helpers... https://t.co/PyEdTYNxxf