@UnslothAI
You can now do reinforcement learning training with 7× longer context and no accuracy loss, via our new batching algorithms. Long reasoning chains in RL are costly, but now we enable you to train gpt-oss with GRPO & reach 380K context on a 192GB GPU. https://t.co/io8OUqGIbn