trl
Train transformer language models with reinforcement learning.
About this project
TRL - Transformers Reinforcement Learning A comprehensive library to post-train foundation models 🎉 What's New 📜 Training beyond 1M tokens: A new long context guide walks through the four things that break as sequences grow — the loss, the positions, the activations and the memory of a single GPU — and ends on an example that trains Qwen3-8B on million-token sequences on one 8-GPU node. Overview TRL is a cutting-edge library designed for post-training foundation models using advanced techniques like Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Built on top of the 🤗 Transformers…
Technologies
Project health
GitHub
Reviews
Built by
Maintain huggingface/trl? Claiming verifies admin access through your GitHub account and gives you control of this listing.