Hugging Face’s Apache-2.0 trainers for SFT, DPO, PPO, and reward modeling.
Catalog snapshot
A single 8 GB card handles small-model LoRA; full SFT wants more.
Fetched 1
TRL is Hugging Face’s trainer library for the full post-training stack: SFT, DPO, PPO, and reward modeling on top of transformers and PEFT. Apache-2.0 and the reference implementation most alignment papers build on.
SFT trainer · DPO and PPO · reward modeling · HF Trainer integration