learn
RLHF
What RLHF means, how it works, and which tools and models on this site relate to it.
Definition
Reinforcement learning from human feedback (RLHF) is a post-training method that steers a model using preference data from people (or from other models).
How it works
Raters compare outputs; a reward model or direct preference method updates the policy. It shapes helpfulness and harmlessness more than raw knowledge.