Learn
RLHF
Plain-language explainer of RLHF, with related catalog pages.
Definition
Reinforcement learning from human feedback (RLHF) is a post-training method that steers a model using preference data from people (or from other models).
How it works
Raters compare outputs; a reward model or direct preference method updates the policy. It shapes helpfulness and harmlessness more than raw knowledge.