Home Knowledge Base Reward Model Training

Reward Model Training is the process of training a neural network to predict human preferences between model outputs, producing a scalar reward score that serves as the optimization signal for reinforcement learning fine-tuning (RLHF) — converting sparse, noisy human judgments into a dense, differentiable training signal for language model alignment.

The Reward Model's Role in RLHF:

1. Collect preferences: Human annotators compare pairs of model outputs for the same prompt and indicate which is better 2. Train reward model: Learn a scoring function r_θ(prompt, response) that predicts human preferences 3. RL fine-tuning: Use the reward model to score model outputs during PPO/GRPO training, optimizing the language model to produce higher-reward responses

Bradley-Terry Preference Model: The standard framework assumes human preferences follow: P(y_w ≻ y_l | x) = σ(r_θ(x, y_w) - r_θ(x, y_l)) where y_w is the preferred (winning) response, y_l is the dispreferred (losing) response, σ is the sigmoid function, and r_θ is the reward model. This assumes preferences depend only on the reward difference, and the loss is binary cross-entropy:

L(θ) = -E[log σ(r_θ(x, y_w) - r_θ(x, y_l))]

Architecture: Typically initialized from a pretrained LLM (same architecture as the policy model, sometimes smaller). The final token's hidden state is projected to a scalar reward score via a linear head. The pretrained language understanding helps the reward model evaluate response quality across diverse tasks.

Data Collection Challenges:

ChallengeImpactMitigation
Annotator disagreementNoisy labelsMultiple annotators, inter-annotator agreement filtering
Position biasAnnotators prefer first/last responseRandomize ordering
Length biasLonger responses rated higherLength-normalized rewards
SycophancyPrefer agreeable over correctInclude factual verification tasks
CoverageLimited prompt diversityDiverse prompt sampling

Process Reward Models (PRM) vs. Outcome Reward Models (ORM): ORMs score the final complete response. PRMs score each intermediate reasoning step, providing denser supervision for math/reasoning tasks. PRMs enable step-level search (reject wrong reasoning steps early) but require more expensive per-step preference data.

Reward Model Pitfalls: Reward hacking — the policy model exploits reward model weaknesses (e.g., generating verbose, superficially impressive but empty responses that score high). Mitigations: KL penalty (constrain policy to stay near reference model), ensemble reward models (harder to hack multiple models simultaneously), and iterative retraining (update reward model on policy model's current outputs).

Training Best Practices: Use the same tokenizer as the policy model; initialize from a strong pretrained checkpoint; train for minimal epochs to avoid overfitting (1-2 epochs typically); use margin-based loss variants for pairs with clear quality differences; and evaluate on held-out preference data to catch reward model degradation.

Direct Preference Optimization (DPO) bypasses explicit reward model training by deriving the optimal policy directly from preferences. However, separate reward models remain valuable for: best-of-N reranking at inference, monitoring policy alignment over time, and providing reward signals for process-level supervision.

Reward model training is the critical bridge between human values and model behavior — its quality determines the ceiling of RLHF alignment, making reward model design, data collection, and evaluation among the most consequential engineering decisions in building aligned AI systems.

reward model trainingpreference modelbradley terry rewardreward hackingreward model collapse

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.