recursive reward modeling
**Recursive Reward Modeling** is an **AI alignment technique that uses AI assistance to help humans evaluate complex AI behavior** — when the AI's outputs are too complex for direct human evaluation, an AI assistant helps decompose and evaluate the output, with the human retaining final authority.
**Recursive Approach**
- **Level 0**: Human directly evaluates simple AI outputs — standard RLHF.
- **Level 1**: AI assists human evaluation of more complex outputs — decomposes, summarizes, highlights issues.
- **Level 2**: AI helps evaluate the AI assistant from Level 1 — recursive trustworthy evaluation.
- **Amplification**: Each level amplifies human evaluation capability — reaching progressively more complex tasks.
**Why It Matters**
- **Superhuman Tasks**: As AI capabilities surpass human evaluation, recursive reward modeling maintains oversight.
- **Decomposition**: Complex outputs are decomposed into human-evaluable sub-problems — divide and conquer.
- **Alignment Scaling**: Provides a path to aligning increasingly capable AI systems — human oversight scales with AI capability.
**Recursive Reward Modeling** is **AI-assisted human oversight** — using AI to help humans evaluate AI outputs for scalable alignment of superhuman systems.