recursive reward modeling

**Recursive Reward Modeling** is an **AI alignment technique that uses AI assistance to help humans evaluate complex AI behavior** — when the AI's outputs are too complex for direct human evaluation, an AI assistant helps decompose and evaluate the output, with the human retaining final authority. **Recursive Approach** - **Level 0**: Human directly evaluates simple AI outputs — standard RLHF. - **Level 1**: AI assists human evaluation of more complex outputs — decomposes, summarizes, highlights issues. - **Level 2**: AI helps evaluate the AI assistant from Level 1 — recursive trustworthy evaluation. - **Amplification**: Each level amplifies human evaluation capability — reaching progressively more complex tasks. **Why It Matters** - **Superhuman Tasks**: As AI capabilities surpass human evaluation, recursive reward modeling maintains oversight. - **Decomposition**: Complex outputs are decomposed into human-evaluable sub-problems — divide and conquer. - **Alignment Scaling**: Provides a path to aligning increasingly capable AI systems — human oversight scales with AI capability. **Recursive Reward Modeling** is **AI-assisted human oversight** — using AI to help humans evaluate AI outputs for scalable alignment of superhuman systems.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account