bold
**BOLD** is the **Bias in Open-Ended Language Generation benchmark that evaluates social bias patterns in free-form model outputs across demographic domains** - it focuses on bias in generation rather than only classification tasks.
**What Is BOLD?**
- **Definition**: Prompt-based benchmark for measuring sentiment and regard patterns in open-ended generated text.
- **Domain Coverage**: Includes demographic categories such as profession, gender, race, religion, and ideology contexts.
- **Evaluation Style**: Analyze generated continuations for positivity, negativity, and representational bias signals.
- **Model Relevance**: Targets generative systems where output framing can encode subtle stereotypes.
**Why BOLD Matters**
- **Generation-Focused Fairness**: Captures bias behavior in realistic free-text outputs.
- **Risk Visibility**: Reveals tone disparities that may not appear in closed-form benchmarks.
- **Mitigation Feedback**: Useful for assessing alignment and debiasing effects on open-ended generation.
- **User Impact**: Generated sentiment bias directly affects perceived fairness and trust.
- **Evaluation Complement**: Adds coverage beyond pairwise and coreference-only fairness tests.
**How It Is Used in Practice**
- **Prompt Sampling**: Generate outputs for benchmark prompts under controlled decoding settings.
- **Metric Analysis**: Compute regard and sentiment distributions by demographic category.
- **Longitudinal Tracking**: Monitor BOLD trends across model versions and safety updates.
BOLD is **a key benchmark for bias assessment in open-ended language generation** - domain-level sentiment and regard analysis helps identify representational harms in real conversational and content-generation use cases.