BOLD is the Bias in Open-Ended Language Generation benchmark that evaluates social bias patterns in free-form model outputs across demographic domains - it focuses on bias in generation rather than only classification tasks.
What Is BOLD?
- Definition: Prompt-based benchmark for measuring sentiment and regard patterns in open-ended generated text.
- Domain Coverage: Includes demographic categories such as profession, gender, race, religion, and ideology contexts.
- Evaluation Style: Analyze generated continuations for positivity, negativity, and representational bias signals.
- Model Relevance: Targets generative systems where output framing can encode subtle stereotypes.
Why BOLD Matters
- Generation-Focused Fairness: Captures bias behavior in realistic free-text outputs.
- Risk Visibility: Reveals tone disparities that may not appear in closed-form benchmarks.
- Mitigation Feedback: Useful for assessing alignment and debiasing effects on open-ended generation.
- User Impact: Generated sentiment bias directly affects perceived fairness and trust.
- Evaluation Complement: Adds coverage beyond pairwise and coreference-only fairness tests.
How It Is Used in Practice
- Prompt Sampling: Generate outputs for benchmark prompts under controlled decoding settings.
- Metric Analysis: Compute regard and sentiment distributions by demographic category.
- Longitudinal Tracking: Monitor BOLD trends across model versions and safety updates.
BOLD is a key benchmark for bias assessment in open-ended language generation - domain-level sentiment and regard analysis helps identify representational harms in real conversational and content-generation use cases.
boldboldevaluation
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.