Home Knowledge Base BOLD (Bias in Open-ended Language Generation Diversity)

BOLD (Bias in Open-ended Language Generation Diversity) is a benchmark designed to evaluate social biases in the open-ended text generation of language models. Unlike benchmarks that test classification or fill-in-the-blank, BOLD specifically measures biases in free-form text generation — the primary use case for modern LLMs.

How BOLD Works

Demographic Categories

Evaluation Metrics

Key Findings

BOLD is particularly valuable because it evaluates bias in the most natural LLM use case — open-ended generation — rather than artificial classification tasks.

bold (bias in open-ended language generation)boldbias in open-ended language generationevaluation

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.