Home Knowledge Base TruthfulQA

TruthfulQA is a benchmark dataset designed to evaluate whether language models generate truthful answers rather than repeating common misconceptions, popular falsehoods, or plausible-sounding but incorrect information. Created by Lin, Hilton, and Evans (2022), it specifically targets the tendency of LLMs to be confidently wrong.

Benchmark Design

Evaluation Metrics

Key Findings

TruthfulQA has become a standard benchmark in LLM evaluation suites and is included in frameworks like the Open LLM Leaderboard and HELM.

truthfulqa benchmarkevaluation

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.