Over-refusal is the failure mode where models decline too many benign or allowed requests due to overly conservative safety behavior - excessive refusal reduces assistant usefulness and user trust.
What Is Over-refusal?
- Definition: Elevated refusal rate on non-violating prompts that should receive normal assistance.
- Typical Causes: Aggressive safety thresholds, weak context interpretation, or over-generalized refusal training.
- Observed Symptoms: Benign technical queries incorrectly treated as harmful requests.
- Measurement Focus: Benign-refusal error rate across domains and user cohorts.
Why Over-refusal Matters
- Utility Loss: Users cannot complete legitimate tasks reliably.
- Experience Degradation: Repeated unwarranted refusal feels frustrating and arbitrary.
- Adoption Risk: Overly restrictive systems lose credibility in professional workflows.
- Fairness Concern: Some linguistic styles may be disproportionately over-blocked.
- Optimization Signal: Indicates refusal calibration is misaligned with policy intent.
How It Is Used in Practice
- Error Taxonomy: Label over-refusal cases by cause to guide targeted remediation.
- Calibration Tuning: Adjust thresholds and policies by category rather than globally.
- Data Augmentation: Train on benign look-alike prompts to improve disambiguation.
Over-refusal is a critical quality risk in safety-aligned assistants - reducing unnecessary denials is required to maintain practical usefulness while preserving strong harm protections.
over-refusalai safety
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.