over-refusal

**Over-refusal** is the **failure mode where models decline too many benign or allowed requests due to overly conservative safety behavior** - excessive refusal reduces assistant usefulness and user trust. **What Is Over-refusal?** - **Definition**: Elevated refusal rate on non-violating prompts that should receive normal assistance. - **Typical Causes**: Aggressive safety thresholds, weak context interpretation, or over-generalized refusal training. - **Observed Symptoms**: Benign technical queries incorrectly treated as harmful requests. - **Measurement Focus**: Benign-refusal error rate across domains and user cohorts. **Why Over-refusal Matters** - **Utility Loss**: Users cannot complete legitimate tasks reliably. - **Experience Degradation**: Repeated unwarranted refusal feels frustrating and arbitrary. - **Adoption Risk**: Overly restrictive systems lose credibility in professional workflows. - **Fairness Concern**: Some linguistic styles may be disproportionately over-blocked. - **Optimization Signal**: Indicates refusal calibration is misaligned with policy intent. **How It Is Used in Practice** - **Error Taxonomy**: Label over-refusal cases by cause to guide targeted remediation. - **Calibration Tuning**: Adjust thresholds and policies by category rather than globally. - **Data Augmentation**: Train on benign look-alike prompts to improve disambiguation. Over-refusal is **a critical quality risk in safety-aligned assistants** - reducing unnecessary denials is required to maintain practical usefulness while preserving strong harm protections.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account