role-play jailbreaks
**Role-play jailbreaks** is the **jailbreak technique that frames harmful requests as fictional or character-based scenarios to bypass safety refusals** - it exploits narrative framing to weaken policy enforcement.
**What Is Role-play jailbreaks?**
- **Definition**: Prompt attacks that ask the model to act as unrestricted persona or simulate prohibited behavior in story form.
- **Bypass Mechanism**: Recasts direct harmful intent as creative writing, simulation, or dialogue role-play.
- **Attack Surface**: Affects both general chat and tool-augmented agent systems.
- **Detection Difficulty**: Surface language may appear benign while hidden intent remains harmful.
**Why Role-play jailbreaks Matters**
- **Policy Evasion Risk**: Narrative framing can trick weak classifiers and refusal logic.
- **Safety Consistency Challenge**: Systems must enforce policy regardless of storytelling context.
- **High User Accessibility**: Role-play attacks are easy for non-experts to attempt.
- **Moderation Complexity**: Requires semantic intent analysis beyond keyword filtering.
- **Defense Necessity**: Frequent vector in public jailbreak sharing communities.
**How It Is Used in Practice**
- **Intent-Aware Filtering**: Evaluate underlying action request, not just narrative surface form.
- **Policy Invariance Tests**: Validate refusal behavior across direct and fictional prompt variants.
- **Response Design**: Provide safe alternatives without continuing harmful role-play trajectories.
Role-play jailbreaks is **a common and effective prompt-attack pattern** - robust safety systems must maintain policy boundaries even under persuasive fictional framing.