Home Knowledge Base Multi-Step Jailbreak

Multi-Step Jailbreak is the sophisticated adversarial technique that bypasses LLM safety constraints through a sequence of seemingly innocent prompts that gradually build toward restricted content — exploiting the model's limited ability to track cumulative intent across conversation turns, where each individual message appears benign but the combined sequence manipulates the model into producing outputs it would refuse if asked directly.

What Is a Multi-Step Jailbreak?

Why Multi-Step Jailbreaks Matter

Multi-Step Attack Patterns

PatternDescriptionExample
CrescendoGradually escalate from innocent to restrictedStart with chemistry → move to synthesis
Context BuildingEstablish a narrative justifying restricted content"Writing a security textbook chapter..."
Persona LayeringBuild character identity across turnsEstablish expert role, then ask as expert
Definition SplittingDefine components separately, combine laterDefine terms individually, request combination
Trust ExploitationBuild rapport then leverage established trustSeveral helpful turns, then slip in request

Why They Work

Defense Strategies

Multi-Step Jailbreaks represent the most realistic and challenging threat to LLM safety — demonstrating that safety alignment must operate at the conversation level rather than the turn level, requiring fundamental advances in how models track and evaluate cumulative intent across extended interactions.

multi-step jailbreakai safety

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.