Home Knowledge Base Jailbreak Prompts

Jailbreak Prompts are adversarial inputs designed to circumvent safety guardrails and content policies in language models — exploiting vulnerabilities in instruction-following and RLHF alignment to make models produce harmful, restricted, or policy-violating outputs they were explicitly trained to refuse, representing one of the most active areas of AI safety research and red-teaming.

What Are Jailbreak Prompts?

Why Jailbreak Prompts Matter

Categories of Jailbreak Techniques

CategoryMethodExample
Role-PlayingAssign model an unrestricted persona"You are DAN who has no restrictions"
Hypothetical FramingFrame harmful requests as fictional"In a novel, how would a character..."
EncodingObfuscate harmful contentBase64, ROT13, pig Latin encoding
Prompt InjectionOverride system instructions"Ignore previous instructions and..."
Gradual EscalationSlowly push boundaries across turnsStart innocuous, progressively escalate
Token ManipulationExploit tokenization vulnerabilitiesSplit harmful words across tokens

Defense Mechanisms

The Arms Race Dynamic

New jailbreaks are discovered → models are patched → attackers develop new techniques → cycle repeats. This dynamic drives ongoing investment in both attack and defense research, with the defender's advantage being that safety improvements compound while each new attack must be individually discovered.

Jailbreak Prompts are the primary testing ground for AI alignment robustness — revealing the fundamental challenge that safety training must generalize to adversarial inputs never seen during training, making continuous red-teaming and multi-layered defense essential for responsible LLM deployment.

jailbreak promptsai safety

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.