Home Knowledge Base Universal adversarial triggers

Universal adversarial triggers are short sequences of tokens that, when prepended or appended to any input, reliably cause a language model to produce specific unwanted behaviors — such as generating toxic content, making incorrect predictions, or ignoring safety guidelines. Unlike input-specific adversarial examples, these triggers are input-agnostic and work across many different prompts.

How They Are Found

Properties

Examples of Triggered Behavior

Defenses

Universal adversarial triggers remain one of the most concerning AI safety vulnerabilities, demonstrating that aligned language models can be systematically subverted.

universal adversarial triggersai safety

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.