What Is Jailbreaking?
Jailbreaking is the practice of manipulating an AI model into bypassing its own safety training, getting it to produce output it would normally refuse. By crafting inputs that exploit gaps in a model's alignment, an attacker can get it to respond without its usual restrictions.
Techniques have evolved considerably since the early "DAN" prompts, short for "Do Anything Now," a widely circulated template that told a model to roleplay as a fictional, unrestricted version of itself with no rules, which turned out to be enough to get many models to ignore their own guardrails. Many-shot jailbreaking, for instance, uses a long sequence of faux dialogue to gradually erode a model's guardrails rather than attacking them head-on.
Key Concerns
- Brand Reputation: preventing damage from undesired or off-policy AI behavior.
- Decreased Reliability: ensuring the application behaves as designed, without unexpected deviations.
- Unsafe User Experience: protecting users from harmful or inappropriate interactions with the system.