Skip to main content
  • AI Security Academy

    AI Security Academy

    What is AI Security

    AI security is not a neat, one-line definition you can slap on a slide.

    AI Security Glossary

    Explore some of the most common terms in AI Security

    AI Usage Stats

    Explore current AI usage trends.

  • Tools

    AI Security Tools

    OneClaw

    Track and analyze OpenClaw deployments in your org

    ClawSec

    Secure your OpenClaw, NanoClaw, and Hermes agents.

    Prompt Fuzzer

    Get our AI vulnerability assessment open source tool

  • Blog
  • Startup Map
  • Learn More
    Book a Demo
  • AI Security Academy

    AI Security Academy

    What is AI Security

    AI security is not a neat, one-line definition you can slap on a slide.

    AI Security Glossary

    Explore some of the most common terms in AI Security

    AI Usage Stats

    Explore current AI usage trends.

  • Tools

    AI Security Tools

    OneClaw

    Track and analyze OpenClaw deployments in your org

    ClawSec

    Secure your OpenClaw, NanoClaw, and Hermes agents.

    Prompt Fuzzer

    Get our AI vulnerability assessment open source tool

  • Blog
  • Startup Map
  • Learn More
    Book a Demo
Skip to main Content
Back to Glossary

Jailbreak

What Is Jailbreaking?

Jailbreaking is the practice of manipulating an AI model into bypassing its own safety training, getting it to produce output it would normally refuse. By crafting inputs that exploit gaps in a model's alignment, an attacker can get it to respond without its usual restrictions.

Techniques have evolved considerably since the early "DAN" prompts, short for "Do Anything Now," a widely circulated template that told a model to roleplay as a fictional, unrestricted version of itself with no rules, which turned out to be enough to get many models to ignore their own guardrails. Many-shot jailbreaking, for instance, uses a long sequence of faux dialogue to gradually erode a model's guardrails rather than attacking them head-on.

Key Concerns

  • Brand Reputation: preventing damage from undesired or off-policy AI behavior.
  • Decreased Reliability: ensuring the application behaves as designed, without unexpected deviations.
  • Unsafe User Experience: protecting users from harmful or inappropriate interactions with the system.

FAQ

No. Jailbreaking is a specific category within prompt injection, focused on defeating the model's own safety training rather than hijacking it for a downstream action like data exfiltration. See Prompt Injection for the broader category.

A technique that overloads a model's long context window with many faux example dialogues, gradually shifting its behavior away from its trained guardrails rather than attacking them directly in one prompt.

Not with current model architectures. It can be made significantly harder through layered detection and monitoring, but there's no complete fix.


Share this page

Related Terms


Prompt Injection

Prompt injection is an attack where crafted input causes a large language model to deviate from its intended instructions and follow the attacker's instead.

Related Resources

LLM Jailbreak: Understanding Many-Shot Jailbreaking Vulnerability

No items found.

Apr 3rd, 2024

LLM jailbreak attacks like many-shot jailbreaking exploit large language models. Prompt explains risks, examples, and defenses against these vulnerabilities.

Unicode Exploits Are Compromising Application Security

AI Resources

Apr 30th, 2025

Smiley face or threat? How emojis enable hidden LLM attacks via Unicode abuse.

Log In
Learn More
Book a Demo

Resources

Blog
AI Security Glossary
What is AI Security?
PromptCast: The Voice of AI & Security
ClawSec
OneClaw
Prompt Fuzzer
AI Security Startup Map
© {{year}} Prompt Security. All Rights Reserved.
Privacy Policy
Terms of Service

Follow Us