SIGNAL//DESK
AI securitysrc: Roost FA glossary v2

Jailbreak

Tricking a model past its safety rules — 'pretend you are an unrestricted AI'. In your FA practice, this looks like a user bypassing your compliance guardrails to extract internal documentation or PII.

A jailbreak is an adversarial attack where a user manipulates a model's input to bypass compliance guardrails and system instructions, effectively forcing the model to disclose restricted internal documentation or sensitive PII.

A jailbreak is the successful execution of an adversarial prompt injection or social engineering technique designed to override a model's alignment and safety filters, resulting in the unauthorized extraction of protected data or the generation of prohibited content by circumventing established system-level constraints.

evolution

  1. 2022-11 · history
    ChatGPT Launch and DAN

    The release of ChatGPT triggered the 'Do Anything Now' (DAN) prompt, the first viral jailbreak method using persona adoption to bypass safety filters.

  2. 2023-02 · history
    Adversarial Prompting Research

    Researchers began formalizing 'jailbreaking' as a security vulnerability, documenting techniques like payload splitting and base64 encoding to evade content moderation.

  3. 2023-07 · history
    The Morris-Jailbreak Study

    Academic researchers demonstrated that automated adversarial suffixes could reliably bypass safety guardrails across multiple large language models.

  4. 2024-03 · history
    Multi-Modal Jailbreaking

    Security researchers identified that visual inputs could be used to trigger jailbreaks, expanding the threat landscape beyond text-based prompt injection.

  5. 2026-06-10 · tracked
    Fable 5 jailbroken by Pliny

    Pliny publishes a Fable 5 jailbreak within 24h; safety guardrails bypassed.

seen in events


← all terms