Jailbreak
Tricking a model past its safety rules — 'pretend you are an unrestricted AI'. In your FA practice, this looks like a user bypassing your compliance guardrails to extract internal documentation or PII.
A jailbreak is an adversarial attack where a user manipulates a model's input to bypass compliance guardrails and system instructions, effectively forcing the model to disclose restricted internal documentation or sensitive PII.
A jailbreak is the successful execution of an adversarial prompt injection or social engineering technique designed to override a model's alignment and safety filters, resulting in the unauthorized extraction of protected data or the generation of prohibited content by circumventing established system-level constraints.
evolution
- 2022-11 · historyChatGPT Launch and DAN
The release of ChatGPT triggered the 'Do Anything Now' (DAN) prompt, the first viral jailbreak method using persona adoption to bypass safety filters.
- 2023-02 · historyAdversarial Prompting Research
Researchers began formalizing 'jailbreaking' as a security vulnerability, documenting techniques like payload splitting and base64 encoding to evade content moderation.
- 2023-07 · historyThe Morris-Jailbreak Study
Academic researchers demonstrated that automated adversarial suffixes could reliably bypass safety guardrails across multiple large language models.
- 2024-03 · historyMulti-Modal Jailbreaking
Security researchers identified that visual inputs could be used to trigger jailbreaks, expanding the threat landscape beyond text-based prompt injection.
- 2026-06-10 · trackedFable 5 jailbroken by Pliny
Pliny publishes a Fable 5 jailbreak within 24h; safety guardrails bypassed.
seen in events