SIGNAL//DESK
AI securitysrc: MITRE ATLAS

System Prompt

A system prompt is like a secret set of house rules given to an AI by its creators. If a bad actor manages to trick the AI into revealing these rules, they can learn how the AI works and find ways to break its safety boundaries.

The system prompt is the foundational set of instructions defining an LLM's persona, operational constraints, and safety protocols. Unauthorized extraction of these instructions via prompt injection allows adversaries to map the model's internal logic and bypass security guardrails.

The system prompt constitutes the developer-defined preamble that establishes the model's behavioral parameters and safety alignment. Adversarial discovery of this context window content facilitates prompt leakage, enabling attackers to reverse-engineer system capabilities and identify vulnerabilities to circumvent safety-critical guardrails.


← all terms