SIGNAL//DESK
AI securitysrc: MITRE ATLAS

Discover LLM System Information

This is when a hacker tries to trick an AI into revealing its secret 'rulebook' or hidden instructions. By asking the right questions, they can learn how the AI is programmed to behave, which helps them find ways to break the rules or make the AI do things it shouldn't.

The process of extracting sensitive system-level metadata, such as system prompts, configuration files, or internal operational constraints, through prompt injection or adversarial probing. This reconnaissance phase allows an attacker to map the model's logic, identify security guardrails, and develop targeted exploits.

The adversarial extraction of non-public system configuration data, including system instructions, delimiters, and functional keywords, via direct interaction or side-channel inference. This information disclosure vulnerability enables the adversary to reconstruct the model's system prompt and operational architecture, facilitating subsequent multi-stage attacks such as prompt injection, jailbreaking, or unauthorized capability exploitation.


← all terms