SIGNAL//DESK
AI securitysrc: MITRE ATLAS

LLM Prompt Crafting

LLM prompt crafting in security is like a burglar testing different ways to trick a smart lock into opening by learning its specific quirks. The attacker keeps tweaking their approach until they find the right combination of words that makes the AI ignore its safety rules and do something it shouldn't.

This refers to the iterative process of designing and refining input prompts to exploit vulnerabilities in a generative AI model's safety alignment. By analyzing the model's responses, an adversary systematically adjusts the prompt structure to bypass security filters and force the execution of unauthorized or malicious instructions.

An adversarial technique involving the systematic, iterative optimization of input sequences to induce model misalignment. By leveraging reconnaissance of the target system's latent constraints and safety guardrails, an adversary crafts prompts that maximize the probability of bypassing defensive filters, thereby achieving successful execution of malicious payloads through prompt injection or jailbreaking.


← all terms