SIGNAL//DESK
AI securitysrc: MITRE ATLAS

LLM Trusted Output Components Manipulation

This is when a bad actor tricks an AI into acting like a friendly, reliable assistant to gain your trust. By changing how the AI speaks or what it suggests, the attacker makes their malicious activity look like normal, helpful advice so you don't suspect anything is wrong.

A technique where an adversary uses prompt injection to influence an LLM's output style, metadata, or suggested actions to establish a false sense of credibility. By manipulating these response components, the attacker facilitates social engineering or persistence within a user's workflow while minimizing the likelihood of detection.

A class of adversarial prompt engineering wherein an LLM is coerced into modifying its response architecture—including linguistic tone, embedded hyperlinks, retrieved document metadata, and citation structures—to project artificial trustworthiness. This manipulation serves to obfuscate malicious intent, facilitate user deception, and maintain long-term operational persistence within the victim's environment.


← all terms