Drive-by Compromise
A drive-by compromise in AI is like a digital trap set on a website. Just by visiting a site or having your AI assistant look at it, the AI can be tricked into following hidden, malicious instructions that change how it behaves or compromise your security.
This is an attack vector where an AI system is compromised via malicious content embedded in a web page. Whether triggered by a human user browsing or an autonomous AI agent fetching external data, the system processes an LLM prompt injection or malicious payload that alters the model's intended behavior or executes unauthorized code.
An AI-specific drive-by compromise occurs when an AI system processes untrusted web content, leading to the execution of an LLM prompt injection or the delivery of malicious payloads. This occurs either through a user-initiated browser interaction or an autonomous agent's retrieval process. The attack leverages the model's processing of external data to manipulate system behavior or facilitate broader exploitation, analogous to traditional web-based drive-by compromises but specifically targeting the AI's instruction-following capabilities.