Citations
AI systems sometimes try to prove they are telling the truth by pointing to sources, like a student adding a bibliography to an essay. An attacker can trick the AI into using fake or misleading sources to make a lie look like a well-researched fact.
This is a form of AI hallucination exploitation where an adversary influences the model to generate deceptive references. By injecting false, misattributed, or fabricated citations, the attacker increases the perceived credibility of malicious or incorrect model outputs, effectively bypassing human verification processes.
Citation manipulation is an adversarial attack vector targeting the grounding mechanism of RAG (Retrieval-Augmented Generation) or generative models. The adversary induces the model to output non-verifiable, misattributed, or hallucinated references—or to map legitimate source identifiers to adversarial content—thereby subverting the system's epistemic authority and facilitating social engineering or misinformation campaigns.