SIGNAL//DESK
AI securitysrc: MITRE ATLAS

AI Attack Staging

AI Attack Staging is like a burglar studying a house's security system and practicing how to bypass the locks before they actually break in. The attacker uses what they know about how the AI works to customize their tools so they are more likely to succeed when they launch their final attack.

AI Attack Staging involves the preparatory phase where an adversary utilizes reconnaissance and system access to refine their attack vectors against an AI model. By leveraging knowledge of the model's architecture or training data, the attacker tailors their inputs—such as adversarial examples or poisoned samples—to maximize the probability of a successful exploit.

The adversary is leveraging their knowledge of and access to the target system to tailor the attack. AI Attack Staging consists of techniques adversaries use to prepare their attack on the target AI model. Techniques can include training proxy models, poisoning the target model, and crafting adversarial data to feed the target model. Some of these techniques can be performed in an offline manner and are thus difficult to mitigate. These techniques are often used to achieve the adversary's end goal.

evolution

  1. 2017-08 · history
    Adversarial Examples in the Physical World

    Research demonstrated that attackers could stage attacks by crafting physical objects to fool image classifiers, establishing the need for pre-attack preparation.

  2. 2019-02 · history
    GPT-2 Release and Misuse Concerns

    The release of large language models highlighted the necessity of staging phases where adversaries curate prompts and fine-tune models to maximize malicious output.

  3. 2023-02 · history
    MITRE ATLAS Framework Expansion

    MITRE formally integrated reconnaissance and resource development tactics into the ATLAS framework, codifying the staging phase of AI-specific cyberattacks.

  4. 2023-10 · history
    Executive Order on AI Safety

    The U.S. government mandated security standards that explicitly address the risks of adversarial staging and model exploitation during the development lifecycle.


← all terms