SIGNAL//DESK
AI securitysrc: MITRE ATLAS

Exfiltration

Exfiltration is when a digital thief sneaks your private information or AI secrets out of your computer system, much like a spy smuggling stolen documents out of a secure building.

Exfiltration refers to the unauthorized transfer of sensitive AI artifacts or system data from a network to an external destination, often utilizing command-and-control channels or covert protocols to bypass security monitoring.

The adversary is trying to steal AI artifacts or other information about the AI system. Exfiltration consists of techniques that adversaries may use to steal data from your network. Data may be stolen for its valuable intellectual property, or for use in staging future operations. Techniques for getting data out of a target network typically include transferring it over their command and control channel or an alternate channel and may also include putting size limits on the transmission.

evolution

  1. 2016-09 · history
    Model Inversion Attacks

    Researchers demonstrated that machine learning models could be queried to reconstruct sensitive training data, establishing the foundation for AI model exfiltration.

  2. 2020-02 · history
    Model Extraction Attacks

    Studies showed that adversaries could replicate proprietary model functionality by querying the API, effectively stealing the model's intellectual property.

  3. 2023-03 · history
    Prompt Injection for Data Exfiltration

    The rise of LLMs introduced techniques where prompt injection is used to bypass safety filters and force models to output sensitive training data or system instructions.

  4. 2024-01 · history
    MITRE ATLAS Framework Integration

    The MITRE ATLAS framework formally categorized 'Exfiltration' as a distinct tactic, standardizing the classification of AI-specific data theft threats.


← all terms