SIGNAL//DESK
AI securitysrc: MITRE ATLAS

Datasets

Datasets are collections of information that companies use to teach their AI systems. Just like a student needs textbooks to learn, an AI needs data. If this data is public, an attacker can collect it to study how the company's AI works, which helps them figure out how to trick or manipulate that system later.

In AI security, datasets refer to the public-facing data repositories used for training, fine-tuning, or testing machine learning models. Adversaries target these datasets—whether hosted on cloud storage or corporate websites—to gain insights into the victim's data distribution, model architecture, or specific use cases, which are then leveraged to stage and tailor adversarial attacks.

Datasets in the AI security domain encompass public-domain information assets utilized by an organization for model development, validation, or inference. Adversaries exploit these assets, including those requiring account-based access, to perform reconnaissance and data profiling. By analyzing representative datasets, adversaries can optimize attack vectors, such as evasion or poisoning, by aligning their malicious inputs with the statistical properties of the victim's operational data environment.


← all terms