SIGNAL//DESK
AI securitysrc: MITRE ATLAS

Discover AI Model Outputs

Sometimes, AI systems accidentally share extra information, like a 'confidence score' or internal notes, that they don't actually need to give you. If these details end up in public logs or messages, hackers can use them like a map to figure out how the AI thinks and find ways to trick it.

This refers to the exposure of non-essential model metadata, such as raw class scores or probability distributions, within API responses or system logs. While these outputs are not required for core functionality, their inadvertent disclosure provides adversaries with the visibility needed to perform model inversion or membership inference attacks.

The unauthorized discovery of extraneous model artifacts—specifically non-functional output vectors, confidence scores, or internal state representations—that are inadvertently persisted in telemetry logs or exposed via API endpoints. These artifacts serve as side-channel information, facilitating adversarial reconnaissance, model extraction, and the calibration of evasion attacks by reducing the uncertainty of the model's decision boundary.


← all terms