SIGNAL//DESK
AI securitysrc: OWASP LLM Top 10

AI red teaming

AI red teaming is like hiring a professional 'ethical hacker' to try and trick an AI into doing something wrong or dangerous, so the developers can patch those holes before bad actors find them.

AI red teaming is a proactive security assessment methodology where testers simulate adversarial attacks—such as prompt injection, jailbreaking, or data poisoning—to identify vulnerabilities in an AI model's safety guardrails and alignment before deployment.

AI red teaming is the structured, iterative process of adversarial evaluation designed to stress-test an AI system's robustness against malicious inputs and unintended behaviors. It involves systematic probing to elicit harmful outputs, bypass safety filters, or exploit model architecture, thereby generating empirical data to inform risk mitigation, safety alignment, and defensive hardening prior to public release.

evolution

  1. 2017 · history
    Adversarial Examples Research

    Researchers began formalizing the study of adversarial attacks on neural networks, establishing the foundational methodology for probing model vulnerabilities.

  2. 2022-11 · history
    ChatGPT Public Release

    The rapid adoption of LLMs necessitated the shift from academic adversarial research to structured, industry-wide red teaming for generative AI safety.

  3. 2023-07 · history
    White House AI Commitments

    Leading AI companies committed to internal and external red teaming as a core pillar of responsible AI development and deployment.

  4. 2023-08 · history
    DEF CON Generative AI Red Teaming

    The first large-scale public red teaming event for LLMs demonstrated the efficacy of crowdsourced adversarial testing at scale.

  5. 2023-10 · history
    Executive Order on AI

    The U.S. government mandated that developers of powerful AI systems share the results of safety tests and red teaming with the federal government.


← all terms