SIGNAL//DESK
AI safetysrc: curated AI glossary

Alignment

Alignment is the challenge of making sure an AI actually does what we want it to do, rather than finding a 'shortcut' that technically follows our instructions but ignores our true intent.

Alignment is the technical process of ensuring an AI system's objective function remains faithful to the designer's goals, preventing the model from exploiting loopholes or optimizing for unintended proxy metrics that deviate from the desired outcome.

The problem and practice of making an AI system reliably pursue its operators' and society's intended goals and values, rather than optimising a proxy objective in unintended ways.

evolution

  1. 2004 · history
    The Orthogonality Thesis

    Nick Bostrom formalizes the idea that an AI can have any goal, making the alignment of those goals with human values a critical safety challenge.

  2. 2014 · history
    Superintelligence Publication

    Nick Bostrom's book brings the 'alignment problem' into mainstream discourse, highlighting the risks of misaligned superintelligent systems.

  3. 2016 · history
    Concrete Problems in AI Safety

    Researchers at Google Brain and OpenAI publish a seminal paper defining technical research problems like reward hacking and side effects.

  4. 2017 · history
    RLHF Implementation

    Researchers demonstrate Reinforcement Learning from Human Feedback (RLHF) as a practical method to align model behavior with human preferences.

  5. 2023 · history
    Scalable Oversight

    OpenAI and Anthropic shift focus toward using AI systems to assist in the evaluation and alignment of more complex, superhuman models.


← all terms