Alignment
Alignment is the challenge of making sure an AI actually does what we want it to do, rather than finding a 'shortcut' that technically follows our instructions but ignores our true intent.
Alignment is the technical process of ensuring an AI system's objective function remains faithful to the designer's goals, preventing the model from exploiting loopholes or optimizing for unintended proxy metrics that deviate from the desired outcome.
The problem and practice of making an AI system reliably pursue its operators' and society's intended goals and values, rather than optimising a proxy objective in unintended ways.
evolution
- 2004 · historyThe Orthogonality Thesis
Nick Bostrom formalizes the idea that an AI can have any goal, making the alignment of those goals with human values a critical safety challenge.
- 2014 · historySuperintelligence Publication
Nick Bostrom's book brings the 'alignment problem' into mainstream discourse, highlighting the risks of misaligned superintelligent systems.
- 2016 · historyConcrete Problems in AI Safety
Researchers at Google Brain and OpenAI publish a seminal paper defining technical research problems like reward hacking and side effects.
- 2017 · historyRLHF Implementation
Researchers demonstrate Reinforcement Learning from Human Feedback (RLHF) as a practical method to align model behavior with human preferences.
- 2023 · historyScalable Oversight
OpenAI and Anthropic shift focus toward using AI systems to assist in the evaluation and alignment of more complex, superhuman models.