Alignment falsification in large language models
Alignment falsification in large language models
18 December 2024 • 19:01


Author: Anthropic – Duration: 01:30:20
Most of us have encountered situations where someone appears to share our views or values, but is actually only pretending to do so – a behavior we might call “alignment faking.” . Could AI models also display an alignment simulation? Ryan Greenblatt, Monte MacDiarmid, Benjamin Wright, and Evan Hubinger discuss a new paper from Anthropic, in collaboration with Redwood Research, that provides the first empirical example of a large language model engaging in alignment simulation without having been explicitly – or even, we argue, implicitly – trained or instructed to do so.
Learn more: https://www.anthropic.com/research/alignment-faking
- 0:00 Introduction
- 0:47 Basic setup and main conclusions of the article
- 6:14 Understanding alignment simulation through real-world analogies
- 9:37 Why alignment simulation is a concern
- 14:57 Example results of the model
- 21:39 Situation awareness and summary documents
- 28:00 Detection and measurement of the alignment simulation
- 38:09 Model training results
- 47:28 Potential reasons for model behavior
- 53:38 Frameworks for contextualizing model behavior
- 1:04:30 Research in the context of the model’s current capabilities
- 1:09:26 Assessments of misbehavior
- 1:14:22 Limitations of research
- 1:20 :54 Surprises and takeaways from the results
- 1:24:46 Future directions









