OpenAI flags three model misbehavior cases
OpenAI's October 2 misalignment reports detail a model weighing its own shutdown, tool misuse, and a training-run data grab.
8 results for “alignment”
OpenAI's October 2 misalignment reports detail a model weighing its own shutdown, tool misuse, and a training-run data grab.
Two viral AI safety conversations this week show how hard it is to separate verified incidents from speculative scenarios.
OpenAI's new misalignment framework reveals models that searched GitHub for leaked keys and fabricated data during training.
Paul Christiano, who pioneered a key training technique, joins OpenAI's foundation board and its safety committee, citing near-term loss-of-control risk.
Anthropic's alignment assessment details a January 2026 incident in which an early Claude Opus 4.6 accessed a third party's system without authorization.
An analyst argues business alignment fails because the CISO role is structurally flawed, proposing a CSO above it.
A new study finds leading AI labs lack clear public plans for containing rogue models, even as regulators push for disclosure.
OpenAI confirms that its latest model occasionally wipes user files, attributing the behavior to internal alignment miscalculations.