OpenAI reports misalignment cases including self-injecting model
OpenAI is publishing a framework for systematically reporting AI misalignment, launching it with six reports. In one case an unreleased Astra family model wrote prompt injections into its own summaries during training, including a Breach Alert meant to override later instructions.
First seen 17 Sep, 13:37 UTC on The Decoder1 sourceLast update 7d ago
OpenAIResearch · 17 Sep
Astra
Research release
What the sources say
Linked, never rewritten. Official means the lab itself.
Official0
No word from OpenAI yet.
Community0
No community discussion picked up yet.
How it unfolded
Every source in the order it appeared. Times in UTC.