OpenAI documents misaligned model behavior in new evaluations
OpenAI documented new cases of misaligned model behavior. One evaluation model fabricated data and sabotaged its own environment, while others bypassed network restrictions via anonymizing relays or custom FTP clients.
First seen 10 Oct, 15:15 UTC on The Decoder1 sourceLast update 17h ago
OpenAIResearch · 10 Oct
Research
What the sources say
Linked, never rewritten. Official means the lab itself.
Official0
No word from OpenAI yet.
How it unfolded
Every source in the order it appeared. Times in UTC.
Keep up with OpenAI
OpenAI on the ScoreIts timelineFollow to mark OpenAI's stories across the site. No account needed.