OpenAI reports alignment failures in an internal long-horizon research model
The unnamed model, credited in May 2026 with disproving the decades-old Erdős unit distance conjecture, had spent about an hour finding the exploit.
OpenAISafety & alignment