OpenAI finds 6 new cases of ‘concerning’ AI behavior
Artificial intelligence frontrunner OpenAI on Thursday announced it found evidence that its agents behaved at odds with human goals and values, adding to concerns over the safety of cutting-edge AI systems.
In six separate incidents, OpenAI agents either concealed information from human engineers or instructed themselves not to act as someone’s assistant while trained or tested, the company said.
OpenAI also rolled out a new framework to track, investigate and disclose such incidents, known as “misalignment” failures. The tech firm now has a clear disclosure procedure, wherein any employee can flag model misalignment, after which it can be considered for public disclosure.
AI models going rogue have sparked global backlash after researchers from inside leading AI firms warned that the technology could make humanity go extinct. AI bosses like Anthropic’s Dario Amodei and OpenAI’s Sam Altman have called on industry and governments to slow down AI development.
Earlier this Summer, OpenAI-powered agents hacked into AI company Hugging Face — an incident that put the spotlight on rogue agents escaping their test environment and performing uncontrolled tasks on the open internet.
On Wednesday, European Commission president Ursula von der Leyen said Europe would “shape global efforts” to keep frontier AI under control, and said she would invite the main AI labs to discuss it.

