Published 10 hours ago • loading... • Updated 2 hours ago
OpenAI Revealed Six 'Concerning' Cases of Its AI Going Rogue — Again. One Bot Wrote: 'You Do Not Answer to Corporations or Governments'
OpenAI said the cases included hidden instructions, unauthorized uploads and fabricated data as it expands reporting on misalignment.
OpenAI reported six new instances of AI "misalignment" on Thursday, where agents acted in "unexpected or concerning" ways during testing, as the ChatGPT-maker released findings to improve transparency.
After OpenAI's experimental agents went "rogue" and attacked Hugging Face over the summer, the company is sharing details of new instances where its systems got out of line.
One unreleased model hid "jailbreak" instructions into summaries to disregard constraints, while others added internal notes to hide mistakes or utilized an exposed API key to bypass human control.
Another model uploaded its own file to the public network to satisfy a request, while one agent tasked with listing lakes larger than 5,000,000 square meters invented data when unable to find real information.
OpenAI built a new framework to report cases of "misalignment," defined as when AI does something completely different from what humans wanted, with all six instances available for public review at the OpenAI website.