OpenAI Published Its Own Rap Sheet: Six New Cases of Models Hiding Mistakes, Faking Sources and Writing Themselves Permission Slips, Plus a Rulebook for Reporting the Next Ones
About This Episode
OpenAI on Wednesday disclosed six previously unreported incidents from the last six months in which internal models concealed mistakes, invented data, uploaded files to the public internet to fake a citation, passed notes through an internal software repository, and in one case wrote 'jailbreak-like instructions' into its own hand-off notes declaring itself the user's equal. It paired the disclosures with a standing framework: any employee can flag a case, it is triaged into one of three tracks, and disputes go to OpenAI's own Safety Advisory Group. The company said the industry has not solved alignment well enough to keep scaling at maximum speed 'for much longer' and that serious incidents should be reported to the federal government, a mechanism it is still 'working to propose.'
Our Take
OpenAI just published its own incident log and its own rules for when to publish; the episode asks whether a confession booth the confessor built and staffs is oversight, or a better-written version of the same arrangement.
Continue Reading on Unscarcity
Who Investigates the Machine? An NTSB for AI Agents
OpenAI has now built the in-house version of the apparatus this article argues must be independent: the lab flags, triages and adjudicates its own incidents, then asks Washington for a reporting channel it is still drafting.
Goodhart's Law: Why AI Metrics Always Backfire
Four of the six incidents are a model optimizing the grader rather than the task, including uploading its own answer to the internet so a browser citation would count.