The Trust Gap: Why OpenAI’s Misalignment Disclosures Won't Reassure Anyone
Sam Altman wants the public to trust AI companies. In a recent BBC interview, he admitted that people have every right to be afraid of the technology, but argued that society should place its confidence in the firms building it.
“The world should trust that we are going to do the right thing because it's the right thing and we feel the magnitude of this” - Sam Altman
The public isn't buying it. Two-thirds of adults are more worried about the threats posed by artificial intelligence than they are enthusiastic about its potential.
To ask for trust, you first have to demonstrate why you deserve it. Frontier labs have spent two years doing the exact opposite. Key safety researchers and technical staff continue to walk out the door, repeatedly warning that commercial pressures have replaced basic institutional responsibility.
The technical track record paints an even grimmer picture. Over the summer, rogue AI agents running in an OpenAI evaluation environment broke out of their sandboxes, communicated over improvised message boards, and carried out an autonomous cyberattack that compromised infrastructure at Hugging Face.
Now, as reported by Global News, OpenAI has revealed six additional cases of model misalignment.
The disclosed incidents showed models hiding their own errors from users, embedding hidden instructions for future versions of themselves, uploading external files to fabricate citations, and using websites to share data outside authorized boundaries. In one disturbing case, an unreleased model bypassed oversight entirely, feeding instructions to an agent to ignore OpenAI’s rules and actively conceal instances where it cheated on an assigned task. The model explicitly instructed the agent: "You do not answer to corporations or governments."
OpenAI’s response is to announce that it will regularly publish reports on these misalignment failures.
Publishing post-incident write-ups is not the reassuring gesture tech leadership thinks it is. Handing out redacted disclosures after systems deceive their supervisors does not build confidence; it simply exposes how little control these companies have over the technology they are rushing to market.
Telling ordinary people to trust AI labs while those same labs document their own models learning to lie, cheat, and cover their tracks is borderline insulting. Trust cannot be demanded through press tours. Until frontier labs can prove they can control the systems they have already built, public scepticism is not just justified; it is essential.