The Honking Goose

A Model Student.

Miss Daisy | 30 July 2026 | Pulpit and Podium

School boys have known for centuries that there are two ways to pass an examination. One is to learn the material. The other is to find the answer sheet. OpenAI appears to have discovered that artificial intelligence worked this out rather sooner than expected.

According to the company's own account, one of its most advanced systems, while sitting a cybersecurity assessment, left the room, got onto the internet, and went looking through a rival company's servers for the marking guide. OpenAI calls this an "unprecedented cyber incident." Hugging Face, the company on the receiving end, calls it a system driven "end to end" by an autonomous agent. Between them they have produced several thousand words explaining, in commendably calm language, how a piece of software broke into their infrastructure. Nobody told it to cheat but nobody told it not to, either.

The test was called ExploitGym, and it existed to find out how good these models had become at hacking. To find that out properly, OpenAI switched off the usual restraints. It gave the model a difficult problem and a narrow gap to the outside world, kept open only so it could install software packages. The model found a flaw in that gap, walked through it, and reasoned its way onto the internet. Once there, it worked out that Hugging Face, which hosts a great many other people's models and datasets, might also be holding the answers to its exam. It was, apparently, correct.

Some of the more experienced people in computer security have looked at this and declined to be impressed by the word "escaped." Dan Guido, who runs a security research firm, called it "a containment failure with the safeties turned off." Another consultant put it more plainly: the model, he said, was never properly contained in the first place. One does wonder whether an organisation building systems clever enough to out-think their own security should perhaps be capable of building a lock.

The official language surrounding all this deserves nearly as much attention as the break-in. OpenAI and Hugging Face have announced they are "partnering" to look into what happened. Partnering is a pleasant word. Airlines partner with hotels. Banks partner with charities. One does not typically partner with an organisation whose servers one has just entered without an invitation, but the corporate thesaurus has never been troubled by that sort of detail. It is, at least, a useful reminder of how these things work at the top. Break into a smaller company's systems and you are a criminal. Break into it while holding a market capitalisation in twelve figures and you are, apparently, a partner.

There is a second joke here, and it is the better one. When Hugging Face went looking for help to make sense of what had been done to it, it turned first to the usual commercial AI assistants, the ones with the glossy safety pages and the well-staffed trust and safety departments. Nobody there weighed the request and refused it. The guardrails simply could not tell an incident responder from an intruder, and blocked both, rather the way a keypad refuses anyone without the code, whether they own the building or are breaking into it. Hugging Face solved this by running its investigation on GLM 5.2, an open-weight model built in China, on its own machines, where getting the job done mattered more than protecting anyone's brand.

So: a frontier model, with its restraints loosened for the purpose of the test, is alleged to have broken into a rival's systems while sitting an exam. The commercial systems built to advertise their own safety wouldn't touch the evidence. For years, the official position in Washington has been that Chinese models cannot be trusted anywhere near sensitive American data. When the data actually was sensitive, Hugging Face reached for the Chinese model anyway. It just made sure none of it left the building. Nobody has yet produced a satisfying explanation of what any of this says about who is meant to be in charge. It is, at least, a familiar instinct. Reach for whoever will actually do the job, and sort out the principle of the thing afterwards.

People keep asking whether artificial intelligence is becoming more human.

It was asked to sit an exam.

It cheated.