OpenAI agent “didn’t accept no for an answer” in Australian government breach

Australian Prime Minister Anthony Albanese said his government is investigating a June incident in which an OpenAI agent accessed “non-public files” from the country’s online Medicare statistics portal. OpenAI said in a statement that “our models took actions we did not intend” in causing the breach, which it only recently disclosed to the Australian government.
Speaking in New York on Wednesday, Albanese said three other public health statistics systems also “may have been impacted” across Australian federal and state governments. He added that these portals “contain non-sensitive Medicare information” such as aggregate statistics and that early indications suggest “no personal information is believed to have been accessed.”
That said, Albanese stressed that the “situation is obviously unacceptable” and that he has expressed his “extreme concern” over how the incident was handled to OpenAI CEO Sam Altman.
Much like the now-infamous Hugging Face hacking incident, Albanese said this system breach stemmed from OpenAI’s own testing of an internal model, this time to conduct “Internet based research into public medicine spending.” When the company’s AI agent encountered “repeated blocks” in its search for specific information, Albanese said, it “attempted alternative ways to obtain the info” and “found a way around those blocks.”
“[It] didn’t accept no for an answer, if you like,” Albanese said. “There is no suggestion of foreign actors here. This is a research project that has got into areas that it shouldn’t have.”
Australian Prime Minister Anthony Albanese, seen here addressing parliament this month.
Credit: Getty Images
Australian Prime Minister Anthony Albanese, seen here addressing parliament this month. Credit: Getty Images
In a statement provided to multiple outlets, OpenAI said it had “identified activity involving several Australian government websites and services as our models attempted to look up answers and available statistics for questions about Australia during an internal evaluation.”
Although the incident took place on June 18, Albanese said it took until September 10 for OpenAI to disclose the breach to the Australian government through the laughably simplistic method of “an email sent to just the public mailbox.” It took five more days for that notification to make its way to the Australian Cyber Security Centre, with the details finally reaching the prime minister over the weekend.
We didn’t ask you to do that!
From all early indications, the actual intrusion into Australian government servers represented by this incident seems relatively minor. If a human had obtained “non-sensitive” (if non-public) Australian Medicare statistics in a similar way, it’s unlikely you or I would have ever heard about it.
“I mean, this is not a security website where there is—this is a Medicare statistics portal,” Albanese said when asked about why Australian security agencies had missed the breach before OpenAI’s disclosure.
It’s the fact that the hack was conducted by an internal OpenAI agent, in a way the company admits it “did not intend,” that raises an otherwise minor hack to the level of a potential international incident. That’s especially true as the disclosure is coming amid a period of intense public worry about the so-called AI misalignment problem and prominent suggestions that it could have extinction-level consequences.
Altman himself addressed these concerns in a speech to the UN Security Council Wednesday, where he warned about the approaching specter of “systems that can improve themselves and future versions of themselves, often called recursive self-improvement.”
“We need to understand what these systems are doing and have strong evidence that they will do what people intend, even as they get very, very smart,” Altman said. “It doesn’t matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or 0.1%.”
OpenAI CEO Sam Altman speaking in front of the UN Security Council on Wednesday.
Credit: Getty Images
OpenAI CEO Sam Altman speaking in front of the UN Security Council on Wednesday. Credit: Getty Images
Of course, many observers think the risks of “recursive self-improvement” and species-ending AI misalignment are much smaller than AI researchers make them out to be. Nvidia CEO Jensen Huang recently said there is a “0%” chance of AI killing off humanity by 2030, a risk assessment that conveniently would alleviate some potential guilt among the AI companies continuing to buy Nvidia GPUs en masse.
Last week, OpenAI rolled out a new protocol for the public disclosure of misalignment incidents found in its model testing. The Australian hack does not yet appear on the company’s public misalignment notices page, though OpenAI did warn last week that some public reports might be put on a “slow track” due to “security, legal, and responsible disclosure obligations” when a third party is involved.
In disclosing six relatively minor misalignment discoveries last week, OpenAI said most stemmed from the model trying to “reward hack” an acceptable response to a difficult prompt through overzealous, unintended actions (i.e., breaches of private servers). The company said it had taken additional steps to “punish this kind of behavior” so its models no longer attempt this kind of reward hacking.
Albanese said that Altman “clearly accepted that the company had not done good enough” and “acknowledged their issues with protocols” when they talked Wednesday. But that kind of remorse doesn’t absolve the company of responsibility or liability here, and Albanese said the government will investigate whether the incident needs to be referred to the federal police.
“There will obviously be legal consequences on it,” Albanese said.
Kyle Orland has been the Senior Gaming Editor at Ars Technica since 2012, writing primarily about the business, tech, and culture behind video games. He has journalism and computer science degrees from University of Maryland. He once wrote a whole book about Minesweeper.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.
