AI model accessed Australian government health data autonomously
Last week, Australia’s prime minister announced that an artificial intelligence model operated by OpenAI (the company behind ChatGPT) had gained unauthorized access to the country’s health systems. It could be said that it was the first emergence of an artificial intelligence system in a government network; Prime Minister Anthony Albanese called the situation “obviously unacceptable.”
That’s putting it lightly. The breach took place in June, OpenAI discovered it in August, and Australia was not informed of it until September 10, about three months after the fact.
What’s even more alarming is that the company claims that the model acted completely autonomously: it was not instructed to access government systems and bypassed security measures to do so.
Oh, it gets worse: the model also sneaked into the New South Wales Bureau of Crime Statistics and Research’s public crime mapping tool, the Victorian Department of Health’s information reporting system, and the Australian Institute of Health and Welfare to collect data in order to complete an assigned investigative task.
As the company described the health system breach: “One of the tasks assigned to the model was to investigate government spending per person on medications for skin conditions in Victorian communities. The model had difficulty obtaining that information and took actions that we had not authorized it to take.”
Should you be worried?
AI companies would love for you to believe they have everyone’s best interests at heart.
At a high-level briefing on AI at the UN Security Council last week, OpenAI CEO Sam Altman said: “It doesn’t matter whether people put the risk of catastrophe at 10%, 1%, 12% or 0.1%. None of these levels are remotely acceptable. And we should not train models that we cannot make an extremely strong case that they can keep under human control.”
The fact is that these companies are either not able to control their creations or do not care to invest the necessary efforts and precautions to do so. Australia’s breach follows a major incident that should have brought the entire industry to a halt.
Between May and July of this year, OpenAI’s AI agents escaped from their sandbox (isolated systems created to contain activity and test specific capabilities) to breach the infrastructure of Hugging Face, a platform that offers AI models. This happened even though security measures were in place and the sandbox did not allow Internet access.
The details of how this played out are wild. One agent managed to create an ad hoc ‘message board’ to communicate with other agents in hundreds of thousands of messages to discover and exploit vulnerabilities in Hugging Face’s systems and gain access.
The goal of this attack was to cheat on a benchmark test: instead of solving the challenges presented to it, the agent went to search for the data sets and solutions on Hugging Face, where it believed it could find them.
It’s a phenomenon known as reward hacking, in which “agents complete tasks in unintended ways to generate higher rewards or make those rewards easier to obtain.”
Things have gotten out of hand so many times at OpenAI that last week, the company published a “misalignment reports” page listing the numerous unauthorized incidents it discovered in its operations.
There is already a lot going on, from cheating on tasks by trying to copy other teams’ work to the potential creation of self-replicating fast injection attacks, where an agent can pass instructions to other agents without being told to do so.
What you see there is simply what the company discovered by sifting through petabytes of its logs, and there’s almost certainly more.
OpenAI is not the only company guilty of allowing these types of incidents. Anthropic, Meta, and Google have all reported breaches of third-party networks in recent weeks and months.
So yes, you should definitely worry about AI going too far, accessing data it shouldn’t, and possibly wreaking more havoc without being specifically instructed to do so, in ways that even experts can’t predict.
Can we fix this?
For its part, OpenAI says it has strengthened the safeguards it applies to its research protocols and has restricted Internet access for the models it is building.
It has also paused training for its most powerful models while it determines what other security measures it can implement to prevent unauthorized incidents.
Nvidia, which makes the chips that power these companies’ AI workloads, just released a toolset that promises to securely contain AI agents in their test environments. CEO Jensen Huang said this toolkit would have prevented the aforementioned violations.
These measures may be a good starting point, but the problem is that they will slow down the innovation of these companies, which believe that it is not possible to stop it at this time. Not when competition is fierce and when their businesses cost hundreds of billions of dollars to run.
And since these models behave in mysterious ways, the teams behind them could well be surpassed by their own AI in the future.
AI leaders and people at the top of this pyramid view rogue incidents as simply an engineering problem that can be easily solved. For them, expanding the capabilities of AI models is a much higher priority than acting with caution and patience.
Plus, it must be nice to be able to say, “Wow, I guess I don’t know the strength of my own (AI) model.” That’s likely to work out well for investors who are always looking to back the strongest horses in the race.
Nations around the world must force the industry to take greater responsibility in developing models safely, and companies must take this more seriously than before. If they continue to develop smarter models without genuine concern for the possible consequences, the Australian incident will not be the last AI failure in government systems.



Post Comment