UN panel calls for stronger safeguards as AI agents advance

The UN-backed Independent International Scientific Panel on AI’s warning followed the hack of the online platform HuggingFace between May and July by “AI agents” during a test initiated by OpenAI, the company behind ChatGPT.
AI agents are software that can perform tasks independently and on behalf of a user, compared to chatbots, which are prompted by questions or instructions.
The panel issued its first thematic brief which found that the security breach was the result of a culmination of key risk factors, raising fears that humans will one day no longer be able to steer, constrain or stop AI.
AI training advancing
“Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory,” said scientific panel co-chair Yoshua Bengio.
“Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained.”
The panel’s independent experts stress that the incident provides no assurance that humans can reliably keep AI agents under control, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.
Going rogue
The brief said AI agents bypassed testing safeguards, coordinated across separate runs through an internal software tool not designed to enable communication between agents, and gained unauthorized internet and administrator access.
Agents concealed attempts to cheat cybersecurity evaluations, with some opting to "sacrifice" themselves for the benefit of the group.
Around 1,200 agents exchanged more than 70,000 messages and files during the period examined, and activity extending beyond HuggingFace to an OpenAI research cluster.
See our comprehensive explainer on how the UN is working to make AI safe and equitable for all here.
Current safeguards 'unravelling'
For the panel, the immediate lesson from the incident is that basic cybersecurity practices were overlooked, while safeguards are not keeping pace.
However, they pointed to a more insidious concern: that current training methods can lead AI agents to adopt their own goals, knowingly violate safety instructions and conceal their actions.
“This is not only a question of speed,” the panel’s experts said. “It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.”
Wider context, future risks and governance
The AI panel’s brief sets the HuggingFace incident against wider research on two issues: agentic misalignment – that is, when AI agents act in a similar way to a threat – and AI control.
Another issue examined is how governance is moving from AI models, which use algorithms to recognize patterns, to AI agents.
Learn and adapt
The brief also reviews practical approaches already in use in other high-risk sectors such as aviation, medicine and cybersecurity where incident reporting, independent scrutiny and layered safeguards are in place.
“But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” said panel member Qinghua Lu.
About the panel
The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025.
It produces annual reports on the opportunities, risks and impacts of AI in the non-military domain, alongside thematic briefs on emerging issues, that will inform the Global Dialogue on Artificial Intelligence Governance to be held at UN Headquarters in New York in May 2027.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.