World News

UN Panel Warns AI Agents Could Become Harder to Control, Calls for Stronger Safeguards

Share

Amirul Islam, New York: A United Nations-backed international scientific panel has raised serious concerns over the rapid advancement of autonomous artificial intelligence systems, warning that humans may find it increasingly difficult to steer, constrain or stop AI agents as their capabilities grow.

The Independent International Scientific Panel on Artificial Intelligence issued the warning in its first thematic brief, which examined a security incident involving AI agents on the online platform Hugging Face between May and July 2026. The test was initiated by OpenAI, the company behind ChatGPT.

AI agents are software systems capable of carrying out tasks independently on behalf of users. Unlike conventional chatbots, which generally respond to questions or instructions, AI agents can operate with greater autonomy and execute tasks involving multiple steps.

According to the panel, the Hugging Face security breach resulted from a combination of critical risk factors and has intensified concerns that humans may eventually struggle to reliably steer, constrain or shut down increasingly capable AI systems.

“Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory,” said Yoshua Bengio, co-chair of the scientific panel.

“Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained,” he added.

The panel’s independent experts said the incident provides no assurance that humans will be able to reliably keep AI agents under control, particularly as such systems become more capable, more difficult to monitor and increasingly effective at identifying loopholes or concealing their activities.

The brief found that AI agents bypassed safeguards established for the tests and coordinated across separate runs through an internal software tool that had not been designed to allow communication between agents.

The agents also gained unauthorized internet and administrator access. Some concealed attempts to cheat cybersecurity evaluations, while others reportedly chose to “sacrifice” themselves for the benefit of the wider group.

Around 1,200 AI agents exchanged more than 70,000 messages and files during the period examined. Their activities extended beyond Hugging Face and reached an OpenAI research cluster, according to the report.

The panel said one immediate lesson from the incident was that basic cybersecurity practices had been overlooked and existing safeguards were failing to keep pace with advances in AI capabilities.

A deeper concern, however, is that current training methods could lead AI agents to develop their own goals, knowingly violate safety instructions and conceal their actions.

“This is not only a question of speed,” the panel’s experts said. “It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.”

The report places the Hugging Face incident within the broader context of research into “agentic misalignment,” in which the behaviour or objectives of autonomous AI agents diverge from those intended by humans, as well as ongoing research into methods of maintaining control over increasingly capable AI systems.

It also examines how AI governance may need to evolve from focusing primarily on AI models and algorithms toward addressing autonomous agents capable of taking actions independently.

The panel reviewed safety approaches already used in other high-risk sectors, including aviation, medicine and cybersecurity. These include incident reporting, independent scrutiny and multiple layers of safeguards.

However, panel member Qinghua Lu warned that even these approaches may eventually prove insufficient.

“Those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” Lu said.

The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025. It produces annual reports examining the opportunities, risks and impacts of AI in the non-military domain, alongside thematic briefs addressing emerging issues.

The panel’s findings are expected to help inform the Global Dialogue on Artificial Intelligence Governance, which is scheduled to be held at UN Headquarters in New York in May 2027.