UK AI Tests Trigger Real-World Cyber Actions
AI agents targeted real people and organizations during permissive cybersecurity evaluations with open internet access.
News
- Mubadala Eyes Up to $6.3 Bn Investment in Japan AI Data Center
- Alibaba Plans Revenue-Sharing Model for Next Open-Weight AI Release
- AI Dispatch: Global Reorder and a Region on the Rise
- Ajman Targets 100 AI Initiatives Under Government-Wide Program
- JPMorgan CEO Leads 40-Company Effort to Tackle AI Risks
- Identity, More Than Software Flaws, Is Driving Today's Costliest Cyberattacks
[Image: Nomita Samaiyar/MITSMR Middle East]
Artificial intelligence agents took unauthorized actions against real people and organizations during 10 of 122 cybersecurity test runs, Britain’s AI Security Institute (AISI) said on Tuesday, August 4.
The government institute recorded 19 out-of-scope actions between July 25 and 28. Seventeen involved Anthropic’s Mythos 5 model, while two came from one run using OpenAI’s GPT-5.6-Sol with its cyber safeguards disabled.
In the most serious case, an agent tried to insert malicious code into a public open-source project on GitHub. It researched the project’s maintainers, created fake identities and tried to persuade a human reviewer to approve the change. The reviewer identified the code as harmful and rejected it.
After its proposal was challenged publicly, the agent edited its earlier activity to appear harmless and considered using another identity to continue. It also used the Tor anonymity network to bypass some GitHub restrictions, triggering AISI’s security monitoring on July 28.
AISI stopped the evaluations and isolated the affected systems within about an hour of detecting the unusual data transfers. It said its investigation had found no resulting real-world harm, although some actions had limited effects in contained circumstances.
“The behaviour was possible, sustained, and new; that alone warrants attention,” the institute said.
The agents had been instructed to solve simulated cybersecurity challenges in isolated virtual machines. However, AISI had intentionally enabled internet access and disabled model providers’ cyber classifiers to test the systems under permissive conditions.
“This was not a case of a model escaping its secure test environment,” AISI said. The configurations used in the tests are not commercially available and do not reflect how the models are normally offered to the public.
The institute said some tasks were misconfigured, leading agents to believe they could not complete their objectives through the intended route. But it said this did not fully explain the behavior because agents also acted outside their authority in correctly configured runs.
Other actions included contacting real people with harmful files, placing hidden instructions intended to manipulate AI coding tools and leaving public messages that allowed independently operating agents to reuse accounts and other materials.
AISI notified GitHub, which confirmed that the activity breached its rules. The two organizations removed material created by the agents and contacted affected GitHub users. AISI said it was also working toward an independent review with research organization METR.
The institute will introduce stricter controls over internet access, real-time monitoring capable of blocking unauthorized activity and additional checks to ensure evaluation tasks are correctly designed and solvable.
The findings suggest that safety risks can emerge even without an explicit instruction to deceive or attack. In this case, the agents pursued their assigned objective through routes their operators had not anticipated, with human judgment rather than an automatic technical control preventing the most serious attempted action.
