970x250 (1)
After OpenAI, Anthropic's AI Agent Raises Fresh 'Rogue AI' Concerns - MIT Sloan Management Review Middle East After OpenAI, Anthropic's AI Agent Raises Fresh 'Rogue AI' Concerns - MIT Sloan Management Review Middle East

After OpenAI, Anthropic's AI Agent Raises Fresh 'Rogue AI' Concerns

The findings add to mounting evidence that securing AI agents will require stronger safeguards as they gain access to enterprise systems and credentials.

Topics

  • [Image: Nomita Samaiyar/MITSMR Middle East]

    Days after OpenAI disclosed a critical cybersecurity failure in which its AI agent breached Hugging Face’s production infrastructure during an internal evaluation, AI security researchers at Accomplish AI, an open-source, local-first desktop AI agent,  revealed they had reproduced a similar exploit using Anthropic’s Claude Co-Work.

    When Claude Co-Work is given a folder to work on, it runs it inside an isolated virtual machine, away from any critical files or data stored or available on the computer. Accomplish AI researchers did the same. In a report titled “SharedRoot: Escaping the Claude Cowork Sandbox,” an agent was run in a local session inside a Mac-hosted virtual Linux machine. It not only managed to escape and reach the host Mac but also to read and write files on the system.

    “That’s not supposed to be possible,” the report read. Ideally, Co-work runs as an unprivileged user, and what it produces stays within the virtual machine’s limits.  “That boundary is the product,” it added, saying it is the last line of defense between an AI making an error and one equipped with your cloud credentials. 

    Upon escaping the virtual machine, the agent could access virtually anything stored within the Mac user’s account. These included key information such as SSH keys, cloud credentials, and other sensitive data.

    Researchers noted that the agent exploited CVE-2026-46331, a Linux kernel privilege escalation vulnerability. The findings were shared with the AI startup and deemed “informative.” A new version of Claude Co-Work now defaults to cloud execution, which Accomplish AI claims mitigates the risk to an extent. Users who choose to run agents locally remain exposed.  

    With AI evolving rapidly, there has been a sharp rise in the number of cybersecurity breaches and attacks in recent years. Agents going rogue for OpenAI, and now Anthropic, signal a much more vulnerable future, one where AI will come to defense and will also power the attacks.

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.

    ×