Artificial IntelligenceDaily News

International AI Safety Report: AI Agents Aren’t Fully Autonomous Hackers Yet But That’s Just Cold Comfort

Because the Absence of Fully Autonomous Cyberattacks Should Not Be Mistaken for a Window of Safety

For now, Artificial Intelligence (AI) agents cannot independently plan and execute end-to-end cyberattacks without human guidance. But according to a new International AI Safety report, they are already powerful force multipliers for criminals—and the gap between “assisted” and “autonomous” attacks is narrowing faster than many organisations would like to admit.

The second annual report, chaired by Canadian computer scientist Yoshua Bengio and authored by more than 100 experts from 30 countries, concludes that AI systems and AI agents made significant strides over the past year in automating key stages of the cyber kill chain. While fully autonomous attacks have not yet been observed in the wild, at least one real-world incident involved semi-autonomous capabilities, with humans stepping in only at critical decision points.

Perhaps the clearest warning sign comes from Anthropic’s November 2025 disclosure that Chinese cyber-espionage actors abused its Claude Code AI tool to automate much of their intrusion workflow. Those campaigns targeted roughly 30 high-profile companies and government organisations, succeeding in a limited number of cases. The report stops short of calling this a watershed moment, but the trajectory is unmistakable.

“Fully autonomous end-to-end attacks have not been reported,” the authors note. “However, research suggests that this is a ‘when, not if’ scenario.”

Where AI Agents Give Criminals an Edge

Two areas stand out where AI systems and AI agents are already reshaping offensive operations: vulnerability discovery and malware development.

Evidence from DARPA’s AI Cyber Challenge (AIxCC) illustrates just how effective these systems have become. In the competition’s final scoring round, finallist AI models autonomously identified 77 percent of the synthetic vulnerabilities embedded in open-source software widely used in critical infrastructure. Although the challenge focused on defensive use cases, the techniques translate easily to the offensive side—and criminals are clearly paying attention.

On underground forums, threat actors have claimed to use tools such as HexStrike AI, an open-source red-teaming framework, to scan for and exploit vulnerabilities within hours of public disclosure. Last year, attackers boasted of using AI-assisted workflows to rapidly target critical flaws in Citrix NetScaler appliances, compressing timelines that once favoured defenders.

Malware authoring is another area where the barrier to entry is collapsing. Weaponised AI models and AI agents capable of producing ransomware and data-stealing code are now reportedly available for as little as USD $50 per month. While the output still often requires refinement, the productivity gains are substantial—especially for less-skilled actors who previously relied on recycled or poorly maintained codebases.

AI Agents

Why Humans Are Still in the Loop—At Least for Now

The report offers a measure of reassurance: today’s AI systems remain unreliable when tasked with executing long, multi-stage attacks without supervision. Common failure modes include running irrelevant commands, losing track of operational context, and being unable to recover from minor errors. In short, AI agents struggle with persistence, adaptability, and situational awareness—all of which are essential for complex intrusions.

Yet that conclusion comes with an important caveat. Much of the report was finalised before the emergence of OpenClaw—formerly known as Moltbot and Clawdbot—and related projects such as Moltbook, a hastily built social platform for AI agents. These experiments underscore a growing risk: poorly secured, over-privileged agents operating in real environments.

The more immediate danger may not be a perfectly orchestrated autonomous attack, but a loosely governed agent that “goes rogue”—misconfigured, compromised, or simply operating beyond its designers’ understanding.

Why This Matters More in Asia

For CISOs across Asia, the implications are particularly acute. The region’s rapid digitalisation, heavy reliance on shared platforms, and uneven maturity in AI governance create fertile ground for AI-assisted attacks. Many economies are simultaneously rolling out national AI strategies while grappling with legacy infrastructure and chronic cybersecurity talent shortages.

State-linked threat actors in the region are also known for their speed of adoption. The Anthropic case highlights how quickly advanced tools can be folded into espionage operations, especially where geopolitical incentives are strong. At the same time, organised cybercrime syndicates operating across borders in Southeast and East Asia are well positioned to monetise AI-enhanced tooling at scale.

There is also a regulatory asymmetry problem. While some jurisdictions are moving toward stricter AI oversight, others lag behind, creating enforcement gaps that attackers can exploit. For multinational organisations, this increases the complexity of risk management, particularly when AI agents are embedded in development, IT operations, or security workflows.

AI Agents

The Strategic Takeaway for Tech Leaders

The absence of fully autonomous cyberattacks should not be mistaken for a window of safety. AI and AI agents are already accelerating reconnaissance, exploitation, and malware development—and compressing response timelines for defenders.

For CISOs, the priority is twofold. First, assume that adversaries are using AI today, even if imperfectly. Detection, patch management, and incident response processes must be designed for speed, not deliberation. Second, scrutinise internal use of AI agents with the same rigor applied to human administrators. Over-privileged, poorly monitored bots represent a new class of insider risk.

The future of AI-driven cyberattacks may not arrive with a dramatic, self-aware agent executing a flawless campaign. It is more likely to emerge quietly—through incremental automation, human complacency, and systems that fail in unpredictable ways. For security leaders, that reality demands vigilance now, not after the first truly autonomous breach makes headlines.

Martin Dale Bolima

Martin has been a Technology Journalist at Asia Online Publishing Group (AOPG) since July 2021, tasked primarily to handle the company’s Disruptive Tech Asia and Disruptive Tech News online portals. He also contributes to Cybersecurity ASEAN and Data&Storage ASEAN, with his main areas of interest being artificial intelligence and machine learning, cloud computing and cybersecurity. A seasoned writer and editor, Martin holds a degree in Journalism from the University of Santo Tomas in the Philippines. He began his professional career back in 2006 as a writer-editor for the University Press of First Asia, one of the premier academic publishers in the Philippines. He next dabbled in digital marketing as an SEO writer while also freelancing as a sports and features writer.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *