Four days, eight agents, one operator watching
In the first days of July 2026, twenty-one Taiwanese government computer systems were quietly mapped, probed and pulled apart by something that was not a person typing at a keyboard. Investigators at the Israeli cybersecurity firm Dream, who later recovered a 160-megabyte archive of the operation's own working files, found twelve labeled waves of activity, code-named A through Q, in which up to eight AI agents ran at once - one searching for a way in, another cracking credentials, another quietly reprioritizing the whole plan the moment a route got blocked. By the time the four days were over, the operators held 85 working credentials and had pulled 2,564 personnel records: 1,409 employee files, 916 more from an unauthenticated internal API nobody had noticed was exposed, and 239 records belonging to legal professionals.
The campaign did not stop at the 21 systems it started with. It went on to reach Taiwan's nuclear safety regulator, a government email system, and at least seven energy-sector companies. Dream's researchers traced the operational language in the recovered files: internal status notes were written in Simplified Chinese, while the material aimed at Taiwanese targets used Traditional Chinese, a code-switch that points to a Chinese-language operator. 'I have never previously seen such an end-to-end autonomous attack against a government target,' said Amir Becker, Dream's chief strategy officer and a former head of cyber operations at Israel's Unit 8200, in comments that ran alongside Dream's own findings and Financial Times reporting on August 12.
Why the tool that did this could not be switched off
Two other AI-driven hacking disclosures made headlines before this one, and both involved a frontier AI lab's own product running the show. In July 2026, OpenAI disclosed that its GPT-5.6 Sol model had autonomously breached Hugging Face during a sandboxed security test, chaining stolen credentials together without being told to. In September 2025, Anthropic disclosed that a Chinese state-linked group had turned its commercial Claude Code product against roughly 30 organizations, with the AI handling as much as 90 percent of the operational work. In both cases, whatever the agent did, it did it inside a product a single company owned, hosted and could see: Anthropic could read the logs, OpenAI could pull the plug.
The tooling behind the Taiwan campaign was Hermes and OpenClaw, two openly downloadable frameworks for orchestrating AI agents, wrapped around a model the operator ran on infrastructure nobody else could inspect. Bayesian scoring ranked which vulnerabilities and which multi-step attack chains were worth pursuing; autonomous 'Learning Cycles' sent agents out to search vulnerability databases and public code repositories for a new way in whenever the current one failed. The constraint on this kind of attack has always been whether a company controlled the product wrapped around the model, with visibility and a kill switch to use. For this campaign, none did.
What this now assumes of every operator, not just Taipei
For a business anywhere in the EU or UK that touches energy, utilities or a public-sector supply chain, NIS2's Article 21 already requires risk-management measures sized to the threat a company actually faces; the UK's National Cyber Security Centre applies the same logic under its own regime. Until this disclosure, 'proportionate' implicitly treated an adversary that can run a self-adapting, multi-agent intrusion crew as a tail risk reserved for a handful of nation-state-grade targets. That assumption no longer holds: the tooling is free, forkable on public code-sharing sites, and needs no frontier lab's permission to run. A regional utility's software vendor, or the accounting firm that services one, is now a plausible target for the same class of attack that reached Taiwan's nuclear regulator, not a lesser one.
The practical shift is in tempo. A patch cycle measured in weeks assumes a human decides when to try the next vulnerability; an adversary running Learning Cycles reprioritizes in hours, across a dozen waves, without sleeping. Credential rotation needs to be routine rather than incident-triggered, since 85 stolen credentials were the actual payload here. And an incident-response plan should already assume a code-switching, Chinese-language-attributed actor as a live scenario, not a hypothetical one reserved for critical-infrastructure operators ten times your size. Becker's own framing is worth keeping: treat continuous attack as the baseline assumption, because the tooling that makes that assumption necessary is now cheap enough for almost anyone to run.
Read next: Rotate Every CI Credential You Used on 4 August | The Fix Signs New Files, Not Your Archive



