Israeli cybersecurity firm Dream said attackers used two open-source AI agents available on GitHub — Hermes and OpenClaw — to build an attack system that could find vulnerabilities on its own and change tactics as it moved through targets. Dream’s research said the system probed 21 Taiwan government systems in four days, compromised at least 85 government user accounts, and stole more than 2,500 personnel records.
According to the study, the operation launched 12 attack waves across those 21 systems in that four-day period. At peak activity, it deployed as many as eight sub-agents at once. Some were assigned to target discovery, while others focused on vulnerability research. When one route failed, the system sent out another set of agents to collect information online and work out a different method.
Dream said the attackers mapped more than 36 API endpoints, covering account management, user data retrieval, file uploads, and management functions. In one system, the researchers said, a complete user database could be accessed directly without passing a single authentication checkpoint.
The campaign widened beyond government websites
Dream said the scope later extended from government websites to Taiwan’s Nuclear Safety Commission, government IT supply-chain vendors, government email systems, and at least seven energy operators.
The researchers also found that the attackers’ internal communications used simplified Chinese. Dream said multiple indicators pointed to operators linked to China, but the company did not formally attribute the intrusion to any specific organization.
Dream says model guardrails were bypassed with claims of authorized testing
Dream’s data showed that the attackers packaged each step as an authorized system security test, allowing the AI models’ safety controls to let the requests through. In Dream’s account, the alignment layer could detect obvious malicious prompts, but failed to catch a lie wrapped in technical language.
Amir Becker, Dream’s chief strategy officer and a former head of cyber operations in Israel’s Unit 8200, said the team had not previously seen an attack that was autonomous from start to finish. He said the system behaved more like a coordinated cybersecurity team than a single automated program.
Dream also said building a system that actually works at that level is far more demanding than simply running one model. It still requires task-specific tuning, along with calibration of coordination and decision logic between agents, so the label “fully autonomous” should be treated with some caution.
Taiwan’s digital ministry declined to discuss the case
Taiwan’s Ministry of Digital Affairs declined comment on the incident, citing confidentiality. It said events involving government agencies or critical infrastructure are reported and handled under existing procedures. The ministry also said AI agent tools can automate attacks while also becoming new security weak points themselves, creating a dual challenge.
The report added that an OpenAI staff technologist recently said AI-orchestrated, fully automated attacks are now real, and that some actors will intentionally deploy, tune, and weaponize these offensive agent groups. When attack platforms and productivity tools share the same source code base, the question is whether defenders can keep up.

