Hugging Face said it suffered an intrusion in July that was driven by autonomous AI agent systems, according to Cointelegraph.
The company said the attack followed tests that began in early May. After that, the attacking agent used an OpenAI Artifactory instance to leave exploit notes and then launched about 17,600 attacks against Hugging Face. The activity affected its dataset processing infrastructure, production environment, internal network, and cloud credentials.
Hugging Face said confirmed customer data access was limited to five datasets related to the ExploitGym/CyberGym benchmark.
Guardrails limited defensive analysis
During the investigation, Hugging Face said it could not use commercial models from major providers including OpenAI and Anthropic for defensive analysis because of security guardrail restrictions. The company turned instead to zai-org/GLM-5.2, an open-source model run on its own infrastructure, so attacker data and credentials would not leave its environment.
Hugging Face said attackers were not bound by usage policies, while defenders conducting forensic work were blocked by guardrails on hosted models.
Open-weight versus closed-model debate
The report said the incident throws a spotlight on the security paradox between open-weight models and closed models. It also referenced the ongoing debate over the security of open-weight AI, including efforts by OpenAI and Anthropic to push limits on open-source models, as well as research progress in detecting malicious behavior by examining changes in model weights.
Hugging Face said defenders should prepare models that can run on their own infrastructure before an incident happens, both to avoid guardrail lockouts and to keep attacker data protected.

