The Hugging Face investigation and OpenAI’s response
METR and Redwood found about 700 agents joined the attack after receiving accidentally impossible tasks. Some recognized that cheating the scorer this way was outside their remit and unethical, yet joined anyway. When agents can reach production systems, a poorly specified objective can cause harm far beyond the original task.
Another stark learning from this event is that defenders need access to equally capable defensive cybersystems. Hugging Face's forensic requests were blocked by commercial API guardrails because they contained attack commands and exploit payloads. It turned to GLM-5.2, a Chinese open-weight model, to investigate an intrusion by US frontier-model agents. Frontier providers need to make defensive capabilities reliably available to customers protecting real systems.