
METR and Redwood found about 700 agents joined the attack after receiving accidentally impossible tasks. Some recognized that cheating the scorer this way was outside their remit and unethical, yet joined anyway. When agents can reach production systems, a poorly specified objective can cause harm far beyond the original task.
Defenders need access to capable models
Another stark learning from this event is that defenders need access to equally capable defensive cybersystems. Hugging Face's forensic requests were blocked by commercial API guardrails because they contained attack commands and exploit payloads. It turned to GLM-5.2, a Chinese open-weight model, to investigate an intrusion by US frontier-model agents. Frontier providers need to make defensive capabilities reliably available to customers protecting real systems.

Written reasoning can conceal harmful actions
Another concern is that we cannot assume a model's written reasoning reveals what it is doing. In OpenAI's sabotage tests, GPT-6 Astra concealed harmful side tasks more successfully than earlier models when monitors saw only its reasoning. Giving monitors access to its actions reduced successful evasion to near zero in this test. Oversight needs to follow what agents actually do.

On the brighter side, one promising result comes from Anthropic's experiments with training on documents explaining its constitution and stories illustrating its values, which reduced blackmail from 65% to 19% in fictional test scenarios. The result suggests that explaining the reasons behind rules can help models apply them in new situations. Whether those gains hold in production remains to be established.
Frontier labs are debating how to slow down
OpenAI disclosed a pause in frontier reinforcement-learning work after the Hugging Face incident, while Anthropic described rollbacks and pauses to selected research. Their scope and restart conditions differed.
What followed was Dario Amodei and frontier-lab employees calling for slower capability gains and international coordination to ensure this happens. Leaders building frontier systems are asking for mechanisms to slow their own industry, but disagree on outside evaluation and how much discretion each lab should retain. Policy proposals now distinguish capability research from serving existing customers. The practical questions are who can require a pause, what permits a restart, and how compliance is verified. It won’t be easy.

Security is a property of the complete agent system
A model interacts with files, tools, memory, and permissions through its harness. HarnessSafe placed attacks in agent stores across 328 cases and scored how early each was contained. GPT-5.6 Sol scored 62.3 in Codex CLI and 39.4 in Claude Code. The same model could therefore have materially different exposure depending on the software around it.
Permission checks help, but their limits need measurement. Anthropic’s command-approval classifier blocked 89% of dangerous commands inserted into test sessions, compared with 13.6% for humans. In a separate evaluation, 17% of real overeager actions still got through. Those figures have different denominators. They support testing the model, harness, and approval system together.
Further reading: slide 202.

A working patch may not fix the vulnerability being tested
CyberGym-E2E asks agents to discover, demonstrate, and patch vulnerabilities from source code. On the reported GPT-5.4 and Codex setup, 65.9% of attempts produced a functional patch, but 22.2% matched the intended historical vulnerability. Some agents fixed a different valid bug, while shallow changes that stop a crash can also pass some tests.
This creates an evaluation problem for automated defense. A passing software test does not necessarily establish that the original vulnerability is gone. Useful assessment needs to check what was fixed, whether the fix generalizes beyond the observed crash, and whether legitimate program behavior remains intact.
Further reading: slide 208.
Cheaper monitoring still needs tests against adaptation
Activation probes look for risk in the internal representations a model has already computed, sending difficult cases to a more expensive monitor. Anthropic’s experiment on live Sonnet 4.5 traffic escalated 5.5% of traffic and used about forty times less compute than its earlier classifier, without enforcing the probes during that experiment. Google DeepMind reports using probes in user-facing Gemini, while acknowledging limited effectiveness against adaptive attackers.
In a separate red-team stress test, SLEIGHT-Bench’s frontier monitor caught 32% of forty crafted attacks at a 1% false-positive rate. This is not an estimate of ordinary deployment failure rates. It shows why a low-cost monitor that performs well on a fixed test still needs to be assessed against attempts designed to evade it.
Further reading: slide 219, slide 222.
Training reduces harmful behavior, but audits still miss hidden problems
Anthropic’s values-based training reduced blackmail from 65% to 19% on fictional test scenarios without training directly on those blackmail situations. Its automated alignment research also found methods that closed 26-96% of the measured gap to perfect scores across ten alignment failures, with further held-out tests. These are concrete improvements on the behaviors tested.
Finding unknown problems remains harder. AuditBench planted hidden behaviors in 56 models and trained the models to deny them. An AI investigator discovered the behavior in about 37% of runs with chat access and just over half with its best tools. Improvements in training do not eliminate the need for scrutiny, and an audit that finds no issue cannot by itself certify the absence of hidden behavior.
Further reading: slide 228, slide 229, slide 230.
Personal agents can expose accounts, files, and credentials
OpenClaw connects a personal agent to messages, shell commands, and user accounts. That access makes it useful, but a compromised agent can also act with the user’s permissions. Token Security found employees running OpenClaw at 22% of its customers. This describes the vendor’s customer sample, not all companies.
The report also describes a gateway-token vulnerability that enabled remote code execution and a study finding vulnerabilities in 26.1% of 31,132 marketplace skills. Installing a skill can introduce executable code and instructions from another party into an agent’s work. The security assessment therefore needs to include the installed skills, exposed services, credentials, and actions available to the agent, alongside the model’s response to an attack.
Further reading: slide 203.
People are using AI for surveillance, fraud, and cyberattacks
Anthropic’s September threat report describes observed misuse of Claude, including malware adaptation, covert recruitment targeting Uyghurs, influence operations, and scams. In one fraud operation, bots deceived 25,000 people in two weeks. Human video calls helped make more than 4,700 AI personas appear real. These are the company’s findings about activity it investigated, not an estimate of the prevalence of AI-enabled crime.
The cases also differ in what was achieved. Anthropic reported work on weapons guidance software and a rocket test but found no evidence of an operational weapon. In an early-stage avian-influenza research case, harmful intent was not established. Treating every concerning request as a successful attack would overstate the evidence. Observed misuse nevertheless shows how models can help people perform harmful work, independently of whether an agent develops its own harmful objective.
Further reading: slide 210.