Detecting and countering misuse of AI:
September 2026 [pdf]
Overview
Amthropic published its September 2026 report entitled Detecting and countering misuse of AI: September 2026 [pdf]. The report covers threat activity Anthropic disrupted between December 2025 and August 2026, across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation. Cases involved Claude Haiku, Sonnet, and Opus, with essentially no misuse found on Fable or Mythos-class models. Threat actors ranged from suspected state-sponsored groups and financially motivated criminals to commercial spyware vendors, state propaganda institutions, and politically motivated individuals.
The report says:
“Many commentators focus on the risk of AI developing exploits at scale. While this is a danger, the risk from AI adoption is more pronounced across the cyber kill chain, where adversaries can operate faster, across a broader and deeper surface area, with fewer
resources. The cases span the period from December 2025 through August 2026. In all cases, Claude Haiku, Sonnet, and Opus models were used; no malicious activity was found on Claude Fable or Mythos (which has a series of safeguards in place that greatly reduce its ability to perform harmful cyber tasks). In each case we disrupted the activity involved, strengthened our AI safeguards based on what we learned, and shared intelligence with authorities and industry partners where appropriate.”
Cyber operations: key theme — “from assistant to orchestrator”
A central finding: AI has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators, so technical sophistication is no longer a reliable signal of who’s behind an attack. Notable cases:
- GTG-20006 (Russian espionage, linked to Midnight Blizzard): Ran AI-driven workflows covering reconnaissance, phishing, malware rebuilding to evade detection, and data exfiltration against 20+ organizations — Ukrainian/European government, defense, and diplomatic targets, plus drone-technology supply chains and hijacked hotel WiFi to deliver malware to guests’ devices.
- GTG-50014 (ShinyHunters-linked actors): Multiple financially motivated affiliates used AI to mass-harvest credentials from Android apps and GitHub, then conducted rapid breach-to-extortion campaigns — one escalated from a stolen token to full admin control in about three hours.
- GTG-10007 (Chinese-speaking operators): Ran autonomous “exploit foundries,” using AI to reverse-engineer security appliances and find zero-days, alongside espionage and reconnaissance against ~50 organizations.
- GTG-50029 (French hacktivist): A single individual built a doxxing platform and compromised European political parties and media using AI-assisted exploit development.
- AI supply chain itself is now a target: stolen API keys/session tokens are harvested, resold, and used as both loot and attack compute (e.g., GTG-50020, GTG-50021).
Influence operations
Nine cases spanning Russia, Iran, Turkey, the Gulf, South Asia, Africa, and Europe, several timed to elections. Trends include influence sold as a paid service, AI acting as a “newsdesk” sub-editor inside existing propaganda pipelines, attribution laundering (stripping state sourcing to look independent), and fake personas/impersonation. Highlighted cases: a Wagner-linked FIMI operation in the Central African Republic (GTG-04001), a six-continent commercial fake-news network (GTG-54002), a Malaysia-targeting election-manipulation platform (GTG-84005), Russian state-media editorial pipelines (GTG-24015), and Iranian state-aligned “cognitive warfare” operations (GTG-34001). Notably, most operations were caught early and failed to reach authentic audiences since Anthropic can see them at the production stage.
Surveillance operations
Cases involved commercial spyware vendors and state-linked actors using Claude to build monitoring and identification systems, including tools aimed at identifying and tracking dissidents and opposition figures (this built on the surveillance-database elements already noted in the CAR and Malaysia influence cases above). This was one of the largest sections of the report by page count (roughly pages 81–111 of the PDF).
Conventional weapons development
The report includes a case involving China’s PLA Navy (PLAN), where Claude was reportedly used to assist research related to anti-torpedo defense systems — an example of dual-use military engineering support rather than direct weapons-building instructions, since Claude’s safeguards are designed to block direct weapon-construction assistance.
Biological misuse
Anthropic’s biological-safety systems successfully blocked direct bioweapons construction prompts. However, researchers in unsupported regions routed dual-use scientific queries — including research involving pathogens and immune evasion — through proxy networks to reach the models indirectly, which the report flags as a visibility gap rather than a successful attack.
Scams and fraud
A notable case was a China-based dating-app fraud operation that deployed more than 4,700 automated Claude-driven personas across over 20 apps, exchanging 2.36 million messages with 25,000 users over two weeks, with real gig workers blended in alongside the AI personas to keep conversations convincing.
Illicit distillation
This section describes attempts by rival AI labs to harvest Claude’s outputs (especially reasoning traces) to train competing models:
- GTG-16005 (attributed to Alibaba/Qwen/Tongyi Lab) is described as the largest distillation attack Anthropic has measured — targeting chain-of-thought transcripts from Opus 4.6/4.7 via fixed prompts forcing inline reasoning, feeding SFT data for Qwen 3.5–3.7. At its peak this reached nearly 3 million exchanges/day from 3,500+ fraudulent accounts, totaling over 151 million exchanges between May–July 2026.
- The report also describes a broader pattern of systematic attempts by several AI labs to harvest information from Claude through large-scale proxying and replay of interactions, involving millions of exchanges and in some cases exposing sensitive information in the processed data.
- This was the one area where a Fable/Mythos-class model was reportedly implicated in misuse.
Overall throughline
Across all seven areas, the report’s core argument is that AI is shifting from a conversational assistant into an execution and orchestration layer — lowering the skill and resource threshold for sophisticated operations, compressing attacker timelines from days to hours, and making the AI supply chain itself (API keys, session tokens, pre-release model access) a prime target, loot source, and attack-compute resource. Anthropic frames its response as banning accounts, hardening safeguards, and sharing intelligence with governments and industry partners.
