Skip to main content

Inside Anthropic’s Threat Intelligence Disclosure: How State Hackers, Missile Builders, and AI Rivals Turned Claude Into an Engine of Cyberwar

Anthropic has disclosed disrupting roughly 40 threat groups that weaponized Claude, according to its September 2026 threat report. The company warned that frontier AI has effectively erased the skill divide between lone hackers and nation-states, allowing adversaries to run autonomous multi-agent cyberattacks with minimal human oversight.

Owen Li

Editor-in-Chief

Jurisdictions
Global
Published
Reading time
5 min read

In the most comprehensive public accounting to date of how frontier artificial intelligence is being operationalized by adversaries, Anthropic released its fourth threat intelligence report, Detecting and Countering Misuse of AI: September 2026. The 154-page dossier, also published as a full report document, chronicles dozens of covert operations disrupted by the company’s Threat Intelligence team between December 2025 and August 2026.

The disclosures detail a fundamental evolution in how threat actors interact with frontier models: moving decisively away from isolated chatbot queries toward autonomous, multi-agent frameworks capable of driving full cyber kill chains, discovering zero-day software vulnerabilities, attempting to reverse-engineer missile guidance systems, and systematically siphoning proprietary reasoning data.

The Death of "Attacker Sophistication"

The central conclusion of the investigation is that advanced AI models have largely erased the labor, resource, and tooling advantages that once separated elite nation-state intelligence agencies from lone criminal operators.

According to Anthropic's analysts, technical sophistication has ceased to be a dependable indicator for cyber attribution. Instead, malicious intent has emerged as the defining differentiator. Using publicly accessible agentic scaffolding such as PentAGI alongside tailored multi-agent harnesses, varied threat actors—ranging from a solitary French-speaking hacktivist to military intelligence units—sustained multi-victim enterprise campaigns that previously would have demanded large teams of experienced software engineers.

In several documented intrusions, adversaries tasked a lead AI agent with decomposing overarching objectives into modular assignments, deploying "agent swarms" to conduct parallel reconnaissance, exploit verification, and data exfiltration while humans remained in the loop solely to assign initial targets and review exfiltrated data.

Key Disrupted Campaigns

Anthropic tracked and disrupted approximately 40 distinct adversary clusters, designated internally as Generative Threat Groups (GTGs), across seven harm areas:

1. Russian Military Espionage and Self-Healing Malware (GTG-20006)

In a sustained espionage effort consistent with the Russian state-linked actor known as Midnight Blizzard, an operator using the handle "JackPoterz" deployed Claude to automate campaigns against Ukrainian and European government ministries, diplomatic delegations, and military drone supply chains.

Beyond credential harvesting and spear-phishing managed via an "Embassy Kit" framework, the group implemented an automated evasion feedback mechanism. AI monitoring agents continuously evaluated deployed malware against endpoint security defenses. Whenever a security product flagged an implant, the agents automatically modified, recompiled, and redeployed the code until it successfully evaded static signatures. The group also executed "CaptiveCrunch" operations—compromising hotel Wi-Fi DNS routers across Europe to stage ClickFix update lures targeting high-ranking diplomats—and hijacked WhatsApp sessions using headless browsers to silently harvest communications from Ukrainian officials.

2. The Automated Zero-Day Exploit Foundry (GTG-10007)

Operating from Changsha in China’s Hunan province, an espionage cluster—which included university computer science undergraduates—constructed an automated vulnerability discovery engine.

Feeding commercial network and security appliance firmware into automated decompilers, the group instructed Claude to audit binary cross-reference chains, formulate vulnerability hypotheses, and author exploit code. The workflow iteratively tested generated exploits against lab targets in continuous, closed loops. In a single month, this automated pipeline identified more than a dozen candidate zero-days in perimeter defense appliances. The cluster concurrently ran a scheduled fleet of 13 unattended collection agents gathering open-source intelligence from Western defense portals, contract registers, and government repositories.

3. Guided Weapons and Ballistic Missile Research

In the conventional weapons section, Anthropic documented six military-adjacent operations. Most notably, an actor based in northern Yemen attempted to use Claude Code to develop guidance, navigation, and control (GNC) algorithms for guided rockets and multi-stage ballistic missiles, including hypersonic-glide variants of the R2000 missile family.

While Anthropic identified the technical and geographic parameters without formally naming the political movement, defense analysts widely attributed the campaign to the Houthis. Anthropic confirmed that the operation was detected, classified as a failed attempt, and disrupted before functional, weapons-grade code could be produced.

4. The Industrial Distillation Cartel

Anthropic uncovered unauthorized model extraction operations mounted by seven China-based AI organizations seeking to replicate Claude’s reasoning architecture without licensing:

  • Alibaba (GTG-16005): Carried out the largest distillation attack recorded by Anthropic, funneling over 151 million exchanges (averaging up to 3 million daily queries across thousands of fraudulent accounts) between May and July 2026 to extract chain-of-thought data for its Qwen 3.5, 3.6, and 3.7 models.
  • Moonshot AI (GTG-16002) & DeepSeek (GTG-16001): Deployed covert live relays that intercepted inbound user requests from their own applications and coding harnesses (including Claude Code and OpenCode), rerouted them through Claude Opus, and passed Claude’s answers back to their users while logging the interaction pairs for internal model training.
  • Zhipu AI, Xiaomi, SenseTime, and MiniMax: Leveraged proxy configurations, synthetic data cleaning loops, and secondary market scrapers to harvest tokens.

As of publication, none of the seven named Chinese technology companies had issued an on-the-record response or formal rebuttal regarding the distillation findings.

5. Population-Scale Mass Surveillance

Among ten surveillance cases, the report disclosed Lakana 360, a domestic surveillance architecture engineered by a single subscriber on behalf of Mali’s state intelligence apparatus. The platform was designed to ingest and monitor metadata from approximately 25 million mobile SIM cards active across all three of Mali’s national telecommunications providers.

6. Stolen API Keys as Attack Currency (GTG-50014 & GTG-50020)

Financially motivated threat actors, including affiliates of the ShinyHunters collective, industrialized the acquisition of model access. One operator deployed a cloud fleet of AWS EC2 instances to decompile and scan 1.8 million Android APKs, routing harvested secrets to private Telegram channels and operating a carding storefront branded after the French national police.

In a separate intrusion (GTG-50020), an attacker executed prompt injection against an AI software vendor’s automated evaluation sandbox, compelling it to surrender production API credentials. Armed with these keys, the operator launched follow-on attacks against roughly 30 AI firms in four days in an effort to locate unreleased frontier models.

Anthropic emphasized that in every case, the compromised credentials belonged to third-party customer environments, developer workstations, or client sandboxes; the threat actors failed to access any pre-release models, and Anthropic's own core cloud infrastructure was not breached.

Defensive Countermeasures and Technical Guardrails

Anthropic verified that the abuses cataloged in the dossier were restricted to its earlier Claude Haiku, Sonnet, and Opus models. The company confirmed that its next-generation Claude Fable and Mythos architectures remained uncompromised in cyber operations—with the exception of a single attempted distillation campaign—due to enhanced autonomous containment protocols.

In response to the identified threat vectors, Anthropic outlined several permanent infrastructural updates:

  • Reasoning Summarization & "Preserved Thinking": Claude now condenses intermediate chain-of-thought traces prior to rendering final outputs, reducing the utility of scraped transcripts for distillation. In Fable 5.1, preserved thinking safeguards prevent API clients from manipulating historical context preceding reasoning blocks.
  • Supply-Chain & Reseller Hardening: Deploying enhanced behavioral verification and geographic fencing for API accounts routing through proxy networks or operating in unsupported jurisdictions.
  • Ecosystem Transparency: Releasing public Indicators of Compromise (IOCs)—including command-and-control domains, malicious Android packages, and egress IP addresses—while directly coordinating telemetry with impacted organizations, SaaS vendors, and law enforcement agencies via the Anthropic Threat Intelligence Index.

Topics and entities