Anthropic disclosed Thursday that it has intercepted efforts by malicious actors to exploit its artificial intelligence models for harmful purposes, including cyberattacks, surveillance, and research that could have facilitated the development of biological weapons.

The company, which is preparing for an initial public offering this fall, said that as AI models become more powerful, even individuals with limited technical expertise can now orchestrate sophisticated threats that were unimaginable just a year ago. In response, Anthropic has implemented stricter safeguards in its latest models to curtail dual-use biological research that could be diverted toward weaponization.

Read also
Technology
Anthropic staff fear AI apocalypse; GOP warns Democrats risk U.S. lead
Anthropic insiders publicly warn AI could kill all humans, prompting GOP calls for guardrails while criticizing Democratic proposals to halt AI development.

“The cases we share here aren’t typical misuse, but rather examples of the most notable and novel threat activity we’ve identified to date,” the company said in its third report on AI misuse since March 2025. The report, which includes excerpts of malicious code and AI prompts, is intended to help governments and competitors identify and prevent similar abuses.

“We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” Anthropic stated.

The report arrives amid growing concerns about AI safety. Just two days before its release, one of Anthropic's own researchers, Jacob Coxon, announced his resignation, warning that the company and its chief rival OpenAI are “racing straight to self-improving superintelligence and gambling with our lives.” Coxon's departure echoes broader fears, both inside and outside the industry, that AI could slip beyond human control.

Attempts to enhance a virus

Between December 2025 and August 2026, Anthropic researchers identified misuse ranging from spyware vendors and politically motivated individuals to state-sponsored propaganda campaigns. Among the most alarming findings was an attempt to use Claude, Anthropic's AI model, to assist in drafting a grant application for scientific funding.

The application involved gain-of-function research on the chikungunya virus, a mosquito-borne pathogen that causes severe fever and joint pain. The proposed research aimed to enhance the virus's transmissibility and immune evasion properties, potentially making it more harmful. While such research could aid in developing vaccines and treatments, Anthropic noted it “could also be used to make the pathogen more dangerous.”

Anthropic said it cannot guarantee that its current models are harmless. The report indicates that none of the identified cases involved its newest, most powerful models—Claude Fable or Mythos—except for one instance of “illicit distillation,” described as an industrial-scale effort to extract a model's capabilities without authorization.

Older models like Claude Opus 4 and Claude Sonnet 4.5, from 2025, were deemed “well below the threshold where they could meaningfully assist a sophisticated user in carrying out dangerous biological research.” Consequently, safeguards on those models were less stringent, focusing mainly on preventing novices from recreating known bioweapons. However, for today's more advanced models, “the evidence is no longer certain,” and Anthropic has applied stricter restrictions on dual-use biological research queries in its newer systems, such as Claude Fable 5.

Call for government oversight

Experts have increasingly urged governments to regulate AI rather than rely on industry self-policing. John Thickstun, an assistant professor of computer science at Cornell University, highlighted the uncomfortable position companies like Anthropic and OpenAI face when they must make “value judgments at societal scale without any kind of democratic or deliberative oversight.”

The report also details influence operations, including groups creating hundreds of fake social media accounts to amplify political messages. Anthropic identified nine such cases originating from Russia, Iran, Turkey, the Persian Gulf, South Asia, Africa, and Europe. While social media platforms can detect these operations after they go live, Anthropic notes it may spot them on Claude while the operation is still being built.

Anthropic emphasized that it has blocked all identified malicious activities, used the incidents to strengthen its safeguards, and shared information with government authorities and industry partners. “We hope that the findings in this report will help other developers recognize similar patterns on their own platforms, give governments and civil society a clearer picture of the threats,” the company said. The report follows a staffer's dire warning about AI risks and comes amid concerns about AI model security.