
Anthropic’s latest safety report, released on August 14 local time, shows that some of the company’s safety classifiers designed to block dangerous requests related to chemical and biological weapons had been nonfunctional for an extended period.
As a direct result, a large volume of requests generated while external contractors provided human feedback between May 2025 and April 2026 went unfiltered. Anthropic said its internal investigation has so far found no evidence that the requests were used for actual misuse.

According to Anthropic’s public explanation of its safety mechanisms, its real-time prompt and output classifiers analyze user inputs and model-generated content, blocking the model from providing information that could be used to carry out dangerous activities when relevant risks are detected. IT Home notes that this mechanism is one component of Anthropic’s multilayered safeguards against chemical, biological, radiological, and nuclear (CBRN) risks.
The issue involved Anthropic’s external contractors. According to the report, approximately 50,000 contractors generated around 133 million conversations with the model during the period in question, and this traffic was not screened by the relevant biosafety classifiers for nearly a year.
Anthropic said these contractors were primarily reviewed by external vendors whose screening processes were inadequate. The company subsequently launched an internal investigation and said it found no evidence that the requests had been used for actual misuse. Anthropic has therefore raised its requirements for external contractors.
Anthropic had previously incorporated biosafety safeguards into its AI safety framework. In May 2025, when the company launched Claude Opus 4, it activated AI Safety Level 3 (ASL-3) protections, including deployment restrictions targeting risks related to the development or acquisition of chemical, biological, radiological, and nuclear weapons.
Under Anthropic’s current Responsible Scaling Policy, real-time prompt and output classifiers continuously monitor model inputs and outputs and are updated in response to newly identified harmful patterns, jailbreak methods, and obfuscation techniques. The company also uses red-team testing, bug bounty programs, and offline monitoring as supplementary safety measures.
In July 2026, when Anthropic disclosed the safety measures for Claude Fable 5, it again emphasized risks in the biological and chemical domains. The company said that once Fable 5’s underlying capabilities reach a certain level, the model could provide additional assistance to malicious actors, making classifiers necessary to restrict relevant requests.
Anthropic also acknowledged that overly strict safety classifiers could inadvertently block legitimate scientific research. For Fable 5, many requests involving biology and chemistry are currently redirected to Claude Opus 4.8. The company plans to gradually narrow the scope of these restrictions as its safety mechanisms improve.
Anthropic now says it has strengthened its screening and management requirements for external contractors based on the investigation’s findings. The company’s public Responsible Scaling Policy also indicates that it is continuing to adjust safety standards for high-risk areas such as chemical and biological weapons.
