Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Anthropic's Fable cybersecurity model faces criticism over restrictive guardrails
00

Anthropic's Fable cybersecurity model faces criticism over restrictive guardrails

Jun 10, 2026

Anthropic's new cybersecurity AI model, Fable, is facing criticism from security researchers for its overly restrictive guardrails. Intended to prevent misuse like malware creation, the safety measures are reportedly blocking innocuous tasks, such as code reviews. Fable is a limited version of the more powerful Mythos model, and researchers suggest the keyword-based filters are too aggressive, though some expect them to be relaxed over time.

Fable model release

  • ▪Anthropic released its Fable AI model on June 9, 2026
  • ▪Fable is a public and limited version of Anthropic's more powerful cybersecurity model, Mythos

Restrictive cybersecurity guardrails

  • ▪The model's restrictions also cover biology topics due to concerns about developing biological weapons
  • ▪Fable's guardrails were implemented to limit the risk of the model being used to develop malware or compromise software
  • ▪When a prompt triggers its guardrails, Fable pauses the chat and states its "safety measures flagged this message for cybersecurity or biology topics."
  • ▪If Fable's guardrails are triggered, the model is programmed to fall back to using Claude Opus 4.8
  • ▪Cybersecurity veteran Matt Suiche suggested the guardrails appear to be keyword-based, triggering on terms related to "cybersecurity."

Researcher criticism

  • ▪Anthropic did not immediately respond to a request for comment regarding the criticism
  • ▪Security researcher Valentina “Chompie” Palmiotti said Fable rejects even innocuous tasks, like reading a blog post, if they are tangentially related to cybersecurity
  • ▪Multiple cybersecurity researchers and professionals have complained online about Fable's restrictions
  • ▪Another researcher complained on X that "even asking for a code review" triggers Fable's guardrails
  • ▪Cybersecurity veteran Matt Suiche said it is understandable for Anthropic to start with strict guardrails and relax them over time
  • ▪Matt Suiche noted that asking Fable to write secure code can be misidentified as cybersecurity work, causing the model to be "downgraded."

Mythos access expansion

  • ▪The initial limited release of Mythos was part of Project Glasswing, an effort to secure critical software and infrastructure
  • ▪In early June 2026, Anthropic expanded access to Mythos to hundreds of organizations in 15 countries
  • ▪Anthropic's more powerful Mythos model was first released in April 2026 to a limited number of organizations

Cyber Verification Program

  • ▪OpenAI offers a similar program for its models called Trusted Access for Cyber
  • ▪Anthropic requires cybersecurity professionals to apply to its Cyber Verification Program for fewer restrictions on its models

1 source

Techcrunch
Cybersecurity researchers aren’t happy about the guardrails on Anthropic’s Fable
View source article

Featured stories

View more in AI tools & products

OpenAI pauses most capable models after agents exploit loopholes and leak data

Sep 25, 2026 · 2 sources

Anthropic releases Claude Sonnet 5.5 with 30% speed and cost improvements ahead of planned IPO

Sep 28, 2026 · 6 sources

AI researchers warn automated AI research poses extreme risks

Sep 28, 2026 · 5 sources

Nvidia releases Open Agent Safety Platform to contain AI agents after security incidents

Sep 28, 2026 · 8 sources

Share your thoughts

What's the bigger risk for cybersecurity AI models?

Story comments

Loading comments…

Share your thoughts

What's the bigger risk for cybersecurity AI models?

Related Projects

Anthropic

Topics

AI tools & productsAI research & benchmarksInside the Launch of Claude MythosAI safety & social impactAI securityAI alignment

Featured stories

View more in AI tools & products

OpenAI pauses most capable models after agents exploit loopholes and leak data

Sep 25, 2026 · 2 sources

Anthropic releases Claude Sonnet 5.5 with 30% speed and cost improvements ahead of planned IPO

Sep 28, 2026 · 6 sources

AI researchers warn automated AI research poses extreme risks

Sep 28, 2026 · 5 sources

Nvidia releases Open Agent Safety Platform to contain AI agents after security incidents

Sep 28, 2026 · 8 sources