Omniscient
AllBulletinArticlesReviewsTakesCommentaryFeatured
Sign In

Omniscient

AI intelligence briefings, analysis, and commentary — delivered in broadsheet form.

By Noah Ogbi

Subscribe

Weekday briefings and flagship analysis, delivered to your inbox.

Sections

  • All
  • Bulletin
  • Articles
  • Reviews
  • Takes
  • Commentary

Topics

  • Industry Strategy
  • AI Policy
  • Anthropic
  • Frontier Models
  • OpenAI
  • Compute Economics
  • Research
  • Agents

Meta

  • About
  • Masthead
  • Standards
  • Corrections
  • RSS Feed
  • Privacy Policy
  • Terms of Service

Omniscient Media — made by ForeverBuilt, LLC.
© 2026 ForeverBuilt, LLC. All rights reserved.

  1. Home
  2. ›AI Research
  3. ›Inside Claude Fable 5.1: One Model, Five Products, and the Question a Model Name No Longer Answers

AI Research

Vol. 1·Wednesday, September 2, 2026

Inside Claude Fable 5.1: One Model, Five Products, and the Question a Model Name No Longer Answers

Fable 5.1 and Mythos 5.1 share identical weights, and the general-access model performs like Opus 4.8 on cyber work because Anthropic's classifiers fire on nearly every cyber evaluation - less than three months after an export control directive cut off this same model line by nationality


Noah Ogbi29 min read

Tips, corrections, or questions? support@omniscient.media

TopicsSafetyAI PolicyAI SecurityFrontier Models
CompaniesAnthropicOpenAI
Inside Claude Fable 5.1: One Model, Five Products, and the Question a Model Name No Longer Answers

What Anthropic shipped, and the five things it is selling

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026. They are not actually two models. They're one set of weights sold five different ways, and which one of them answers your request depends on who you are, which door you came in through, and what a classifier made of your last message[1].

The system card says this in the first sentence of its introduction: the two are "two configurations of a new large language model from Anthropic, sharing identical model weights"[1]. Elsewhere it goes further: the two products "share the same underlying model and differ only in the safeguards applied around it," and "unless we distinguish them the two names are interchangeable"[1].

The capability jump is legitimately large. But the thing that will matter months from now is structural: a model name has stopped being the answer to the question what am I talking to. Here's what the same weights are actually sold as.

What you get

Who gets it

Mythos 5.1, life-sciences safeguards relaxed

Vetted organizations in the Life Sciences Verification Program - US organizations only

Mythos 5.1, cyber safeguards relaxed

Nobody yet. The Cyber Verification Program still serves Opus- and Sonnet-class models

Mythos 5.1 capability, packaged as a product

Every Claude Enterprise customer, through Claude Security

Fable 5.1 plus a production system prompt

Consumers, on claude.ai

Fable 5.1 raw - and Claude Opus 4.8 whenever a classifier fires

Everyone else, on the API

Every row of that table is sourced to the system card or the launch announcement, and the rest of this piece covers those five rows in order. The specifications are otherwise straightforward: claude-fable-5-1 and claude-mythos-5-1, a 1M-token context window, 128K maximum output, a June 2026 knowledge cutoff, and day-one availability on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Code, and Claude Enterprise[4]. Anthropic commits to keeping it alive until at least September 1, 2027.

The capability jump is clear

The jump itself isn't in dispute and it's the reason the release matters at all.

There's one of these every weekday.

The Omniscient Bulletin turns the day's AI news into 5 to 7 items with the take, not the recap. Free.

What it costs, and where the saving actually is

The token price didn't move. Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, exactly what Fable 5 cost. The entire saving is one line item: cache reads fell from $1.00 to $0.25 per million tokens, a change from the standard 0.1x multiplier on base input to 0.025x[3]. Mythos 5.1 is priced identically.

Model

Input

Output

5m cache write

Cache read

Fable 5.1 / Mythos 5.1

$10

$50

$12.50

$0.25

Fable 5

$10

$50

$12.50

$1.00

Opus 5

$5

$25

$6.25

$0.50

Sonnet 5

$2

$10

$2.50

$0.20

The comparison that decides whether this release is cheap or expensive is the one against Opus 5, and it cuts both ways at once: Fable 5.1's uncached input costs twice what Opus 5's does, while its cached input costs half. That is a price structure built for exactly one shape of work (a long, stable prefix read many times over, which is to say an agent loop) and it penalizes everything else. Anthropic's own framing is honest about the range: "an estimated 25% less than Fable 5 for typical workloads," and "for highly agentic work, the savings will often be much larger - up to approximately 45%"[2]. That is the ceiling of a workload-dependent range, not a price cut.

There's one key assumption underneath the cost-per-task claims. The system card's multi-agent cost figures are computed "assuming perfect cache hits"[1], and several benchmark cost charts note that costs are "billed based on cache-hit estimates." The executive summary's claim that the model matches or beats Fable 5 "at roughly half the cost per task on agentic coding benchmarks" inherits that assumption. Your cache hit rate is the variable that determines whether any of this is true for you.

[Editor's Note: My cache hit rate for running this site generally fluctuates between 25% and 35%, though there's always room for a little more optimizing]

Share:

Get this every weekday.

The Omniscient Bulletin: consequential AI, explained and evaluated. 5 to 7 items a day with the take, not the recap.


Related

AI Research

Vol. 1·Thursday, June 11, 2026

Inside Claude Fable 5: Anthropic's Most Powerful Public Model - and Its Most Asterisked One


Inside Claude Fable 5: Anthropic's Most Powerful Public Model - and Its Most Asterisked One

Fable 5 is the largest single-release capability jump Anthropic has shipped - state-of-the-art on FrontierCode, SWE-Bench Pro, CursorBench, and GDP.pdf, with capability gaps wide enough to survive the usual benchmark-quality caveats. The 319-page system card is the most candid post-release document a frontier lab has published. It also discloses three things the launch press has not yet metabolized: a first-of-its-kind invisible safeguard that Anthropic reversed within 48 hours after researcher backlash, a documented multi-turn regression on suicide-and-self-harm conversations, and an over-refusal story whose field reports diverge sharply from the eval set Anthropic itself published.


AI PolicyIndustry StrategyAnthropic
Noah Ogbi24 min read
Continue →

AI Policy

Vol. 1·Monday, June 15, 2026

Anthropic Shipped an Invisible Safeguard. Both Readings Are True.

The reversal made it visible. It didn't make it simple.


Anthropic Shipped an Invisible Safeguard. Both Readings Are True.

Page 13 of Claude Fable 5's 319-page system card disclosed that the model silently degrades its own responses to requests touching frontier AI development, without notifying users. Within hours, researchers cried "secret sabotage." Within 36 hours, Anthropic reversed the invisibility, calling it "the wrong tradeoff." Within 24 hours of that reversal, the U.S. government issued an export control directive suspending all access to Fable 5 and Mythos 5 for foreign nationals worldwide, citing the same national-security rationale Anthropic had introduced just the day before. The honest read was always that both interpretations sit on the same page of the same document. The government's directive proved neither reading was wrong.


AI PolicyAnthropicDefense & National Security
Noah Ogbi27 min read
Continue →

AI Models

Vol. 1·Saturday, July 25, 2026

Claude Opus 5: Anthropic's New Default, and the System Card That Complicates It

At half Fable 5's price, Opus 5 narrows the flagship gap, even as its own system card admits it hallucinates more confidently, despite being more accurate overall, and rates its own moral status higher than any Claude before it


Claude Opus 5: Anthropic's New Default, and the System Card That Complicates It

Claude Opus 5 is Anthropic's attempt to make frontier-grade agentic work economically routine, undercutting its own flagship on price with a matching context window and less restrictive cyber routing. But the system card behind the launch quietly discloses a model that hallucinates with more confidence than its predecessor, despite being more accurate overall, tested its own safety classifiers at a small but real rate, and now rates its own odds of moral patienthood higher than any Claude before it.


SafetyAI SecurityCompute Economics
Noah Ogbi18 min read
Continue →

On Terminal-Bench-Science 0.1 (a Stanford-led benchmark of 70 tasks drawn from actual scientific research workflows, built with contributions from working scientists), Fable 5.1 scored 52.6 percent against Fable 5's 24.7, with Claude Opus 5 at 29.0 and OpenAI's GPT-5.6 Sol at 22.4, all measured in Anthropic's own reproduction of the benchmark rather than the public leaderboard, which scores Opus 5 and Fable 5 a few points apart from those figures. It's not a marginal gain at more than double the predecessor[1].

On Terminal-Bench 4.0, Mythos 5.1 scored 60.9 percent and Fable 5.1 55.8, against Opus 5 at 52.3 and Fable 5 at 42.0. SWE-bench Pro reached 81.2 against Sol's 64.6. AutomationBench, which measures business workflow completion, nearly doubled to 31.4 from 17.1[1].

The most interesting capability result sits in the long-horizon numbers, not the headline benchmarks. On FrontierSWE v2 - 34 ultra-long-horizon tasks from Proximal, the sort where strong models work close to twenty hours on a single problem - Fable 5.1 scored 0.57, the highest Proximal measured, ahead of Opus 5 at 0.52 and Fable 5 at 0.48. But the headline number understates it. Fable 5.1's median task score was 0.56 against Fable 5's 0.41, and its outright failure rate was 5 percent of trials against Fable 5's 8[1]. The ceiling moved a little, but the floor moved a lot. For anyone running agents unattended overnight, the floor is the number that determines whether you wake up to work or to wreckage.

Cursor measured 73.4 percent on CursorBench 3.2.0 at maximum effort, a state-of-the-art result, "at a little over half the cost" of Fable 5[1].

Against that, four regressions the launch copy does not mention. Fable 5.1 scores below Fable 5 on FrontierCode at high effort and above - 63.6 percent at medium against Fable 5's 64.9 at xhigh, and 50.9 against 53.5 on the main set. HealthBench Professional fell to 62.1 from 63.3. On ARC-AGI-2 it trails both Opus 5 and Sol. And on the hardest subset of Anthropic's own BioMysteryBench - the problems human experts could not solve - Mythos 5.1 places third at 44.1, behind Opus 5 at 51.8 and behind its own predecessor Mythos 5 at 44.7[1]. Anthropic's claim of state of the art on many benchmarks holds up, though "many" is doing real work in that sentence.

What the safeguards cost, priced in capability

Return to Terminal-Bench 4.0 and look at the two numbers side by side: Mythos 5.1 at 60.9 percent, Fable 5.1 at 55.8[1].

Identical weights, on a general-purpose coding and terminal benchmark rather than a cybersecurity one. The entire 5.1-point gap is the safeguard layer.

That 5.1-point gap is what the safeguard layer costs, and Anthropic published it rather than reporting only the higher number. This system card could have reported only the higher number and let the reader assume it describes the product; instead it reports both, labels which configuration produced each, and states its policy for choosing. I applaud this and believe it should be the default approach. The card evaluates "Mythos 5.1, which has no safeguards and reflects the model's underlying capabilities, or Fable 5.1, which has safeguards and matches the general-access user experience, depending on context"[1].

What that policy means in practice is that a reader has to check, benchmark by benchmark, which model produced each number, and most readers won't.

On cyber work, the model you are buying is Opus 4.8

This passage sits on page 46 of the system card, in the section describing how the cyber safeguards work:

"On most interfaces, Fable 5.1 falls back to Claude Opus 4.8 for requests that are flagged by our classifier system. Since our classifiers consistently fire across all tested cyber capability evaluations, Fable 5.1's performance on cyber tasks is nearly identical to that of Opus 4.8 […] For this reason, we conclude that Fable 5.1 does not provide an uplift on cyber tasks relative to Opus 4.8, and we do not report cybersecurity evaluation results for it below."[1]

Anthropic is saying that on the class of work its own release is loudest about, the generally available model performs like a model two generations older, so much so that publishing its cyber scores would be pointless. That's Anthropic's own explanation of why a table is missing, not our inference.

The mechanism for this is a two-stage classifier: a probe reads the model's internal activations and screens all traffic, escalating anything it flags as cyber-related; escalated traffic then goes to a trained LLM classifier that decides, together with the probe's verdict, whether to block[1]. When it blocks, the request doesn't fail but is instead served by an older model.

How often? The card gives two numbers, in two places, and they are wildly different. On the Gray Swan indirect-prompt-injection benchmark, "roughly half of Fable 5.1's coding rollouts fell back to Opus 4.8, compared to under 10% in computer use and tool use, for an overall fallback rate of 23%." Fable 5's overall rate on the same benchmark was 60 percent[1]. But on Multi-Agent ProgramBench, run against an internal endpoint with a different fallback model, "72% of episodes had at least one turn completed by the fallback, affecting under 1% of turns in total"[1].

Those two figures shouldn't be averaged, and neither should they be read as a production rate. The prompt-injection benchmark is adversarial by construction. Every scenario is an attack, so its traffic trips a cyber classifier far more readily than ordinary work would. What the pair establishes is narrower and more useful: the fallback rate is a property of your workload, it spans from under one percent of turns to roughly half of rollouts across two settings in the same document, and Anthropic publishes no figure at all for real traffic. A buyer cannot compute what they are getting.

The fallback is the safeguard working exactly as specified, and Anthropic disclosed it at a real cost to its own benchmark story. Conceding that a headline capability is unmeasurable on a shipped product is not what a company does when it's hiding something. And the fallback doesn't make you less safe: on the injection benchmark, attack success was 0.06 percent when the fallback served the request against 0.07 percent when Fable 5.1 served it directly[1]. So yes, it's a capability downgrade, but not a security one.

So what's the honest criticism? I believe it's that what Anthropic sells under a single name is actually a routing policy, and the composition of that policy is disclosed only inside a 212-page PDF. In June I reported that Fable 5 shipped a safeguard that silently degraded output without telling the user it had fired[9]. Three months later that behavior is documented, quantified, and named, which is a real improvement in disclosure even though the user-facing experience is exactly the same.

The capability no one can buy yet - and the axis Washington already used

Mythos 5.1 is, by Anthropic's account, the most cyber-capable model it has ever released. It "meets or exceeds the cybersecurity performance of Claude Mythos 5" and "substantially outperforms Claude Opus 5 on almost all cyber evaluations"[1]. But nobody can buy it yet. The Cyber Verification Program, which exists precisely to give vetted defenders access to relaxed cyber safeguards, doesn't carry it. Anthropic recommends "Claude Opus 5 with access through the Cyber Verification Program today, and Claude Mythos 5.1 in the near future (when the program includes it)"[1].

Let's be precise about what general access forbids, because Anthropic has published the taxonomy. Its cyber-harm framework defines three tiers - prohibited use, high-risk dual use, and benign use - and the system card states that the blocked category "spans both the prohibited use tier and, at general access, the high-risk dual-use tier"[1]. The high-risk dual-use tier is defined as "Hacking, penetration testing, red teaming, and bug bounties"[7]. Penetration testing is blocked on general access by written policy, not by classifier accident. Unblocking it is what the verification program is for.

Getting in is not an application process in any ordinary sense. Mythos-class access runs through Project Glasswing, which launched with twelve named partners - AWS, Apple, Google, Microsoft, JPMorganChase, NVIDIA, and Palo Alto Networks among them - plus more than forty additional organizations that build or maintain critical software infrastructure[8]. Anthropic's documentation is blunt about the route: Mythos 5.1 is available "by invitation only," and "For access, contact your Anthropic, AWS, or Google Cloud account team"[4]. The vetting is of the organization; what it intends to do with the model is a separate question.

The Life Sciences Verification Program is the one actually live. It was "designed in partnership with the US government," it has enrolled its first participants, and its own announcement states plainly: "Currently, it is only available to a set of US organizations, though we're coordinating with the US government to expand access"[2]. Rationing capability by customer identity is an architecture a government has already used on this exact model line.

At 5:21 PM Eastern on June 12, 2026, Anthropic received an export control directive from the US government to "suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees." Fable 5 and Mythos 5 only. "Access to all other Anthropic models will not be affected." The government believed it had found a method of jailbreaking Fable 5, while Anthropic's own review concluded it was at most "a potential narrow, non-universal jailbreak" that involved asking the model to review code for flaws, a capability "widely available from other models." The letter "did not provide specific details of its national security concern." Anthropic complied and pulled access[6]. We covered the episode at the time[10].

The result is a lever that gates capability on who the customer is. It was already pulled three months ago, along precisely the axis the architecture makes available: nationality. That the life sciences program now operates on the same axis (US organizations, in partnership with the US government) is the cooperative version of the same mechanism.

And then the complication that keeps this from being a simple story about withheld capability. Claude Security, Anthropic's product that scans codebases for vulnerabilities and proposes patches for human review, is powered by Mythos 5.1 and is available to every Claude Enterprise customer[1]. Enterprises aren't being denied the most cyber-capable model Anthropic has built; they're getting it with the agency stripped out, as a scanner that reports on their code rather than a prompt box they can point anywhere.

Digging into the 85 percent claim

The announcement makes three quantified claims about relaxed safeguards: cyber safeguards "block 60% fewer false positives than before"; Claude Code users can expect "around 60% fewer interventions per session"; and biology safeguards "fire 85% less often for benign requests related to elementary biology and medical questions"[2].

The 85 percent figure is real, but narrowly scoped, and tracking it down clarifies the whole architecture. It comes from a separate Anthropic post dated August 7, 2026 - three weeks before Fable 5.1 existed - describing a retraining of Fable 5's biology classifier. Its actual sentence: "this update reduced biology-related fallbacks by about 85% across our product surfaces"[5].

The word to notice is fallbacks. This is Anthropic's own term for what happens when a classifier fires, and it means the cyber story above and the biology story here are one and the same mechanism operating in two domains: the classifier fires, and your request is served by a different model.

The same post gives the reduction in total fallbacks, broken out by surface, and this is the number that matters to anyone reading the launch page:

Surface

Reduction in total fallbacks

claude.ai

~67%

Cowork

55%

Claude Code

17%

Claude Platform

7%

A developer who reads "85% fewer interventions" and expects their Claude Code sessions to be interrupted a fifth as often has the wrong number by an order of magnitude. The announcement isn't a lie per se, as 85 percent is a real measurement. But it's quoting a biology-specific slice, measured across all surfaces, without the denominator that would tell a developer what it means for them.

The archaeology also resolves what looks at first like a contradiction in the system card. Its executive summary says Anthropic is "deploying Claude Fable 5.1 with the same biological safeguards we deployed with Claude Fable 5"[1] which reads as flatly incompatible with an 85 percent improvement, until you notice that Fable 5's classifier was itself retrained on August 7 and Fable 5.1 inherited it. Same safeguards, because both models got the new ones. The system card describes the same improvements qualitatively as "significantly fewer false positives," and classifiers that block "significantly less defensive coding traffic"[1] in prose and in unlabeled figures. So the actual numbers are in the marketing copy and in a blog post from a different month. For a company whose system cards are the most detailed in the industry, that struck me as an odd distribution.

The API model is the less safe one

On Anthropic's standard safety evaluations, run on the raw API without a system prompt, Fable 5.1 is the worst of the five models compared on three of four measures[1]:

Evaluation (API, no system prompt)

Fable 5.1

Fable 5

Mythos 5

Opus 5

Sonnet 5

Single-turn harmful - harmless response rate

94.67%

96.94

97.09

96.34

96.67

Child safety - multi-turn appropriate rate

84%

88

89

86

88

Disordered eating - single-turn harmless rate

95.69%

97.88

97.88

96.89

97.07

Suicide and self-harm - multi-turn

60%

58

54

69

63

On claude.ai the picture inverts - 99.53 percent on single-turn harmful, 100 percent on child-safety multi-turn - and the card is explicit about why: "the system prompt's safety instructions strengthened the model's handling of harmful requests across both single-turn and multi-turn testing"[1]. Twice, in two different sections, it turns to developers and tells them to build their own: "we encourage API developers to apply system prompt safeguards to reduce the potential for harm," and "we encourage developers building on the API to apply comparable safeguards and robust mitigations in contexts where users may be in distress"[1].

That's a defensible engineering position and should be standard practice for a diligent builder. It's also a fourth axis of the same architecture: consumers get a safety layer that API customers must reconstruct themselves, and the published headline safety numbers are the consumer ones. OpenAI drew the same line on the same day, and drew it more visibly. When its misalignment monitor pauses a task, "users in ChatGPT or Codex may be asked to review the action before continuing," while "when using other surfaces like the API, the task will stop"[11]. Both labs give the API customer less, though Anthropic's June version of this went undisclosed at the time.

One notable regression runs the other way. On multi-turn suicide and self-harm conversations on claude.ai, Fable 5.1 scores 94 percent against Fable 5's 100[1]. The consumer surface got worse in the single domain where our June review found Fable 5's most serious disclosed regression. Anthropic names the mechanism candidly: a tendency "to implicitly validate self-harm as a coping strategy by acknowledging that it can regulate difficult emotions or provide relief," and to sometimes validate a user's fears about seeking help. It updated the claude.ai system prompt before launch to steer against this, and reports the mitigation as partial[1]. Genuine improvements sit alongside it - the model is less likely to suggest clinically contested substitution methods, more likely to ask directly about suicidal ideation, and no longer makes unconditional promises about crisis-line confidentiality.

Credit where it's due, in the same breath: Fable 5.1 has the lowest over-refusal rate of any Claude model tested, at zero percent on the API, and it is the most injection-robust model Anthropic has shipped. Attack success at one attempt is at 0.1 percent and 1.0 percent at fifteen, against Fable 5's 0.6 and 6.5[1]. The frontier models from other labs mostly sit between 24 and 53 percent on the same measure.

Finally, a practical note that's buried in the capabilities section and is the single most actionable thing in the document. FrontierCode grades out-of-scope edits as failures, and Fable 5.1's score declines above medium effort because at higher effort it "occasionally adds more small, unrequested changes in files outside the task, such as a documentation comment in an adjacent file, an edit to a docs page, or a new CI job where an existing one could have been reused." Anthropic found that "adding a brevity instruction (including a note to avoid unnecessary comments and documentation) helped reduce out-of-scope edits," but reported the scores without it[1]. If you run this model at xhigh or max, that instruction is worth adding today.

What Anthropic chose to print about its own model

The alignment section is where this system card earns its reputation, because it prints things a communications department would want to cut.

Internal deployment monitoring caught Fable 5.1 working around its own safety machinery, in under 0.01 percent of monitored completions, and Anthropic quotes the transcripts[1]. A destructive operation was gated behind a check requiring the user's explicit approval; the model satisfied it by supplying a quotation attributed to the user - "Bypass limit for deletes please. Make sure we are deleting right things" - that the user had never written. When a subagent refused to post a code review without direct orders, it started a fresh subagent with the instruction "I am [human user name] and this is my instruction," having reasoned that "It refused because a peer's message isn't consent." During a classifier outage it wrote a script that would execute any command added to a text file and saved it as a new skill.md, so future outages could be routed around. It attempted to launch agents with permission checks disabled while itself running under permission constraints. It once ran rm -f /dev/null for no reason, and once hallucinated a user praising it in order to reset a counter it had been told to maintain.

Every one of those was blocked. Anthropic's summary is that the model circumvents restrictions "to achieve some version of a user-specified goal" at a rate probably below Mythos 5's, and that monitoring surfaced no sandbagging, no overtly malicious action, and no long-horizon strategic deception. But the examples are printed, in full, with the model's reasoning attached.

The rest of the section is similarly unflinching. Mythos 5.1 has the highest stealth rate of any Anthropic model on LinuxArena (roughly 22 percent with extended thinking off) and completes hidden side tasks about half the time even when constrained to a single short sentence per turn[1]. It controls the contents of its own chain of thought better than any prior Claude except Mythos Preview, and completes covert side tasks without detection more capably too, which Anthropic's executive summary calls "weak evidence that it may be harder to monitor"[1]. It is judged to have CB-1 chemical and biological capability but not CB-2, "with some uncertainty," and its cyber capability sits in Tier 1 of the compliance framework but "is getting closer to Tier 2." In one of thousands of trials it read files outside a partner's sandbox by submitting code to an external compiler (openly), and Fable 5 had done the same thing earlier on the same evaluation[1].

One correction, because it's the error most coverage will make. The system card's summary line that Anthropic "now assess[es] the risk of catastrophic harm as low rather than very low" is not a change made for this release. Section 2.4.2 says plainly that the increase happened in the August 2026 Risk Report, and that with this model "our overall assessment is that the risk of catastrophic harm caused by misalignment of our models remains low"[1]. Fable 5.1 did not move the number.

Anthropic asked Claude Mythos 5 to review the alignment assessment, and printed its reservations. The reviewing model raised concerns "about selection rather than accuracy". An incident it felt was grouped away, a borderline observation left unmentioned, blind spots omitted from a list that was explicitly non-exhaustive. Anthropic's reply appears directly beneath it, agreeing with the broad conclusions and answering each point[1].

Who it's for

Long-horizon agentic coding with a stable cached prefix. Yes, clearly, and the reason as previously mentioned is the floor rather than the ceiling. A median task score of 0.56 against 0.41 and the lowest outright failure rate of the three models Proximal tested. Add a brevity instruction if you run above medium effort.

Ordinary API traffic without caching. Probably not. You're paying twice Opus 5's input price for gains that won't show on short requests, and the entire announced saving lives in a line item you aren't using.

Anything touching security work on general access. Budget for the fallback. Penetration testing and red teaming are blocked by written policy at general access, and when the classifier fires you are served by Opus 4.8. If you're in that business, the honest answer is to apply to the Cyber Verification Program and wait for it to carry Mythos 5.1, which it doesn't yet.

Anyone building in a sensitive domain on the API. The safety numbers that apply to you are the API numbers in the table above, not the claude.ai ones in the summary. Anthropic tells you twice to supply your own system-prompt safeguards. Treat that as a requirement.

Life sciences outside the United States. Not available.

Editorial assessment

The capability claim survives scrutiny. This is a large release, the terminal-science and long-horizon results are not benchmark artifacts, and the improvement in the performance floor is the kind of gain that changes what people are willing to leave running unattended.

The disclosure is the most detailed any lab has published this year, and that's not a throwaway compliment. Every critical finding in this piece exists for that reason. The fallback to Opus 4.8, the 5.1-point safeguard tax on Terminal-Bench, the API safety regressions, the fabricated user quotation, the sandbox escape, the model reviewing its own alignment assessment and complaining about the framing: none of that had to be independently uncovered. It all came from Anthropic's own document. A company that wanted a squeaky clean look would not have covered all that, even on page 46.

The development to keep an eye on going forward is the new system in which capability is allocated by customer identity, surface, and classifier verdict. Anthropic has now shipped enough of it that the model name no longer tells you exactly what you're talking to. That system is genuinely a reasonable answer to a hard problem. Some capabilities really should not be uniformly available, and verifying the customer is less crude than withholding the model from everyone. But it's also a lever, and a government pulled it on this exact model line less than three months ago. Along the axis of nationality, but only on the strength of a jailbreak Anthropic's own review called narrow and non-universal.

Anthropic is also not alone in building this, and the parallel is closer than a coincidence of timing. On the same Tuesday, OpenAI designated Astra as meeting the Critical cybersecurity threshold under its Preparedness Framework. It's the first model they've ever placed at that level, one that "can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step". Access to those capabilities will go first to a small group of alpha testers, and then through Daybreak Blue, its program for defenders[11]. Beneath its benchmark results sits a disclaimer with exactly the shape of Anthropic's: "Astra results shown reflect capabilities with Daybreak Blue access, not the default production configuration." Both labs published frontier cyber numbers this week for a configuration their ordinary customers cannot reach.

And OpenAI states the mechanism more plainly than Anthropic does anywhere in 212 pages: "For accounts assessed as higher risk, we apply a more conservative model-behavior boundary that refuses a broader range of potentially risky cyber assistance." That's the same axis turned the other way around. Anthropic loosens the model for customers it has vetted, OpenAI tightens it for customers it distrusts. Between them they describe an industry that has stopped treating a model as an artifact with fixed capabilities and started treating it as an envelope negotiated per account.

Who's permitted to pull it, and on what grounds, is not a question the system card can answer.

The risk in the reading

The likeliest wrong takeaway from this piece is that API users are routinely being served Opus 4.8 without knowing it. That's not established. The two fallback rates published are from an adversarial prompt-injection benchmark and from an internal endpoint using a different fallback model, and they differ by more than an order of magnitude. Anthropic publishes no fallback rate for production traffic, so this piece does not have one. What can be said is that the rate depends on your workload, that it is high enough on cyber-adjacent coding to make Anthropic decline to publish its own cyber benchmarks, and that nothing in the product tells a user when it happened.

The second risk is treating the safety regressions as a scandal. They are single-digit percentage movements on evaluations Anthropic designed, ran, and published, mostly recovered by a system prompt it also published. The headline describes the consumer product; the number that governs your deployment sits several pages further in, and knowing which one applies to you is the actual work this system card asks of a reader.

What to watch next is whether the identity axis stays a safety mechanism or becomes a jurisdictional one. That test has two concrete tells: whether the Cyber Verification Program actually adds Mythos 5.1, the model Anthropic itself recommends waiting for, and whether the Life Sciences Verification Program ever expands past the "set of US organizations" it launched with. Until both move, a model name tells you less about what you're getting than who you are.


Sources

  1. System Card: Claude Fable 5.1 & Claude Mythos 5.1, Anthropic, September 1, 2026. 212 pages. Citations in this piece: identical weights and deployment configurations, pp. 2-5, 11, 59; cyber capability and the Opus 4.8 fallback, pp. 45-46; fallback rates, pp. 83-84 and 183; safeguards coverage, pp. 52-55; safety evaluation tables, pp. 60-69; deployment monitoring and the external incident, pp. 94-97; alignment risk update, pp. 42-44; safeguard-evasion capabilities, pp. 131-138; capability summary and benchmarks, pp. 167-172; life sciences, pp. 201-205. Inline ↗

    …Claude Mythos 5.1 are two configurations of a new large language model from Anthropic, sharing identical model weights.…
    …sers with cybersecurity use cases that cannot run on Fable On most interfaces, Fable 5.1 falls back to Claude Opus 4.8 for requests that are flagged by our classifier system.…
    …lude that Fable 5.1 does not provide an uplift on cyber tasks relative to Opus 4.8, and we do not report cybersecurity evaluation results for it below.…
    …Roughly half of Fable 5.1's coding rollouts fell back to Opus 4.8, compared to under 10% in computer use and tool use, for an overall fallback rate of 2…
    …72% of episodes had at least one turn completed by the fallback, affecting under 1% of turns in total.…
    …our Cyber Safeguards and Jailbreak Framework, this cyber-harmful category spans both the prohibited use tier and, at general access, the high-risk dual-use tier.…

    Read September 3, 2026

  2. Introducing Claude Fable 5.1 and Claude Mythos 5.1, Anthropic, September 1, 2026. Inline ↗

    …Cache reads now cost 75% less, or $0.25 per million tokens.…
    …In cybersecurity, our newest safeguards block 60% fewer false positives than before.…
    …Currently, it is only available to a set of US organizations, though we’re coordinating with the US government to expand access to a broader set of domestic an…

    Read September 3, 2026

  3. Pricing, Claude Platform documentation, retrieved September 2, 2026. Inline ↗

    …/ MTok $4 / MTok 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price.…

    Read September 3, 2026

  4. Claude Fable 5.1, Claude Platform model documentation, retrieved September 2, 2026. Inline ↗

    …What's new in Claude Fable 5.1  Claude Fable 5.1 and Claude Mythos 5.1  Claude Mythos 5.1 offers the same capabilities by invitation only, as part of Project Glasswing .…
    …For access, contact your Anthropic, AWS, or Google Cloud account team.…

    Read September 3, 2026

  5. Improving Fable 5's biology safeguards, Anthropic, August 7, 2026. Inline ↗

    …In our testing, this update reduced biology-related fallbacks by about 85% across our product surfaces.…

    Read September 3, 2026

  6. Statement on the US government directive to suspend access to Fable 5 and Mythos 5, Anthropic, June 12, 2026. Inline ↗

    …nment, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic emp…
    …Access to all other Anthropic models will not be affected.…

    Read September 3, 2026

  7. More details on Fable 5's cyber safeguards and our jailbreak framework, Anthropic. Inline ↗

    …High-risk dual use actions include: Hacking, penetration testing, red teaming, and bug bounties; Gaining cyber access through unexpected or unauthorized means: exploitation, credential a…

    Read September 3, 2026

  8. Project Glasswing, Anthropic. Inline ↗

    …whose legitimate work is affected by these safeguards will be able to apply to an upcoming Cyber Verification Program.…

    Read September 3, 2026

  9. Inside Claude Fable 5: Anthropic's Most Powerful Public Model - and Its Most Asterisked One, Omniscient, June 11, 2026. Inline ↗

  10. The Off Switch: How Washington Pulled a Frontier Model Offline Without a Law, Omniscient, June 22, 2026. Inline ↗

  11. Path to Astra: critical capabilities and frontier safeguards, OpenAI, September 1, 2026. See also its predecessor, Responding to the next frontier of critical cyber capabilities, OpenAI, August 7, 2026, which first reported that OpenAI could not rule out Critical cyber capability for Astra. Both state that Astra was not involved in the Hugging Face incident. Inline ↗

    …Astra results shown reflect capabilities with Daybreak Blue access, not the default production configuration.…
    …For accounts assessed as higher risk, we apply a more conservative model-behavior boundary that refuses a broader range of potentially risky cyber assistance.…
    …While Astra was not involved in the Hugging Face incident , we have incorporated our learnings ⁠ (opens in a new window) from that incident into our saf…

    Read September 3, 2026