Anthropic's 212-page system card says what the launch page does not. Fable 5.1 and Mythos 5.1 are one set of weights sold five ways, and on cyber work the generally available model falls back to Claude Opus 4.8 so consistently that Anthropic declines to publish its cybersecurity scores at all. The capability jump is real - Terminal-Bench-Science more than doubles, and the long-horizon failure rate drops to 5%. But the safeguard layer costs 5.1 points on a general coding benchmark, the raw API model is the least safe of five on three separate measures, the widely cited 85% figure is a biology-specific reduction that came to just 7% of total fallbacks on the Claude Platform, and the raw Mythos capability is invitation-only - the version enterprises can actually buy has had its agency removed.
Dario Amodei is right that capability-tiered testing isn't regulatory capture, it's a tax on being biggest. But investor David Sacks has a real counter: that same tax only holds if smaller labs actually clear the queue faster, and neither Amodei nor Gavin Baker is asking whether any lab can prove its safety claims at all.