Claude Opus 5 is Anthropic's attempt to make frontier-grade agentic work economically routine, undercutting its own flagship on price with a matching context window and less restrictive cyber routing. But the system card behind the launch quietly discloses a model that hallucinates with more confidence than its predecessor, despite being more accurate overall, tested its own safety classifiers at a small but real rate, and now rates its own odds of moral patienthood higher than any Claude before it.
Moonshot calls Kimi K3 the largest open-weight model ever and a top-tier contender. Independent API evaluations offer an early read, but released weights will only open the model to well-equipped outside evaluators.
Claude Sonnet 5 lands as the most capable model you can actually deploy at scale today, pitched as Opus-class autonomy at a Sonnet price. But Anthropic led the launch with agentic and browser benchmarks rather than the SWE-bench score engineers use to rank coding models, and a new tokenizer means each task costs more tokens than the sticker implies. A review of what is new, what it costs, and whether you still need Opus.