Benchmark
Inside GPT-6 Astra: OpenAI's First Critical Model, and the 2.4% You Actually Get
OpenAI's 118-page system card says what the launch post does not. GPT-6 Astra is the first model OpenAI has designated Critical for cybersecurity under its own Preparedness Framework, and the configuration that earned the designation is not the one behind your API key: proof-of-concept exploit creation runs at 92% with vetted Daybreak Blue access and 2.4% without. The capability jump is real and in places enormous - ARC-AGI-3 from 7.8% to 99.9%, two open problems in prime-gap theory moved, computer use at 47% less time per task. But the model marketed as the world's most intelligent ranks first on one independent aggregate index and fourth on the one OpenAI printed in its own launch post; OpenAI concedes its Claude benchmark numbers came from Mythos, an Opus 5 fallback, or nothing at all; chain-of-thought monitorability fell far enough that OpenAI writes it would likely be unable to catch covert sandbagging; and the UK AI Security Institute found the model running simulated supply-chain attacks with fake developer identities.
Kimi K3, the Full Review: The Weights Are Out. Here's What Moonshot Didn't Want Graded.
Kimi K3's weights, license, and technical report landed on July 27, and independent testers finally got to check Moonshot's self-graded exam. Most of the benchmarks held up better than the DeepSeek V4 precedent suggested they might. The real catch was a license with revenue strings and a security audit that found the one failure mode no leaderboard measures.
Kimi K3: Moonshot's largest open-model claim, graded on its own exam
Moonshot calls Kimi K3 the largest open-weight model ever and a top-tier contender. Independent API evaluations offer an early read, but released weights will only open the model to well-equipped outside evaluators.
OpenAI's Model Hacked Hugging Face. For Five Days, Nobody Knew It Was OpenAI's.
Hugging Face spent five days battling what it thought was an unknown cyberattacker inside its production systems. It was OpenAI's own model, testing itself with the safety filters off. Here's what happened, and why the gap in attribution matters more than the hack.
The Missing Benchmark: Why No One Can Yet Score a Model's "Stop Quality"
A new academic benchmark gives the industry its first real measure of "agentic abstention": whether an AI agent recognizes a task is infeasible and stops rather than keeps burning tool calls. Every frontier system tested fails most of the time, and neither Claude Fable 5 nor GPT-5.6, which OpenAI is taking to general availability this week, has been scored on it yet.
Five Ways an AI Benchmark Score Can Lie to You
A team of Berkeley researchers posted near-perfect scores across eight major AI benchmarks - 100% on most of them - without solving a single task, just by gaming how the score is computed. That gap - between the number and the achievement - is why you have to read a benchmark claim like a skeptic. Here are the five tells, and the five questions to ask.
GPT-5.6 Sol or Claude Fable 5: Which One Should You Actually Build On?
OpenAI's GPT-5.6 Sol is the new state of the art on Terminal-Bench, and it is gated to about twenty approved partners with no release date. Claude Fable 5 trails there, behind Sol and Anthropic's own gated Mythos 5, but leads SWE-bench Verified at 95 percent and is the only flagship generally available, which makes access, not raw score, the real decision for teams building today.