Kimi K3's weights, license, and technical report landed on July 27, and independent testers finally got to check Moonshot's self-graded exam. Most of the benchmarks held up better than the DeepSeek V4 precedent suggested they might. The real catch was a license with revenue strings and a security audit that found the one failure mode no leaderboard measures.
Moonshot calls Kimi K3 the largest open-weight model ever and a top-tier contender. Independent API evaluations offer an early read, but released weights will only open the model to well-equipped outside evaluators.
Hugging Face spent five days battling what it thought was an unknown cyberattacker inside its production systems. It was OpenAI's own model, testing itself with the safety filters off. Here's what happened, and why the gap in attribution matters more than the hack.
A new academic benchmark gives the industry its first real measure of "agentic abstention": whether an AI agent recognizes a task is infeasible and stops rather than keeps burning tool calls. Every frontier system tested fails most of the time, and neither Claude Fable 5 nor GPT-5.6, which OpenAI is taking to general availability this week, has been scored on it yet.
A team of Berkeley researchers posted near-perfect scores across eight major AI benchmarks - 100% on most of them - without solving a single task, just by gaming how the score is computed. That gap - between the number and the achievement - is why you have to read a benchmark claim like a skeptic. Here are the five tells, and the five questions to ask.
OpenAI's GPT-5.6 Sol is the new state of the art on Terminal-Bench, and it is gated to about twenty approved partners with no release date. Claude Fable 5 trails there, behind Sol and Anthropic's own gated Mythos 5, but leads SWE-bench Verified at 95 percent and is the only flagship generally available, which makes access, not raw score, the real decision for teams building today.