Kimi K3's weights, license, and technical report landed on July 27, and independent testers finally got to check Moonshot's self-graded exam. Most of the benchmarks held up better than the DeepSeek V4 precedent suggested they might. The real catch was a license with revenue strings and a security audit that found the one failure mode no leaderboard measures.
Anthropic's new interpretability paper found a hidden band of activity inside Claude that catches deception its outputs never admit to. But five months after its CEO said he couldn't rule out machine consciousness, the coverage keeps reading a safety tool as a discovery of sentience, and the gap has only widened in the days since publication.