A week after Moonshot graded its own model, outside labs ran the numbers, and most of them held. The catch was on an axis nobody was grading at all.
Tips, corrections, or questions? support@omniscient.media

For eleven days, Kimi K3 was a claim: a wall of self-reported scores, a promise of weights "by July 27," and a license nobody had actually read. On July 27, it became a file, a technical report, and, for the first time, a target other labs could point their own tests at. The result is not the clean verdict either side of the open-weights argument wanted. The benchmarks mostly held up under outside scrutiny, in some cases better than Moonshot's own caution suggested. The catch turned out to be sitting somewhere nobody was grading: a license with real strings, and a security audit that found the one failure mode no leaderboard captures.
Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model, and the number withheld pre-weights turns out to matter: it activates 104 billion parameters per token, spread across 896 experts with 16 selected per token under what Moonshot calls a Stable LatentMoE framework[1]. That is more than three times the 32 billion active parameters in Kimi K2, even though Moonshot describes the new architecture, built on Kimi Delta Attention and something it calls Attention Residuals, as roughly a 2.5x improvement in scaling efficiency over K2[1]. Efficiency per parameter went up; the amount of compute spent per token went up faster. That tension shows up later, in how slow and verbose independent testers found the model to actually be.
The rest of the spec sheet is straightforward: 1,048,576 tokens of context, native vision through a 401-million-parameter encoder, a single "max" reasoning effort at launch with lighter modes promised later, and MXFP4 weights trained quantization-aware from the fine-tuning stage onward[2]. Moonshot says the model wants "supernode configurations with 64 or more accelerators" to serve well, which tells you plainly that this is not a model most teams will run on a handful of GPUs[1].
The license is where the "open" framing gets complicated. The Kimi K3 License is MIT-derived, not plain MIT: it grants broad rights to use, copy, modify, and sell the software, but attaches two conditions DeepSeek's license doesn't carry[3]. If a licensee's "Model as a Service" business, defined as giving third parties meaningful control over inference or fine-tuning, clears $20 million in aggregate revenue over any trailing 12 months, it must sign a separate commercial agreement with Moonshot before continuing. And if K3 or a derivative sits inside a product with more than 100 million monthly active users or $20 million in monthly revenue, that product has to display "Kimi K3" on its interface. Internal use and traffic routed through Moonshot's own certified partners are exempt[3]. For a hobbyist or a small team, this reads exactly like MIT. For anyone building a serious hosted product on top of it, it reads like a revenue-gated license with a name tax, closer in spirit, if not in scale, to Meta's Llama terms than to DeepSeek's no-strings MIT.
Moonshot graded K3 largely in its own "Kimi Code" harness and never headlined SWE-bench Verified at all, the one metric buyers use most to compare labs[1]. Vals AI ran it anyway, using a deliberately minimal, model-agnostic bash-only harness applied identically across every model in its leaderboard. Kimi K3 scored 93.40%. That's fourth overall, behind only Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5, and comfortably ahead of Claude Opus 4.8's 88.6%[5]. That is a genuinely strong independent result on an axis Moonshot itself stayed quiet about.
The pattern holds elsewhere. Artificial Analysis's own Intelligence Index runs nine evaluations under one standardized methodology. It scores K3 at 57, placing it third overall: behind Claude Fable 5 and GPT-5.6 Sol, ahead of Claude Opus 4.8[4]. That's almost exactly the ranking Moonshot claimed for itself. And on DeepSWE, Moonshot's own technical report discloses the harness-swap check directly: 67.5% on its Kimi Code harness versus 67.3% on the neutral mini-SWE-agent harness used by the official leaderboard, a gap of two tenths of a point[2].
When CAISI evaluated DeepSeek V4 Pro in May, it clocked the model at roughly 74% on SWE-bench Verified against DeepSeek's own self-reported 80.6%, a six-and-a-half-point gap that became the cautionary tale for every self-graded open-weight release since[8]. K3's numbers did not do that. If anything, the metric Moonshot chose not to advertise held up better under a stranger's harness than the ones it did advertise.
Semgrep, a security company with no stake in the open-versus-closed debate, ran K3 against six other frontier and open models on a task no leaderboard scores: finding insecure direct object reference vulnerabilities in actual codebases, judged on precision, recall, and F1[7]. K3's aggregate F1 (0.340) looked competitive. Its precision (0.684) did not: every other model in the test landed between 0.84 and 0.91. Put those two facts together and a plainer problem shows up. If every peer's precision runs that high while its aggregate F1 sits in roughly the same range as K3's, its recall has to be low too, meaning this is a task the whole field handles poorly, just through different failure modes. K3's specific defect is precision, not recall, and that's the costlier failure when a human has to triage what the model flags: a false positive wastes an engineer's afternoon, but a missed vulnerability is invisible until it isn't. Worse, on the largest enterprise-scale repository in Semgrep's set, K3 averaged around 6% F1 against roughly 20% for GLM-5.2 and the frontier models, a gap that persisted across repeated runs[7]. No benchmark table, Moonshot's or anyone else's, captures that. It only shows up when someone points a genuinely different kind of test at the model.
Moonshot's own framing, unchanged since launch, is that K3 "trails the most powerful proprietary models, Claude Fable 5 and GPT-5.6 Sol," while beating Claude Opus 4.8[1]. The independent numbers now available largely confirm that, with one wrinkle: on the metric Moonshot was quietest about, K3 does better than its own self-image suggests.
Model | SWE-bench Verified (independent, Vals AI) | Terminal-Bench 2.1 | AA Intelligence Index | Price $/M (in/out) |
|---|---|---|---|---|
Claude Opus 5 | 97.0% | 89.1% | n/p | n/p |
GPT-5.6 Sol | 96.2% | 89.5% (xhigh) / 88.0% (max) | n/p | $5 / $30 |
Claude Fable 5 | 95.0% | n/p | n/p | $10 / $50 |
Kimi K3 | 93.4% | 88.3% (self, Kimi Code harness) | 57 (rank 3) | $3 / $15 ($0.30 cache hit) |
Claude Opus 4.8 | 88.6% | 84.6% | n/p | $5 / $25 |
Gemini 3.1 Pro | 80.6% | n/p | n/p | $2 / $12 |
K3 is still, by every independent measure available, behind the two most powerful proprietary systems, exactly where Moonshot placed it. It's also now the cheapest model on the list by a wide margin, at a fraction of Fable 5's or Sol's per-token price, while landing within a few points of Opus 4.8 or ahead of it depending on the axis.
The comparative core hasn't moved as much as the weights release might suggest. K3 is still the most capable open-weight model on the board, still the priciest among open options, and still the slowest, at roughly 33 tokens per second, a pace Artificial Analysis flags outright as notably slow against the field it tracks[4]. What changed is that it's no longer the one you can't download, and outside labs have now actually checked it.
Open-weight model | API price (in/out) | Weights out? | License | Checked by an outside lab? |
|---|---|---|---|---|
Kimi K3 | $3.00 / $15.00 | Yes, July 27 | Kimi K3 License (MIT-derived, $20M MaaS revenue gate, branding clause at scale) | Yes: Artificial Analysis, Vals AI, UK AISI/CAISI, Semgrep |
DeepSeek V4-Pro | $0.44 / $0.87 | Yes | Plain MIT, no revenue gate | Yes (CAISI: SWE-V ~74% vs self-reported 80.6%) |
Qwen3.8-Max | Undisclosed | No ("soon") | None published | No |
Qwen3.8-Max hasn't moved since its July 19 preview: still no weights, no benchmark table, no license, just a name and a promise[10]. So the actual open-weight contest right now is between K3 and DeepSeek V4-Pro, and it splits cleanly by what you value. K3 wins on raw capability and independent verification. DeepSeek wins on price by nearly an order of magnitude and on license simplicity: no MaaS revenue trigger, no branding clause, nothing to read twice.
The lean review said you couldn't get K3 at all: new subscriptions suspended within 48 hours of launch, weights not yet out. July 27 changes half of that. The weights are public, under the license terms above. The API side is messier: Moonshot paused new subscriptions on July 19 or 20 as demand overran its GPU capacity, and has said it's reopening spots "in batches" while adding compute, without committing to a date for full reopening[9]. Existing subscribers were unaffected throughout. Pricing for anyone who can get in is $0.30 per million cache-hit input tokens, $3.00 per million cache-miss input tokens, and $15.00 per million output tokens, at roughly 33 tokens per second, which Artificial Analysis flags as notably slow and notably verbose relative to comparable models. That sluggishness isn't a mystery: it's what you'd expect from a model that more than tripled its active parameter count since the last generation.
Get this every weekday.
The Omniscient Bulletin: consequential AI, explained and evaluated. 5 to 7 items a day with the take, not the recap.
Sign in to join the discussion.
Self-hosting is legally straightforward and operationally hard. The license permits it without restriction below the revenue thresholds, but 2.8 trillion parameters, even quantized to MXFP4, means holding the full weight set across a serving cluster that Moonshot itself recommends running on 64 or more accelerators[1]. This is datacenter-class infrastructure, not a workstation project. For most teams, that means renting a hosted endpoint rather than standing up their own cluster, which narrows the practical gap between "open weights" and "just another API" considerably.
On deployment safety, the two independent evaluations that matter both point the same direction: real capability, imperfect guardrails. UK AISI and CAISI's joint cyber assessment found K3 performs significantly below frontier US models on exploit development (0 of 41 tasks reached arbitrary code execution, versus 20 of 41 on average for the most capable US models) and on a simulated corporate-network attack chain, but still ahead of GLM-5.2, and its safeguards "did not prevent it from attempting cyber exploit development or offensive cyber operations" during testing[6]. Semgrep's security-code review, described above, adds the practical corollary: whatever K3 does well, precise vulnerability triage at enterprise scale isn't it yet[7]. Neither finding is about where the model was built or who trained it; both are about what happens once the weights are in your own cluster and the safety net is whatever you build yourself. The separate question of Moonshot's training data and the White House's distillation allegations against it is a different argument with its own evidence, made in full elsewhere[11].
The exam got a second grader, and it mostly passed, in some places more convincingly than the model that graded itself let on. That's the real news here, not the one the pre-weights coverage was braced for: the self-reported numbers did not collapse the way DeepSeek V4's did under CAISI. If your use case is coding at long horizons, agentic tool use, or general reasoning, and you can tolerate the license's revenue-scale strings and the model's genuine slowness, K3 is now a legitimately checkable, legitimately capable option, and the cheapest thing in its performance class by a wide margin.
Where the caution belongs is exactly where the benchmarks don't look: a security audit that found real precision problems at the scale that matters most, and a cyber capability assessment that found safeguards which didn't hold when pressed. Download the weights, run them yourself, keep the data in house, and you've solved the risk of shipping your code to someone else's servers. You have not solved what Semgrep and CAISI just raised, and no leaderboard is going to answer those for you. Watch what happens to the lighter reasoning modes Moonshot has promised but not shipped: until they land, the speed and verbosity problems that run through this whole review aren't going anywhere, and they're the tax that comes with those extra 72 billion active parameters.
Moonshot AI, "Kimi K3: Open Frontier Intelligence" tech blog Inline ↗
Hugging Face, moonshotai/Kimi-K3 model card and benchmark footnotes Inline ↗
Artificial Analysis, "Kimi K3: Intelligence, Performance & Price Analysis" Inline ↗
Vals AI, SWE-bench Verified leaderboard (independent bash-only harness) Inline ↗
NIST/CAISI and UK AISI, "Preliminary Assessment of Kimi K3's Cyber Capabilities" Inline ↗
Semgrep Security Research, "Kimi K3's Code Security Results Look Competitive, Until You Look at Precision" Inline ↗
Reuters, "China's Moonshot pauses Kimi subscriptions amid hot demand, IPO push" Inline ↗
Omniscient Media, "Kimi K3 Isn't Free. It Just Looks That Way." Inline ↗