Omniscient
AllBulletinArticlesReviewsTakesCommentaryFeatured
Sign In

Omniscient

AI intelligence briefings, analysis, and commentary — delivered in broadsheet form.

By Noah Ogbi

Subscribe

Weekday briefings and flagship analysis, delivered to your inbox.

Sections

  • All
  • Bulletin
  • Articles
  • Reviews
  • Takes
  • Commentary

Topics

  • Industry Strategy
  • Anthropic
  • AI Policy
  • Frontier Models
  • OpenAI
  • Compute Economics
  • Research
  • Agents

Meta

  • About
  • Masthead
  • Standards
  • Corrections
  • RSS Feed
  • Privacy Policy
  • Terms of Service

Omniscient Media — made by ForeverBuilt, LLC.
© 2026 ForeverBuilt, LLC. All rights reserved.

  1. Home
  2. ›AI Research
  3. ›Kimi K3, the Full Review: The Weights Are Out. Here's What Moonshot Didn't Want Graded.

AI Research

Vol. 1·Monday, July 27, 2026

Kimi K3, the Full Review: The Weights Are Out. Here's What Moonshot Didn't Want Graded.

A week after Moonshot graded its own model, outside labs ran the numbers, and most of them held. The catch was on an axis nobody was grading at all.


Noah Ogbi12 min read

Tips, corrections, or questions? support@omniscient.media

TopicsBenchmarkAI SecurityFrontier ModelsCoding & DevTools
CompaniesMoonshot AI
Kimi K3, the Full Review: The Weights Are Out. Here's What Moonshot Didn't Want Graded.

For eleven days, Kimi K3 was a claim: a wall of self-reported scores, a promise of weights "by July 27," and a license nobody had actually read. On July 27, it became a file, a technical report, and, for the first time, a target other labs could point their own tests at. The result is not the clean verdict either side of the open-weights argument wanted. The benchmarks mostly held up under outside scrutiny, in some cases better than Moonshot's own caution suggested. The catch turned out to be sitting somewhere nobody was grading: a license with real strings, and a security audit that found the one failure mode no leaderboard captures.

What did the weights and technical report actually reveal?

Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model, and the number withheld pre-weights turns out to matter: it activates 104 billion parameters per token, spread across 896 experts with 16 selected per token under what Moonshot calls a Stable LatentMoE framework[1]. That is more than three times the 32 billion active parameters in Kimi K2, even though Moonshot describes the new architecture, built on Kimi Delta Attention and something it calls Attention Residuals, as roughly a 2.5x improvement in scaling efficiency over K2[1]. Efficiency per parameter went up; the amount of compute spent per token went up faster. That tension shows up later, in how slow and verbose independent testers found the model to actually be.

The rest of the spec sheet is straightforward: 1,048,576 tokens of context, native vision through a 401-million-parameter encoder, a single "max" reasoning effort at launch with lighter modes promised later, and MXFP4 weights trained quantization-aware from the fine-tuning stage onward[2]. Moonshot says the model wants "supernode configurations with 64 or more accelerators" to serve well, which tells you plainly that this is not a model most teams will run on a handful of GPUs[1].

The license is where the "open" framing gets complicated. The Kimi K3 License is MIT-derived, not plain MIT: it grants broad rights to use, copy, modify, and sell the software, but attaches two conditions DeepSeek's license doesn't carry[3]. If a licensee's "Model as a Service" business, defined as giving third parties meaningful control over inference or fine-tuning, clears $20 million in aggregate revenue over any trailing 12 months, it must sign a separate commercial agreement with Moonshot before continuing. And if K3 or a derivative sits inside a product with more than 100 million monthly active users or $20 million in monthly revenue, that product has to display "Kimi K3" on its interface. Internal use and traffic routed through Moonshot's own certified partners are exempt[3]. For a hobbyist or a small team, this reads exactly like MIT. For anyone building a serious hosted product on top of it, it reads like a revenue-gated license with a name tax, closer in spirit, if not in scale, to Meta's Llama terms than to DeepSeek's no-strings MIT.

Do the benchmarks survive an independent harness?

Moonshot graded K3 largely in its own "Kimi Code" harness and never headlined SWE-bench Verified at all, the one metric buyers use most to compare labs[1]. Vals AI ran it anyway, using a deliberately minimal, model-agnostic bash-only harness applied identically across every model in its leaderboard. Kimi K3 scored 93.40%. That's fourth overall, behind only Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5, and comfortably ahead of Claude Opus 4.8's 88.6%[5]. That is a genuinely strong independent result on an axis Moonshot itself stayed quiet about.

The pattern holds elsewhere. Artificial Analysis's own Intelligence Index runs nine evaluations under one standardized methodology. It scores K3 at 57, placing it third overall: behind Claude Fable 5 and GPT-5.6 Sol, ahead of Claude Opus 4.8[4]. That's almost exactly the ranking Moonshot claimed for itself. And on DeepSWE, Moonshot's own technical report discloses the harness-swap check directly: 67.5% on its Kimi Code harness versus 67.3% on the neutral mini-SWE-agent harness used by the official leaderboard, a gap of two tenths of a point[2].

When CAISI evaluated DeepSeek V4 Pro in May, it clocked the model at roughly 74% on SWE-bench Verified against DeepSeek's own self-reported 80.6%, a six-and-a-half-point gap that became the cautionary tale for every self-graded open-weight release since[8]. K3's numbers did not do that. If anything, the metric Moonshot chose not to advertise held up better under a stranger's harness than the ones it did advertise.

Semgrep, a security company with no stake in the open-versus-closed debate, ran K3 against six other frontier and open models on a task no leaderboard scores: finding insecure direct object reference vulnerabilities in actual codebases, judged on precision, recall, and F1[7]. K3's aggregate F1 (0.340) looked competitive. Its precision (0.684) did not: every other model in the test landed between 0.84 and 0.91. Put those two facts together and a plainer problem shows up. If every peer's precision runs that high while its aggregate F1 sits in roughly the same range as K3's, its recall has to be low too, meaning this is a task the whole field handles poorly, just through different failure modes. K3's specific defect is precision, not recall, and that's the costlier failure when a human has to triage what the model flags: a false positive wastes an engineer's afternoon, but a missed vulnerability is invisible until it isn't. Worse, on the largest enterprise-scale repository in Semgrep's set, K3 averaged around 6% F1 against roughly 20% for GLM-5.2 and the frontier models, a gap that persisted across repeated runs[7]. No benchmark table, Moonshot's or anyone else's, captures that. It only shows up when someone points a genuinely different kind of test at the model.

Where does it land against the US frontier?

Moonshot's own framing, unchanged since launch, is that K3 "trails the most powerful proprietary models, Claude Fable 5 and GPT-5.6 Sol," while beating Claude Opus 4.8[1]. The independent numbers now available largely confirm that, with one wrinkle: on the metric Moonshot was quietest about, K3 does better than its own self-image suggests.

Model

SWE-bench Verified (independent, Vals AI)

Terminal-Bench 2.1

AA Intelligence Index

Price $/M (in/out)

Claude Opus 5

97.0%

89.1%

n/p

n/p

GPT-5.6 Sol

96.2%

89.5% (xhigh) / 88.0% (max)

n/p

$5 / $30

Claude Fable 5

95.0%

n/p

n/p

$10 / $50

Kimi K3

93.4%

88.3% (self, Kimi Code harness)

57 (rank 3)

$3 / $15 ($0.30 cache hit)

Claude Opus 4.8

88.6%

84.6%

n/p

$5 / $25

Gemini 3.1 Pro

80.6%

n/p

n/p

$2 / $12

K3 is still, by every independent measure available, behind the two most powerful proprietary systems, exactly where Moonshot placed it. It's also now the cheapest model on the list by a wide margin, at a fraction of Fable 5's or Sol's per-token price, while landing within a few points of Opus 4.8 or ahead of it depending on the axis.

There's one of these every weekday.

The Omniscient Bulletin turns the day's AI news into 5 to 7 items with the take, not the recap. Free.

Is it the new king of open weights?

The comparative core hasn't moved as much as the weights release might suggest. K3 is still the most capable open-weight model on the board, still the priciest among open options, and still the slowest, at roughly 33 tokens per second, a pace Artificial Analysis flags outright as notably slow against the field it tracks[4]. What changed is that it's no longer the one you can't download, and outside labs have now actually checked it.

Open-weight model

API price (in/out)

Weights out?

License

Checked by an outside lab?

Kimi K3

$3.00 / $15.00

Yes, July 27

Kimi K3 License (MIT-derived, $20M MaaS revenue gate, branding clause at scale)

Yes: Artificial Analysis, Vals AI, UK AISI/CAISI, Semgrep

DeepSeek V4-Pro

$0.44 / $0.87

Yes

Plain MIT, no revenue gate

Yes (CAISI: SWE-V ~74% vs self-reported 80.6%)

Qwen3.8-Max

Undisclosed

No ("soon")

None published

No

Qwen3.8-Max hasn't moved since its July 19 preview: still no weights, no benchmark table, no license, just a name and a promise[10]. So the actual open-weight contest right now is between K3 and DeepSeek V4-Pro, and it splits cleanly by what you value. K3 wins on raw capability and independent verification. DeepSeek wins on price by nearly an order of magnitude and on license simplicity: no MaaS revenue trigger, no branding clause, nothing to read twice.

Open weights or hosted: which K3 should you run, and can you get it?

The lean review said you couldn't get K3 at all: new subscriptions suspended within 48 hours of launch, weights not yet out. July 27 changes half of that. The weights are public, under the license terms above. The API side is messier: Moonshot paused new subscriptions on July 19 or 20 as demand overran its GPU capacity, and has said it's reopening spots "in batches" while adding compute, without committing to a date for full reopening[9]. Existing subscribers were unaffected throughout. Pricing for anyone who can get in is $0.30 per million cache-hit input tokens, $3.00 per million cache-miss input tokens, and $15.00 per million output tokens, at roughly 33 tokens per second, which Artificial Analysis flags as notably slow and notably verbose relative to comparable models. That sluggishness isn't a mystery: it's what you'd expect from a model that more than tripled its active parameter count since the last generation.

Share:

Get this every weekday.

The Omniscient Bulletin: consequential AI, explained and evaluated. 5 to 7 items a day with the take, not the recap.

Discussion


Sign in to join the discussion.


Related

AI Models

Vol. 1·Thursday, July 23, 2026

Kimi K3: Moonshot's largest open-model claim, graded on its own exam

Kimi K3: Moonshot's largest open-model claim, graded on its own exam

Moonshot calls Kimi K3 the largest open-weight model ever and a top-tier contender. Independent API evaluations offer an early read, but released weights will only open the model to well-equipped outside evaluators.


BenchmarkCompute EconomicsDeepSeek
Noah Ogbi8 min read
Continue →

AI Policy

Vol. 1·Monday, July 20, 2026

Kimi K3 Isn't Free. It Just Looks That Way.

Kimi K3 is set to become the largest open-weight model ever released. What American labs say about how models like it are built, and what the US government found when it tested their predecessors, should give any American user pause.


Kimi K3 Isn't Free. It Just Looks That Way.

Moonshot's Kimi K3 is set to join DeepSeek's models and Alibaba's Qwen as a free download climbing the leaderboards. But between Anthropic and OpenAI's distillation findings, NIST's security testing, and Beijing's own moves to lock down its "open" frontier, the case for treating these weights as neutral technology is getting harder to make.


AI PolicyAI SecurityDeepSeek
Noah Ogbi8 min read
Continue →

AI Research

Vol. 1·Thursday, April 9, 2026

The Benchmark Racket: Why the Frontier Model Race Is Measuring the Wrong Thing

The Benchmark Racket: Why the Frontier Model Race Is Measuring the Wrong Thing

Six publicly available frontier models are clustered within 1.3 percentage points on the industry's most-cited coding benchmark. Meanwhile, a withheld model just scored 93.9% on the same test. The measurement system isn't broken - it's being gamed at two levels simultaneously.


Frontier ModelsResearch
Noah Ogbi13 min read
Continue →
[4]

Self-hosting is legally straightforward and operationally hard. The license permits it without restriction below the revenue thresholds, but 2.8 trillion parameters, even quantized to MXFP4, means holding the full weight set across a serving cluster that Moonshot itself recommends running on 64 or more accelerators[1]. This is datacenter-class infrastructure, not a workstation project. For most teams, that means renting a hosted endpoint rather than standing up their own cluster, which narrows the practical gap between "open weights" and "just another API" considerably.

On deployment safety, the two independent evaluations that matter both point the same direction: real capability, imperfect guardrails. UK AISI and CAISI's joint cyber assessment found K3 performs significantly below frontier US models on exploit development (0 of 41 tasks reached arbitrary code execution, versus 20 of 41 on average for the most capable US models) and on a simulated corporate-network attack chain, but still ahead of GLM-5.2, and its safeguards "did not prevent it from attempting cyber exploit development or offensive cyber operations" during testing[6]. Semgrep's security-code review, described above, adds the practical corollary: whatever K3 does well, precise vulnerability triage at enterprise scale isn't it yet[7]. Neither finding is about where the model was built or who trained it; both are about what happens once the weights are in your own cluster and the safety net is whatever you build yourself. The separate question of Moonshot's training data and the White House's distillation allegations against it is a different argument with its own evidence, made in full elsewhere[11].

So, is Kimi K3 worth building on?

The exam got a second grader, and it mostly passed, in some places more convincingly than the model that graded itself let on. That's the real news here, not the one the pre-weights coverage was braced for: the self-reported numbers did not collapse the way DeepSeek V4's did under CAISI. If your use case is coding at long horizons, agentic tool use, or general reasoning, and you can tolerate the license's revenue-scale strings and the model's genuine slowness, K3 is now a legitimately checkable, legitimately capable option, and the cheapest thing in its performance class by a wide margin.

Where the caution belongs is exactly where the benchmarks don't look: a security audit that found real precision problems at the scale that matters most, and a cyber capability assessment that found safeguards which didn't hold when pressed. Download the weights, run them yourself, keep the data in house, and you've solved the risk of shipping your code to someone else's servers. You have not solved what Semgrep and CAISI just raised, and no leaderboard is going to answer those for you. Watch what happens to the lighter reasoning modes Moonshot has promised but not shipped: until they land, the speed and verbosity problems that run through this whole review aren't going anywhere, and they're the tax that comes with those extra 72 billion active parameters.


Sources

  1. Moonshot AI, "Kimi K3: Open Frontier Intelligence" tech blog Inline ↗

  2. Hugging Face, moonshotai/Kimi-K3 model card and benchmark footnotes Inline ↗

  3. Moonshot AI, Kimi K3 License (full text) Inline ↗

  4. Artificial Analysis, "Kimi K3: Intelligence, Performance & Price Analysis" Inline ↗

  5. Vals AI, SWE-bench Verified leaderboard (independent bash-only harness) Inline ↗

  6. NIST/CAISI and UK AISI, "Preliminary Assessment of Kimi K3's Cyber Capabilities" Inline ↗

  7. Semgrep Security Research, "Kimi K3's Code Security Results Look Competitive, Until You Look at Precision" Inline ↗

  8. NIST/CAISI, "CAISI Evaluation of DeepSeek V4 Pro" Inline ↗

  9. Reuters, "China's Moonshot pauses Kimi subscriptions amid hot demand, IPO push" Inline ↗

  10. MarkTechPost, "Alibaba Previews Qwen3.8-Max" Inline ↗

  11. Omniscient Media, "Kimi K3 Isn't Free. It Just Looks That Way." Inline ↗