Omniscient
AllBulletinArticlesReviewsTakesCommentaryFeatured
Sign In

Omniscient

AI intelligence briefings, analysis, and commentary — delivered in broadsheet form.

By Noah Ogbi

Subscribe

Weekday briefings and flagship analysis, delivered to your inbox.

Sections

  • All
  • Bulletin
  • Articles
  • Reviews
  • Takes
  • Commentary

Topics

  • Industry Strategy
  • AI Policy
  • Anthropic
  • OpenAI
  • Frontier Models
  • Compute Economics
  • Research
  • Agents

Meta

  • About
  • Masthead
  • Standards
  • Corrections
  • RSS Feed
  • Privacy Policy
  • Terms of Service

Omniscient Media — made by ForeverBuilt, LLC.
© 2026 ForeverBuilt, LLC. All rights reserved.

  1. Home
  2. ›AI Models
  3. ›Kimi K3: Moonshot's largest open-model claim, graded on its own exam

AI Models

Vol. 1·Thursday, July 23, 2026

Kimi K3: Moonshot's largest open-model claim, graded on its own exam

Moonshot's 2.8-trillion-parameter model is a serious hosted contender. Its weights, license, and reproducible local testing are due July 27.


Noah Ogbi8 min read

Tips, corrections, or questions? support@omniscient.media

TopicsBenchmarkCompute EconomicsFrontier ModelsCoding & DevTools
CompaniesDeepSeekMoonshot AI
Kimi K3: Moonshot's largest open-model claim, graded on its own exam

There's one of these every weekday.

The Omniscient Bulletin turns the day's AI news into 5 to 7 items with the take, not the recap. Free.

So should you wait for July 27?

For most people, yes. K3 asks to be taken seriously on numbers that cannot yet be reproduced on local weights, runs on an API that independent testing finds slow and verbose, and sits behind a subscription service Moonshot has limited after demand pressed against available capacity[4][11]. Existing subscribers retained access, and Moonshot says it will reopen spots in batches as it adds capacity[11].

Who should still poke at it now: teams doing front-end or agentic coding, where K3's clearest outside result lives, that can treat every score as provisional and absorb the token bill. They should also think about where the data goes. Any hosted model routes prompts through a provider's servers, so that part is not unique to K3. What is specific now is that the hosted API is the only way to use K3. For regulated or sensitive work, procurement teams should assess Moonshot's terms, retention practices, and applicable jurisdiction before sending data. The safe default is boring: if the data is sensitive and you cannot yet run the model yourself, wait.

For other users, DeepSeek V4 Pro offers a cheaper, downloadable alternative today, with a real capability cost visible in the available index. July 27 should widen the evidence base for K3. It will not make the model easy for everyone to run.

For the other half of the K3 question, who built it, on whose data, and what a US government lab found when it tested models like it, see the companion piece, Kimi K3 isn't free, it just looks that way.

July 27 is when this review really begins. That is the day the weights are due, the technical report should expose more of the evidence behind Moonshot's claims, and well-equipped independent evaluators can begin testing the released artifact in their own harnesses. We will put it against its rivals then, hands on, and report back. Until the file exists, "the largest open model ever" is a headline Moonshot can write and no one else can yet grade.


Sources

  1. Moonshot AI, Kimi K3 announcement: architecture, availability, benchmark methodology, deployment guidance, and July 27 weight-release promise Inline ↗

  2. Moonshot API documentation, Kimi K3 pricing Inline ↗

  3. Artificial Analysis, Kimi K3: Intelligence Index, price, speed, and verbosity

Share:

Get this every weekday.

The Omniscient Bulletin: consequential AI, explained and evaluated. 5 to 7 items a day with the take, not the recap.

Discussion


Sign in to join the discussion.


Related

AI Policy

Vol. 1·Monday, July 20, 2026

Kimi K3 Isn't Free. It Just Looks That Way.

Kimi K3 is set to become the largest open-weight model ever released. What American labs say about how models like it are built, and what the US government found when it tested their predecessors, should give any American user pause.


Kimi K3 Isn't Free. It Just Looks That Way.

Moonshot's Kimi K3 is set to join DeepSeek's models and Alibaba's Qwen as a free download climbing the leaderboards. But between Anthropic and OpenAI's distillation findings, NIST's security testing, and Beijing's own moves to lock down its "open" frontier, the case for treating these weights as neutral technology is getting harder to make.


AI PolicyAI SecurityDeepSeek
Noah Ogbi8 min read
Continue →

AI Research

Vol. 1·Thursday, April 9, 2026

The Benchmark Racket: Why the Frontier Model Race Is Measuring the Wrong Thing

The Benchmark Racket: Why the Frontier Model Race Is Measuring the Wrong Thing

Six publicly available frontier models are clustered within 1.3 percentage points on the industry's most-cited coding benchmark. Meanwhile, a withheld model just scored 93.9% on the same test. The measurement system isn't broken - it's being gamed at two levels simultaneously.


Frontier ModelsResearch
Noah Ogbi13 min read
Continue →

AI Research

Vol. 1·Friday, April 3, 2026

Gemini 3.1 Pro Reviewed: Google's Reasoning Reversal


Gemini 3.1 Pro Reviewed: Google's Reasoning Reversal

Google DeepMind's Gemini 3.1 Pro arrived with the strongest independently verified reasoning scores of any frontier model. Three weeks later, GPT-5.4 changed the picture. A benchmark-by-benchmark assessment of where Gemini still leads, where it has fallen behind, and what the competitive gap actually looks like on verified data.


Frontier ModelsGoogle
Noah Ogbi16 min read
Continue →

Moonshot AI calls Kimi K3 the largest open-weight model ever built, and to make the case the company published a wall of benchmarks showing K3 ahead of several named rivals. Read two lines down in Moonshot's own announcement and the concession is right there: K3's "overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol"[1]. So the pitch is not that K3 is the best model on the market. It is Moonshot's claim to the best open model, pending a release that will let outside researchers inspect the artifact itself.

The catch is how it was measured: Moonshot selected the headline comparisons and reported many results through a mixture of its own and model-specific harnesses. The weights are not public yet. K3 is available through Moonshot's hosted products and API, but it cannot be independently run or audited locally. This is the read before the more complete review that July 27 may make possible.

What did Moonshot ship, and what did it hold back?

K3 is a 2.8-trillion-parameter mixture-of-experts model that routes each token through 16 of 896 experts, with a 1-million-token context window and native vision input[1]. Its API uses reasoning and, at launch, defaults to max thinking effort; Moonshot says low- and high-effort modes will follow in subsequent updates[1]. Moonshot prices it at $3 per million uncached input tokens and $15 per million output tokens, while cached input costs $0.30 per million tokens[3].

Moonshot says the full weights and a technical report with further detail on K3's architecture, training, and evaluations will arrive by July 27, 2026[1]. Until then, outside evaluators can inspect only the hosted model's behavior, not the artifact or training account behind it. That distinction matters: an API benchmark can be independently gathered, but it cannot establish what a released checkpoint will do under another team's inference stack, quantization, or agent scaffold.

A download will not make K3 locally practical for most teams. Moonshot recommends deploying it on supernode configurations with 64 or more accelerators[1]. The release should let well-equipped labs and inference providers test the weights under different setups. It will not turn a 2.8-trillion-parameter model into a laptop project.

Can you trust the benchmarks?

K3's headline scores should be read as vendor-reported results, not a neutral ranking. Moonshot's table mixes its KimiCode harness with Claude Code, Codex, and results drawn from external leaderboards; on several tests, it reports the best available rival score across harnesses[1]. Those choices do not make the results useless. They do make direct rank-order comparisons fragile, because agent scaffolding, token budgets, and tool configuration can move an outcome as much as the underlying model. We keep a field guide to the ways a benchmark score misleads.

The independent reads that exist are necessarily API-based: with no public weights, no one can run K3 locally in a neutral harness. Within that limit, the picture is strong. Artificial Analysis currently places K3 fourth of 186 models on its Intelligence Index, with a score of 57[4]. Arena.ai's Frontend Code Arena put it first, at 1,679 points, ahead of Claude Fable 5[5]. That is K3's clearest outside result, but it is one specialized arena, not a general verdict on coding.

There are caveats. Artificial Analysis calls K3 notably slow and very verbose, measuring 35.2 output tokens per second[4]. Developer Simon Willison recorded a simple SVG task that used 13,241 reasoning tokens and cost about 25 cents[6]. Moonshot's own architecture claims aim to improve efficiency at scale; the hosted behavior available today is what buyers have to judge.

A separate government evaluation offers a useful precedent, not a rule. When the US government's AI standards center tested DeepSeek V4 Pro, it recorded 74 percent on SWE-bench Verified and concluded:

"DeepSeek V4 scores better on DeepSeek's self-reported evaluations than on CAISI evaluations"[7].

Vendor-reported results and outside evaluations can diverge, even for the same model. K3's weights have not yet been released, so outside teams cannot test that artifact themselves.

Is it even the cheap option?

Ask where Kimi K3 fits and the word that comes back is cheap. The price tag only half-agrees.

Against the US frontier, K3's $3-in, $15-out list price sits at the workhorse tier, not the flagship one. It undercuts Opus 4.8 at $5/$25 and Fable 5 at $10/$50. Its list price matches Claude Sonnet 5's $3/$15, but through August 31, 2026, Sonnet's introductory rate is $2/$10[8]. Today, then, K3 costs 50 percent more than Sonnet on both input and output. The trade on offer is not frontier capability at a discount. It is a model Moonshot says trails the two leading proprietary systems, priced above a temporarily discounted mid-tier alternative.

DeepSeek V4 Pro makes the sharper open-model comparison. On Artificial Analysis's same Intelligence Index, K3 scores 57 while DeepSeek V4 Pro at max effort scores 44[4][12]. That is a meaningful capability gap in the available API evaluation, not a footnote to be hidden. OpenRouter lists DeepSeek at $0.435 per million input tokens and $0.87 per million output tokens, roughly seven times cheaper for input and seventeen times cheaper for output than K3[9].

DeepSeek's weights are publicly downloadable under an MIT license, and CAISI has evaluated the released model[10]. Artificial Analysis lists it as a 1.6-trillion-parameter mixture-of-experts model with 49 billion active parameters per token[12]. That does not settle a capability contest. It makes the tradeoff plain: K3 looks stronger in the available API tests; DeepSeek is cheaper, faster in Artificial Analysis's measurements, downloadable, and already exposed to outside testing.

Model

API price (in / out)

Weights available?

License

Independent evaluation on released weights?

Kimi K3

$3.00 / $15.00

Promised by July 27

Not yet disclosed

No

DeepSeek V4 Pro

$0.435 / $0.87

Yes

MIT

Yes, CAISI

K3 may prove more capable when its weights arrive. For now, the available API evaluations make it a serious contender, but they do not erase the capability-price-reproducibility tradeoff.

Inline ↗
  • Arena.ai, Kimi K3 Frontend Code Arena result Inline ↗

  • Simon Willison, Kimi K3 SVG-task token use and cost Inline ↗

  • NIST CAISI, Evaluation of DeepSeek V4 Pro Inline ↗

  • Anthropic Claude Platform, current Fable 5, Opus 4.8, and Sonnet 5 pricing Inline ↗

  • OpenRouter, DeepSeek V4 Pro listed input and output pricing Inline ↗

  • DeepSeek on Hugging Face, DeepSeek V4 Pro weights and MIT license Inline ↗

  • Kimi.ai statement: subscription capacity limits and planned reopening Inline ↗

  • Artificial Analysis, DeepSeek V4 Pro: Intelligence Index, throughput, and parameter details Inline ↗