Skip to content
See the World Through Science

Kimi K3 Is the Biggest Open-Weight Model Ever Built. The Benchmarks Say Near-Frontier, Not First Place.

AI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Rows of servers inside a data center, the kind of hardware used to train large AI models
A data center. Representative image for large-model training; not Moonshot AI's own hardware.Robert (Flickr) · CC-BY-2.0

When a new large language model launches, the release blog almost always comes with a table of benchmarks colored in its own favor. Kimi K3's did too. Moonshot AI, the Beijing lab behind the Kimi family, published a grid on July 16 showing its new model beating Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 on most of 35 tests. The launch was covered within hours by Bloomberg, CNBC, Fortune and a wall of others.

The more useful question is what happens when someone other than the vendor runs the tests. On that, the picture is clear and a little more sober: Kimi K3 is genuinely near the frontier, and genuinely not at the top of it.

What the outside boards found

Artificial Analysis, an independent evaluation firm that scores models on a battery it controls rather than one the labs supply, places Kimi K3 fourth out of 187 models on its Intelligence Index, with a score of 57. That is well above the field average of 31, and it puts K3 in elite company. It is also, by the same board's numbers, behind Anthropic's Claude Fable 5 (60) and OpenAI's GPT-5.6 Sol (59). Near the summit, in other words, but below the two models currently sitting on it.

On a separate, private long-horizon evaluation meant to mimic sustained knowledge work, Artificial Analysis clocked K3 at an Elo of 1547, a large jump over the earlier Kimi K2.6 and behind only Fable 5. The independent developer Simon Willison, who tested the model on release day, noted the same split most reviewers landed on: strong, clearly frontier-adjacent, and not the outright leader the vendor grid implies.

Coding is where K3 looks best. On Arena, the crowd-voted head-to-head board, it debuted at the top of the Frontend Code arena, ahead of even Fable 5 on that specific task. A single arena win is not a general ranking, but it lines up with the broader read that K3's real strength is agentic and coding work rather than a clean sweep across every category.

So the honest one-line summary is the reframe, not the launch headline: independent tests place Kimi K3 among the world’s best models, just behind the very top performers. The claim that it beats Opus 4.8 and GPT-5.5 is Moonshot's own; the independent placement is near-frontier, not first.

The engineering underneath

What makes the result notable is less the ranking than how it was reached. K3 is a high-sparsity mixture-of-experts model: of its 896 experts, only 16 fire on any given token, so although the model totals 2.8 trillion parameters, only a small fraction do the work of any single prediction. That is how a model this large stays affordable to run.

Two newer pieces are Moonshot's own. Kimi Delta Attention, the lab says, decodes up to 6.3 times faster on million-token contexts, and a technique it calls attention residuals lifts training efficiency by about 25% for under 2% added compute. The context window runs to a million tokens, roughly a shelf of books held in working memory at once. These figures come from Moonshot and have not been independently reproduced, so read them as vendor claims about the plumbing rather than settled results. The end-to-end scores above are what the outside boards actually measured.

"Open" is a promise, not yet a fact

The label doing a lot of work in the headlines is "open-weight." At 2.8 trillion parameters, K3 would be the largest openly released model to date, ahead of DeepSeek and other Chinese open models that have driven the field this year. That is a real milestone for anyone who wants to run or study a near-frontier model without a corporate gatekeeper.

The catch is timing. As of the July 16 launch, the weights were not downloadable. Moonshot has promised to release them by July 27. Until they land, "open" is a pledge, not something a researcher can act on. Every claim about openness here should carry that asterisk: the model is currently reachable only through Moonshot's API, and the open release is a date on a calendar, not a file you can fetch.

The price signal

There is one more shift worth flagging, because it cuts against the story the West has told itself about Chinese AI. Kimi K3 costs $3 per million input tokens and $15 per million output tokens, with a discounted $0.30 rate on cached input. That is a sharp jump from K2.6, which ran roughly $0.95 per million input tokens, and it puts K3 in the same price band as Anthropic's mid-range Claude Sonnet line rather than the deep-discount tier Chinese labs made their name on.

Compounding the cost, K3 currently ships with a single reasoning setting, "max," and no cheaper mode. One early tester watched it burn more than 13,000 reasoning tokens answering a trivial prompt. For simple queries, that combination of a premium price and a single high-effort mode makes the model expensive in a way its predecessors were not. The era of reflexively cheap frontier models from China may be closing, at least at the top of the range.

None of that dims the underlying achievement. A 2.8-trillion-parameter model that independent boards rank among the top handful in the world, built by a lab most Western readers had not heard of two years ago, is a real marker of how fast the gap has narrowed. It is just worth reading the scoreboard the outside evaluators kept, not the one that shipped with the press release.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Kimi K3 Is the Biggest Open-Weight Model Ever Built. The Benchmarks Say Near-Frontier, Not First Place.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.