Artificial Intelligence
10 min readJuly 23, 2026

Kimi K3 Review: Open Weights, 1M Context, $3/$15

Moonshot released K3's 2.8T weights, technical report, and training infrastructure. Current pricing and access, with the launch-period benchmark screenshots kept in context.

Paras Tiwari
Paras TiwariFounder, Spectrum AI Labs
Three benchmark charts comparing Kimi K3 with other frontier AI models on intelligence, speed, and cost per task

Get weekly AI tool reviews

We test tools so you don't have to. No spam.

TL;DR

Kimi K3 is Moonshot AI's 2.8 trillion parameter flagship with 104 billion active parameters, native multimodal input, and a one-million-token context window. Moonshot released the weights, technical report, and training infrastructure on July 27. The official API price is $3 per million cache-miss input tokens, $0.30 for cached input, and $15 for output. K3 is available through Kimi's app, Kimi Work, Kimi Code, API, and enterprise products. The benchmark screenshots below capture the launch period; product access and prices were rechecked against Moonshot's official pages on September 1.

Kimi K3, Verified
Updated September 1, 2026
  • Moonshot AI announced Kimi K3 on July 17, 2026: 2.8 trillion parameters, 1 million token context, native multimodal input. The announcement post passed 23.5 million views on X.
  • Moonshot released the weights, technical report, and training infrastructure on July 27, 2026 under the Kimi K3 License.
  • Official API pricing, rechecked September 1: $3 per million cache-miss input tokens, $15 per million output, and $0.30 for cached input.
  • Artificial Analysis scores K3 at 57 on its Intelligence Index, 4th of 186 models, behind Claude Fable 5 (60) and GPT-5.6 Sol (59).
  • The same testing measured 35 output tokens per second, the slowest of the major flagships, and 130 million tokens generated during the evaluation against a 63 million average.
  • Cost per task landed at $0.95, under GPT-5.6 Sol at $1.04 and well under Claude Fable 5 at $2.75, despite the verbosity.
  • On Code Arena's WebDev leaderboard (483,895 total votes), K3 held rank 1 with a preliminary score of 1679, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618.
  • K3 runs on kimi.com, Kimi Work 3.1 or later, Kimi Code, the Kimi API, and Kimi Enterprise. Consumer access uses plan credits; Moonshot does not document an unlimited free K3 allowance.

Kimi K2.6 made Moonshot the open-weight lab to watch. Kimi K3 is the first time one of its models walks into the frontier conversation and does not need the word "open" to justify being there. We captured the launch benchmarks and tested the app in July. For this September update, we treated those screenshots as historical evidence and checked every current price, availability, and model specification against Moonshot's own pages.

2.8T
total parameters
104B active per forward pass
57
intelligence index
4th of 186, Artificial Analysis
$3 / $15
per 1M tokens
$0.30 cached input
35
tokens per second
slowest flagship measured

What Moonshot Shipped

A frontier-sized model whose weights are now public.

Moonshot announced K3 on July 17 with three architecture claims: Kimi Delta Attention for up to 6.3x faster decoding in million-token contexts, Attention Residuals for roughly 25% better training efficiency at under 2% extra cost, and a design brief of long-horizon agentic coding and self-evolving workflows. The model went live the same day on kimi.com, Kimi Work, Kimi Code, and the API.

Moonshot AI's Kimi K3 announcement post on X with 57.1K likes, listing 2.8 trillion parameters, 1 million context, and benchmark charts

The launch detail that mattered most was the date attached to the weights. K3 first shipped through Kimi's products and API, with a July 27 release commitment. Moonshot met it: the company published the weights, technical report, and training infrastructure on July 27. The screenshot below records the promise before it was fulfilled.

Reddit comment with 394 upvotes quoting Moonshot's commitment that the full model weights will be released by July 27, 2026

At 2.8 trillion parameters, with 104 billion active on each forward pass, K3 is data-center scale. The official model card lists a 1,048,576-token context window and native vision. The practical value of the release is hosting choice, research, fine-tuning, and downstream work rather than a normal desktop setup.

Moonshot's Own Coding Numbers

The vendor chart is unusually honest. K3 loses some of it.

Vendor benchmark charts usually show a clean sweep. Moonshot published one where its own model loses several panels, which is a reason to take the wins more seriously.

Moonshot's official coding benchmark chart comparing Kimi K3 with GPT-5.6 Sol, Claude Fable 5, Opus 4.8, GPT-5.5, and GLM-5.2 across six coding benchmarks

The honest read, panel by panel: K3 leads Program Bench (77.8) and SWE Marathon (42.0, with every model scoring low), and sits half a point behind GPT-5.6 Sol on Terminal Bench 2.1 (88.3 against 88.8). It trails Claude Fable 5 by five points on FrontierSWE (81.2 against 86.6) and sits third on DeepSWE behind both Sol and Fable 5. The line that builds the most trust: on Kimi Code Bench 2.0, Moonshot's own internal benchmark, Fable 5 beats K3, 76.9 to 72.9. A lab publishing a chart where a rival wins its home benchmark is rare, and it makes the rest of the chart easier to believe.

What Independent Testing Measured

Fourth of 186 models. Three points off the frontier.

Artificial Analysis, which runs the same evaluation suite across every major model, scores K3 at 57 on its Intelligence Index. That places it 4th of 186 models measured, behind Claude Fable 5 at 60 and GPT-5.6 Sol at 59, and ahead of Grok 4.5, GLM-5.2, and Gemini 3.6 Flash.

Artificial Analysis Intelligence Index bar chart showing Kimi K3 at 57, third bar after Claude Fable 5 at 60 and GPT-5.6 Sol at 59

For an open-weight model, this is the closest the category has ever been to the closed frontier. The r/LocalLLaMA joke that China is now six days behind the west refers to the gap between GPT-5.6's launch and this result. Two years ago that gap was measured in model generations.

The Slow Part Nobody Puts in the Headline

35 tokens per second, and it never stops thinking.

The same testing measured K3 at 35 output tokens per second. That is the slowest of the major flagships, about half the speed of GPT-5.6 Sol and DeepSeek V4 Pro, and nearly eight times slower than Gemini 3.6 Flash.

Artificial Analysis output speed chart showing Kimi K3 last at 35 tokens per second while Gemini 3.6 Flash leads at 271

K3 also always reasons before answering, and it reasons at length. Generating the answers for the Intelligence Index took K3 130 million tokens against a 63 million average for comparable models. In a chat that shows up as waiting. In an agent pipeline it shows up as latency per step, which compounds across a long run. If your workload is interactive and speed-sensitive, this single chart may decide the question for you.

Deciding between K3, Claude, and GPT-5.6? Get a personalized pick in 60 seconds.

Take the Free Quiz

The Arena Result Everyone Screenshotted

Rank 1 on WebDev, with a preliminary tag doing real work.

The screenshot that carried K3 across social media is Code Arena's WebDev leaderboard, where humans vote blind on front-end builds. K3 took rank 1 with a score of 1679, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618. The thread below hit 2,000 upvotes on r/LocalLLaMA in a day.

Reddit post with 2,000 upvotes showing the Code Arena WebDev leaderboard with kimi-k3 at rank 1 above claude-fable-5 and gpt-5.6-sol

Two caveats before you repeat the headline. The score carries a Preliminary tag, and K3 had 1,757 votes against 2,505 for Fable 5 and 4,722 for GLM-5.2, so the ranking can move as votes accumulate. Arena votes also reward what demos well. The result is real, and it is still one leaderboard rather than a coronation.

Using K3 in Kimi's Products

Current access is documented by product and plan, not by an assumed free tier.

Moonshot lists K3 in the Kimi app and on kimi.com, Kimi Work 3.1 or later, Kimi Code, the Kimi API, and Kimi Enterprise. Kimi Code exposes K3 through its model selector, while Kimi's consumer help pages describe model use in credits. The official documentation does not promise a permanent unlimited free K3 tier.

Kimi.com model picker showing K2.6, K3 described as flagship all-rounder, K3 Swarm, and a thinking effort setting

The screenshot below is from our July test account after its quota was exhausted. It records what that account showed on that date; it is not proof of a current product-wide fallback rule. For a fresh evaluation, check the model label, available credits, and plan terms in your own session before judging K3.

Kimi.com composer with the free quota banner visible and K3 High selected as the model

Treat account screenshots as dated evidence

Kimi's plans and credit rules can change. This July screenshot shows one account state, not a promise about today's free access. Use Moonshot's current membership and product documentation for access, then confirm the model label in the response you receive.

Kimi K3 vs Claude Sonnet 5: Current API Prices

K3 is $3 / $15. Sonnet 5 stayed at $2 / $10.

K3's API rate is $3 per million cache-miss input tokens and $15 per million output, with cached input at $0.30. Claude Sonnet 5 is cheaper per token: Anthropic made its $2 input and $10 output rate permanent on August 10 instead of raising it in September. GPT-5.6 Sol is currently $4 input, $0.40 cached input, and $20 output under OpenAI's promotion through at least November 21.

Kimi K3 vs Claude Sonnet 5 vs GPT-5.6 Sol, September 1, 2026

Kimi K3Claude Sonnet 5GPT-5.6 Sol
Price per 1M in / out$3 / $15$2 / $10$4 / $20 promotional
Context window1M1M1.05M
Intelligence Index57 (4th)not in the top chart shown above59 (2nd)
Output speed35 tok/sfaster in practice, not measured here63 tok/s
Cost per task (measured)$0.95$2.29 (see our Sonnet 5 review)$1.04
Open weightsreleased July 27nono

Source: Current prices from Moonshot, Anthropic, and OpenAI official pages, checked September 1, 2026. Cost per task and speed are launch-period measurements from Artificial Analysis.

Per-token price is only half the bill, because K3 is verbose. Artificial Analysis measured it at $0.95 per task on their workload, cheaper than GPT-5.6 Sol at $1.04 and about a third of Claude Fable 5 at $2.75. The verbosity eats some of the sticker advantage and still leaves K3 the cheapest of the three on measured work.

Artificial Analysis cost per task chart showing Kimi K3 at $0.95, GPT-5.6 Sol at $1.04, and Claude Fable 5 at $2.75

For a monthly figure: at 30 messages a day, roughly 0.45M input and 0.72M output tokens a month, K3 works out to about $12.15. Claude Sonnet 5 works out to about $8.10 at its permanent rate. You can run your own usage through our AI cost calculator, or answer six questions in the model picker and get a shortlist instead of a spreadsheet.

What r/LocalLLaMA Makes of It

The gap that used to be a year is now six days.

The community reaction is less about any single benchmark and more about the release calendar. GPT-5.6 shipped on July 9. K3 matched or passed it on several coding measures on July 17. The top comment in the 2,000-upvote thread did the math:

Top Reddit comment with 872 upvotes reading: So China is now 6 days behind the west

The replies under it are the practical wish list: excitement that the weights are coming, and hope that K3 traces get distilled into models small enough to run locally. That is the substance behind the joke. A frontier-adjacent model with released weights becomes training material for everything below it, which is why a 2.8T model nobody can self-host still matters to a subreddit about local models.

Who Should Actually Use Kimi K3

Agent builders and weight watchers yes, chat users probably not.

Pick K3 if you run long agentic coding jobs and bill by the task. The measured cost per task is the best of the frontier trio, the 1M context is real, and the benchmark profile leans toward multi-step autonomous work. Latency matters less when nobody is watching the tokens stream.

Pick it too if the weights are the point. Moonshot released them on July 27 with the technical report and training infrastructure. That gives teams a route to inspect and host the model under the Kimi K3 License, provided they have the infrastructure for a 2.8T-parameter mixture-of-experts model.

Skip it for interactive chat if latency matters more than long-horizon work. The launch-period speed test below measured 35 output tokens per second, so run a current trial on your own provider before committing. For a lower API rate, Claude Sonnet 5 is now $2 / $10 permanently. If you want the task-by-task breakdown across the whole lineup, our model guide covers it.

Sources

Everything above traces to one of these.

FAQ

Is Kimi K3 free to use?

K3 is available on kimi.com and in Kimi products, but current consumer usage is credit-based and depends on the plan. Moonshot's official pages do not document a permanent unlimited free K3 allowance. The API is pay as you go at $3 input and $15 output per million tokens, with cached input at $0.30.

Is Kimi K3 better than Claude?

Close, cheaper, and slower rather than better. Independent testing has K3 three points behind Claude Fable 5 on intelligence, and Moonshot's own chart shows Fable 5 winning FrontierSWE and Moonshot's internal benchmark. K3 wins Program Bench, SWE Marathon, and the early Code Arena voting, at about a third of Fable 5's measured cost per task.

Are the Kimi K3 weights available?

Yes. Moonshot released the weights, technical report, and training infrastructure on July 27, 2026. The official model card lists 2.8 trillion total parameters, 104 billion active parameters, a 1,048,576-token context window, and the Kimi K3 License.

What hardware do you need to run Kimi K3 locally?

Realistically, more than any individual has. A 2.8 trillion parameter model is data-center scale even heavily quantized. The practical local story is downstream: distillations and fine-tunes built from the released weights, in the 30B range the community is already asking for.

Paras Tiwari
Written by
Paras Tiwari
Founder, Spectrum AI Labs

Founder of Spectrum AI Labs — testing AI tools and models, and writing up what actually ships.

More about Paras →

Not Sure Which Model Fits Your Stack?

We compare, test, and price the current AI models so you do not have to. Get a free 15-minute call and we will map your workload to the right one.