Kimi K3 is Moonshot AI's 2.8 trillion parameter flagship with 104 billion active parameters, native multimodal input, and a one-million-token context window. Moonshot released the weights, technical report, and training infrastructure on July 27. The official API price is $3 per million cache-miss input tokens, $0.30 for cached input, and $15 for output. K3 is available through Kimi's app, Kimi Work, Kimi Code, API, and enterprise products. The benchmark screenshots below capture the launch period; product access and prices were rechecked against Moonshot's official pages on September 1.
- Moonshot AI announced Kimi K3 on July 17, 2026: 2.8 trillion parameters, 1 million token context, native multimodal input. The announcement post passed 23.5 million views on X.
- Moonshot released the weights, technical report, and training infrastructure on July 27, 2026 under the Kimi K3 License.
- Official API pricing, rechecked September 1: $3 per million cache-miss input tokens, $15 per million output, and $0.30 for cached input.
- Artificial Analysis scores K3 at 57 on its Intelligence Index, 4th of 186 models, behind Claude Fable 5 (60) and GPT-5.6 Sol (59).
- The same testing measured 35 output tokens per second, the slowest of the major flagships, and 130 million tokens generated during the evaluation against a 63 million average.
- Cost per task landed at $0.95, under GPT-5.6 Sol at $1.04 and well under Claude Fable 5 at $2.75, despite the verbosity.
- On Code Arena's WebDev leaderboard (483,895 total votes), K3 held rank 1 with a preliminary score of 1679, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618.
- K3 runs on kimi.com, Kimi Work 3.1 or later, Kimi Code, the Kimi API, and Kimi Enterprise. Consumer access uses plan credits; Moonshot does not document an unlimited free K3 allowance.
Kimi K2.6 made Moonshot the open-weight lab to watch. Kimi K3 is the first time one of its models walks into the frontier conversation and does not need the word "open" to justify being there. We captured the launch benchmarks and tested the app in July. For this September update, we treated those screenshots as historical evidence and checked every current price, availability, and model specification against Moonshot's own pages.
What Moonshot Shipped
A frontier-sized model whose weights are now public.
Moonshot announced K3 on July 17 with three architecture claims: Kimi Delta Attention for up to 6.3x faster decoding in million-token contexts, Attention Residuals for roughly 25% better training efficiency at under 2% extra cost, and a design brief of long-horizon agentic coding and self-evolving workflows. The model went live the same day on kimi.com, Kimi Work, Kimi Code, and the API.

The launch detail that mattered most was the date attached to the weights. K3 first shipped through Kimi's products and API, with a July 27 release commitment. Moonshot met it: the company published the weights, technical report, and training infrastructure on July 27. The screenshot below records the promise before it was fulfilled.

At 2.8 trillion parameters, with 104 billion active on each forward pass, K3 is data-center scale. The official model card lists a 1,048,576-token context window and native vision. The practical value of the release is hosting choice, research, fine-tuning, and downstream work rather than a normal desktop setup.
Moonshot's Own Coding Numbers
The vendor chart is unusually honest. K3 loses some of it.
Vendor benchmark charts usually show a clean sweep. Moonshot published one where its own model loses several panels, which is a reason to take the wins more seriously.

The honest read, panel by panel: K3 leads Program Bench (77.8) and SWE Marathon (42.0, with every model scoring low), and sits half a point behind GPT-5.6 Sol on Terminal Bench 2.1 (88.3 against 88.8). It trails Claude Fable 5 by five points on FrontierSWE (81.2 against 86.6) and sits third on DeepSWE behind both Sol and Fable 5. The line that builds the most trust: on Kimi Code Bench 2.0, Moonshot's own internal benchmark, Fable 5 beats K3, 76.9 to 72.9. A lab publishing a chart where a rival wins its home benchmark is rare, and it makes the rest of the chart easier to believe.
What Independent Testing Measured
Fourth of 186 models. Three points off the frontier.
Artificial Analysis, which runs the same evaluation suite across every major model, scores K3 at 57 on its Intelligence Index. That places it 4th of 186 models measured, behind Claude Fable 5 at 60 and GPT-5.6 Sol at 59, and ahead of Grok 4.5, GLM-5.2, and Gemini 3.6 Flash.

For an open-weight model, this is the closest the category has ever been to the closed frontier. The r/LocalLLaMA joke that China is now six days behind the west refers to the gap between GPT-5.6's launch and this result. Two years ago that gap was measured in model generations.
The Slow Part Nobody Puts in the Headline
35 tokens per second, and it never stops thinking.
The same testing measured K3 at 35 output tokens per second. That is the slowest of the major flagships, about half the speed of GPT-5.6 Sol and DeepSeek V4 Pro, and nearly eight times slower than Gemini 3.6 Flash.

K3 also always reasons before answering, and it reasons at length. Generating the answers for the Intelligence Index took K3 130 million tokens against a 63 million average for comparable models. In a chat that shows up as waiting. In an agent pipeline it shows up as latency per step, which compounds across a long run. If your workload is interactive and speed-sensitive, this single chart may decide the question for you.
Deciding between K3, Claude, and GPT-5.6? Get a personalized pick in 60 seconds.
The Arena Result Everyone Screenshotted
Rank 1 on WebDev, with a preliminary tag doing real work.
The screenshot that carried K3 across social media is Code Arena's WebDev leaderboard, where humans vote blind on front-end builds. K3 took rank 1 with a score of 1679, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618. The thread below hit 2,000 upvotes on r/LocalLLaMA in a day.

Two caveats before you repeat the headline. The score carries a Preliminary tag, and K3 had 1,757 votes against 2,505 for Fable 5 and 4,722 for GLM-5.2, so the ranking can move as votes accumulate. Arena votes also reward what demos well. The result is real, and it is still one leaderboard rather than a coronation.
Using K3 in Kimi's Products
Current access is documented by product and plan, not by an assumed free tier.
Moonshot lists K3 in the Kimi app and on kimi.com, Kimi Work 3.1 or later, Kimi Code, the Kimi API, and Kimi Enterprise. Kimi Code exposes K3 through its model selector, while Kimi's consumer help pages describe model use in credits. The official documentation does not promise a permanent unlimited free K3 tier.

The screenshot below is from our July test account after its quota was exhausted. It records what that account showed on that date; it is not proof of a current product-wide fallback rule. For a fresh evaluation, check the model label, available credits, and plan terms in your own session before judging K3.

Treat account screenshots as dated evidence
Kimi's plans and credit rules can change. This July screenshot shows one account state, not a promise about today's free access. Use Moonshot's current membership and product documentation for access, then confirm the model label in the response you receive.
Kimi K3 vs Claude Sonnet 5: Current API Prices
K3 is $3 / $15. Sonnet 5 stayed at $2 / $10.
K3's API rate is $3 per million cache-miss input tokens and $15 per million output, with cached input at $0.30. Claude Sonnet 5 is cheaper per token: Anthropic made its $2 input and $10 output rate permanent on August 10 instead of raising it in September. GPT-5.6 Sol is currently $4 input, $0.40 cached input, and $20 output under OpenAI's promotion through at least November 21.
Kimi K3 vs Claude Sonnet 5 vs GPT-5.6 Sol, September 1, 2026
| Kimi K3 | Claude Sonnet 5 | GPT-5.6 Sol | |
|---|---|---|---|
| Price per 1M in / out | $3 / $15 | $2 / $10 | $4 / $20 promotional |
| Context window | 1M | 1M | 1.05M |
| Intelligence Index | 57 (4th) | not in the top chart shown above | 59 (2nd) |
| Output speed | 35 tok/s | faster in practice, not measured here | 63 tok/s |
| Cost per task (measured) | $0.95 | $2.29 (see our Sonnet 5 review) | $1.04 |
| Open weights | released July 27 | no | no |
Source: Current prices from Moonshot, Anthropic, and OpenAI official pages, checked September 1, 2026. Cost per task and speed are launch-period measurements from Artificial Analysis.
Per-token price is only half the bill, because K3 is verbose. Artificial Analysis measured it at $0.95 per task on their workload, cheaper than GPT-5.6 Sol at $1.04 and about a third of Claude Fable 5 at $2.75. The verbosity eats some of the sticker advantage and still leaves K3 the cheapest of the three on measured work.

For a monthly figure: at 30 messages a day, roughly 0.45M input and 0.72M output tokens a month, K3 works out to about $12.15. Claude Sonnet 5 works out to about $8.10 at its permanent rate. You can run your own usage through our AI cost calculator, or answer six questions in the model picker and get a shortlist instead of a spreadsheet.
What r/LocalLLaMA Makes of It
The gap that used to be a year is now six days.
The community reaction is less about any single benchmark and more about the release calendar. GPT-5.6 shipped on July 9. K3 matched or passed it on several coding measures on July 17. The top comment in the 2,000-upvote thread did the math:

The replies under it are the practical wish list: excitement that the weights are coming, and hope that K3 traces get distilled into models small enough to run locally. That is the substance behind the joke. A frontier-adjacent model with released weights becomes training material for everything below it, which is why a 2.8T model nobody can self-host still matters to a subreddit about local models.
Who Should Actually Use Kimi K3
Agent builders and weight watchers yes, chat users probably not.
Pick K3 if you run long agentic coding jobs and bill by the task. The measured cost per task is the best of the frontier trio, the 1M context is real, and the benchmark profile leans toward multi-step autonomous work. Latency matters less when nobody is watching the tokens stream.
Pick it too if the weights are the point. Moonshot released them on July 27 with the technical report and training infrastructure. That gives teams a route to inspect and host the model under the Kimi K3 License, provided they have the infrastructure for a 2.8T-parameter mixture-of-experts model.
Skip it for interactive chat if latency matters more than long-horizon work. The launch-period speed test below measured 35 output tokens per second, so run a current trial on your own provider before committing. For a lower API rate, Claude Sonnet 5 is now $2 / $10 permanently. If you want the task-by-task breakdown across the whole lineup, our model guide covers it.
Sources
Everything above traces to one of these.
- Moonshot AI: Kimi K3 announcement, architecture, access, and API pricing
- Moonshot AI: Kimi K3 weights, report, and infrastructure release
- Moonshot AI: official Kimi K3 model card
- Moonshot AI: Kimi Code model availability
- Moonshot AI: membership and credit rules
- Anthropic API release notes: Sonnet 5 standard pricing
- OpenAI: GPT-5.6 Sol model and promotional pricing
- Artificial Analysis, Kimi K3 model page: Intelligence Index, speed, verbosity, and cost per task, captured July 23, 2026
- Code Arena WebDev leaderboard as of July 16, 2026: 483,895 votes across 98 models
- r/LocalLLaMA threads of July 17, 2026, including the 2,029-upvote Code Arena thread
- Our own kimi.com session on July 23, 2026, for the dated app screenshots; not used as proof of current plan behavior
FAQ
Is Kimi K3 free to use?
K3 is available on kimi.com and in Kimi products, but current consumer usage is credit-based and depends on the plan. Moonshot's official pages do not document a permanent unlimited free K3 allowance. The API is pay as you go at $3 input and $15 output per million tokens, with cached input at $0.30.
Is Kimi K3 better than Claude?
Close, cheaper, and slower rather than better. Independent testing has K3 three points behind Claude Fable 5 on intelligence, and Moonshot's own chart shows Fable 5 winning FrontierSWE and Moonshot's internal benchmark. K3 wins Program Bench, SWE Marathon, and the early Code Arena voting, at about a third of Fable 5's measured cost per task.
Are the Kimi K3 weights available?
Yes. Moonshot released the weights, technical report, and training infrastructure on July 27, 2026. The official model card lists 2.8 trillion total parameters, 104 billion active parameters, a 1,048,576-token context window, and the Kimi K3 License.
What hardware do you need to run Kimi K3 locally?
Realistically, more than any individual has. A 2.8 trillion parameter model is data-center scale even heavily quantized. The practical local story is downstream: distillations and fine-tunes built from the released weights, in the 30B range the community is already asking for.

Founder of Spectrum AI Labs — testing AI tools and models, and writing up what actually ships.
More about Paras →Not Sure Which Model Fits Your Stack?
We compare, test, and price the current AI models so you do not have to. Get a free 15-minute call and we will map your workload to the right one.



