Artificial Intelligence
10 min readJuly 23, 2026

Kimi K3 Review: An Open-Weight Flagship at Claude Sonnet Prices

Moonshot's 2.8T-parameter Kimi K3 lands three points behind Claude Fable 5 on independent testing at a fraction of the price, and it is also the slowest flagship measured. The honest review, with screenshots.

Paras Tiwari
Paras TiwariFounder, Spectrum AI Labs
Three benchmark charts comparing Kimi K3 with other frontier AI models on intelligence, speed, and cost per task

Get weekly AI tool reviews

We test tools so you don't have to. No spam.

TL;DR

Kimi K3 is Moonshot AI's new 2.8 trillion parameter flagship, announced July 17, 2026, with open weights promised by July 27. Independent testing puts it 4th of 186 models on intelligence, three points behind Claude Fable 5, at $3 input and $15 output per million tokens. That is exactly what Claude Sonnet 5 will charge from September. It tops the Code Arena WebDev leaderboard in early human voting and wins several of Moonshot's own agentic coding benchmarks. It is also the slowest flagship measured, one of the most verbose, and the free tier on kimi.com quietly downgrades you to K2.6 when your quota runs out. Every claim in this Kimi K3 review comes with a screenshot or a primary source.

Kimi K3, Verified
Updated July 23, 2026
  • Moonshot AI announced Kimi K3 on July 17, 2026: 2.8 trillion parameters, 1 million token context, native multimodal input. The announcement post passed 23.5 million views on X.
  • Full open weights are promised by July 27, 2026, which would make K3 the largest open-weight model released to date.
  • API pricing, checked on platform.kimi.ai on July 23: $3 per million input tokens, $15 per million output, $0.30 for cached input.
  • Artificial Analysis scores K3 at 57 on its Intelligence Index, 4th of 186 models, behind Claude Fable 5 (60) and GPT-5.6 Sol (59).
  • The same testing measured 35 output tokens per second, the slowest of the major flagships, and 130 million tokens generated during the evaluation against a 63 million average.
  • Cost per task landed at $0.95, under GPT-5.6 Sol at $1.04 and well under Claude Fable 5 at $2.75, despite the verbosity.
  • On Code Arena's WebDev leaderboard (483,895 total votes), K3 held rank 1 with a preliminary score of 1679, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618.
  • K3 runs on kimi.com, Kimi Work, Kimi Code, and the API. The free tier has a quota; when it runs out, the app falls back to K2.6.

Kimi K2.6 made Moonshot the open-weight lab to watch. Kimi K3 is the first time one of its models walks into the frontier conversation and does not need the word "open" to justify being there. We spent a day verifying the numbers, capturing the benchmarks first-hand, and running the app, including the part where it quietly stopped serving us K3. This review covers what held up.

2.8T
parameters
largest open-weight to date
57
intelligence index
4th of 186, Artificial Analysis
$3 / $15
per 1M tokens
Sonnet 5's standard rate
35
tokens per second
slowest flagship measured

What Moonshot Shipped

A frontier-sized model with a ten day countdown on its weights.

Moonshot announced K3 on July 17 with three architecture claims: Kimi Delta Attention for up to 6.3x faster decoding in million-token contexts, Attention Residuals for roughly 25% better training efficiency at under 2% extra cost, and a design brief of long-horizon agentic coding and self-evolving workflows. The model went live the same day on kimi.com, Kimi Work, Kimi Code, and the API.

Moonshot AI's Kimi K3 announcement post on X with 57.1K likes, listing 2.8 trillion parameters, 1 million context, and benchmark charts

The launch detail that matters most is the one with a date on it. K3 shipped as API and app only, with the full weights promised by July 27, 2026. The community noticed, and the receipts are already pinned in the top threads:

Reddit comment with 394 upvotes quoting Moonshot's commitment that the full model weights will be released by July 27, 2026

At roughly 2.8 trillion parameters, a full release would make K3 the largest open-weight model ever published. Nobody is running that at home. The point of weights at this scale is hosting choice, fine-tuning rights, and distillation into smaller models, and the community is already planning for exactly that.

Moonshot's Own Coding Numbers

The vendor chart is unusually honest. K3 loses some of it.

Vendor benchmark charts usually show a clean sweep. Moonshot published one where its own model loses several panels, which is a reason to take the wins more seriously.

Moonshot's official coding benchmark chart comparing Kimi K3 with GPT-5.6 Sol, Claude Fable 5, Opus 4.8, GPT-5.5, and GLM-5.2 across six coding benchmarks

The honest read, panel by panel: K3 leads Program Bench (77.8) and SWE Marathon (42.0, with every model scoring low), and sits half a point behind GPT-5.6 Sol on Terminal Bench 2.1 (88.3 against 88.8). It trails Claude Fable 5 by five points on FrontierSWE (81.2 against 86.6) and sits third on DeepSWE behind both Sol and Fable 5. The line that builds the most trust: on Kimi Code Bench 2.0, Moonshot's own internal benchmark, Fable 5 beats K3, 76.9 to 72.9. A lab publishing a chart where a rival wins its home benchmark is rare, and it makes the rest of the chart easier to believe.

What Independent Testing Measured

Fourth of 186 models. Three points off the frontier.

Artificial Analysis, which runs the same evaluation suite across every major model, scores K3 at 57 on its Intelligence Index. That places it 4th of 186 models measured, behind Claude Fable 5 at 60 and GPT-5.6 Sol at 59, and ahead of Grok 4.5, GLM-5.2, and Gemini 3.6 Flash.

Artificial Analysis Intelligence Index bar chart showing Kimi K3 at 57, third bar after Claude Fable 5 at 60 and GPT-5.6 Sol at 59

For an open-weight model, this is the closest the category has ever been to the closed frontier. The r/LocalLLaMA joke that China is now six days behind the west refers to the gap between GPT-5.6's launch and this result. Two years ago that gap was measured in model generations.

The Slow Part Nobody Puts in the Headline

35 tokens per second, and it never stops thinking.

The same testing measured K3 at 35 output tokens per second. That is the slowest of the major flagships, about half the speed of GPT-5.6 Sol and DeepSeek V4 Pro, and nearly eight times slower than Gemini 3.6 Flash.

Artificial Analysis output speed chart showing Kimi K3 last at 35 tokens per second while Gemini 3.6 Flash leads at 271

K3 also always reasons before answering, and it reasons at length. Generating the answers for the Intelligence Index took K3 130 million tokens against a 63 million average for comparable models. In a chat that shows up as waiting. In an agent pipeline it shows up as latency per step, which compounds across a long run. If your workload is interactive and speed-sensitive, this single chart may decide the question for you.

Deciding between K3, Claude, and GPT-5.6? Get a personalized pick in 60 seconds.

12 models · Personalized picks · 60 seconds

Take the Free Quiz

The Arena Result Everyone Screenshotted

Rank 1 on WebDev, with a preliminary tag doing real work.

The screenshot that carried K3 across social media is Code Arena's WebDev leaderboard, where humans vote blind on front-end builds. K3 took rank 1 with a score of 1679, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618. The thread below hit 2,000 upvotes on r/LocalLLaMA in a day.

Reddit post with 2,000 upvotes showing the Code Arena WebDev leaderboard with kimi-k3 at rank 1 above claude-fable-5 and gpt-5.6-sol

Two caveats before you repeat the headline. The score carries a Preliminary tag, and K3 had 1,757 votes against 2,505 for Fable 5 and 4,722 for GLM-5.2, so the ranking can move as votes accumulate. Arena votes also reward what demos well. The result is real, and it is still one leaderboard rather than a coronation.

Using K3 on Kimi.com, Including the Catch

Three tiers, a thinking dial, and a silent downgrade.

On kimi.com, the model picker offers K2.6 for fast chat, K3 as the flagship all-rounder, and K3 Swarm for batch and search-heavy jobs, plus a thinking-effort setting. The free tier includes K3 under a usage quota.

Kimi.com model picker showing K2.6, K3 described as flagship all-rounder, K3 Swarm, and a thinking effort setting

Here is the catch we ran into ourselves. Our test account's free quota was exhausted, and the composer still let us select K3 High and submit a prompt. The response came back fast and looked fine. The model label on the reply said K2.6. The app had downgraded the request without an error, a modal, or any visible warning beyond the quota banner.

Kimi.com composer with the free quota banner visible and K3 High selected as the model

Check the label before you judge the model

If you are evaluating K3 on the free tier, verify the model tag on each response before drawing conclusions. With the quota spent, kimi.com serves K2.6 while the picker still shows K3 selected. Plenty of "K3 is underwhelming" takes are probably K2.6 reviews.

Kimi K3 vs Claude Sonnet 5: The Pricing Coincidence

$3 in, $15 out. The exact number Claude lands on in September.

K3's API rate is $3 per million input tokens and $15 per million output, with cached input at $0.30. Claude Sonnet 5 currently runs an introductory $2 and $10, and moves to $3 and $15 on September 1. From that day, the open-weight flagship and Anthropic's everyday default cost the same per token.

Kimi K3 vs Claude Sonnet 5 vs GPT-5.6 Sol, July 23, 2026

Kimi K3Claude Sonnet 5GPT-5.6 Sol
Price per 1M in / out$3 / $15$2 / $10 until Aug 31, then $3 / $15$5 / $30
Context window1M1M1.1M (per Code Arena listing)
Intelligence Index57 (4th)not in the top chart shown above59 (2nd)
Output speed35 tok/sfaster in practice, not measured here63 tok/s
Cost per task (measured)$0.95$2.29 (see our Sonnet 5 review)$1.04
Open weightspromised July 27nono

Source: Pricing from platform.kimi.ai and Anthropic's published rates, July 23, 2026. Cost per task and speed from Artificial Analysis. Blank cells stay blank rather than guessed.

Per-token price is only half the bill, because K3 is verbose. Artificial Analysis measured it at $0.95 per task on their workload, cheaper than GPT-5.6 Sol at $1.04 and about a third of Claude Fable 5 at $2.75. The verbosity eats some of the sticker advantage and still leaves K3 the cheapest of the three on measured work.

Artificial Analysis cost per task chart showing Kimi K3 at $0.95, GPT-5.6 Sol at $1.04, and Claude Fable 5 at $2.75

For a monthly figure: at 30 messages a day, roughly 0.45M input and 0.72M output tokens a month, K3 works out to about $12.15. Claude Sonnet 5 costs $8.10 on intro pricing today and the same $12.15 from September. You can run your own usage through our AI cost calculator, which we updated with K3 and the rest of the July lineup, or answer six questions in the model picker and get a shortlist instead of a spreadsheet.

What r/LocalLLaMA Makes of It

The gap that used to be a year is now six days.

The community reaction is less about any single benchmark and more about the release calendar. GPT-5.6 shipped on July 9. K3 matched or passed it on several coding measures on July 17. The top comment in the 2,000-upvote thread did the math:

Top Reddit comment with 872 upvotes reading: So China is now 6 days behind the west

The replies under it are the practical wish list: excitement that the weights are coming, and hope that K3 traces get distilled into models small enough to run locally. That is the substance behind the joke. A frontier-adjacent model with released weights becomes training material for everything below it, which is why a 2.8T model nobody can self-host still matters to a subreddit about local models.

Who Should Actually Use Kimi K3

Agent builders and weight watchers yes, chat users probably not.

Pick K3 if you run long agentic coding jobs and bill by the task. The measured cost per task is the best of the frontier trio, the 1M context is real, and the benchmark profile leans toward multi-step autonomous work. Latency matters less when nobody is watching the tokens stream.

Pick it too if the weights are the point. If Moonshot delivers on July 27, K3 becomes the strongest open-weight base ever released, with everything that implies for hosting choice and for the smaller distilled models that will follow. Our open-source coding model ranking is due a rewrite the day that happens.

Skip it for interactive chat. At 35 tokens per second with always-on reasoning, K3 feels slow in a way benchmark tables do not show, and the free-tier quota plus the silent K2.6 fallback make casual evaluation misleading. For everyday chat the honest comparison is Claude Sonnet 5, which is cheaper until September and far snappier, or GPT-5.6 Sol inside ChatGPT Plus. If you want the task-by-task breakdown across the whole lineup, our model guide covers it.

Sources

Everything above traces to one of these.

  • Moonshot AI's K3 announcement on X, July 17, 2026, and the linked tech blog at kimi.com/blog/kimi-k3
  • platform.kimi.ai pricing documentation, checked July 23, 2026
  • Artificial Analysis, Kimi K3 model page: Intelligence Index, speed, verbosity, and cost per task, captured July 23, 2026
  • Code Arena WebDev leaderboard as of July 16, 2026: 483,895 votes across 98 models
  • r/LocalLLaMA threads of July 17, 2026, including the 2,029-upvote Code Arena thread
  • Our own kimi.com session on July 23, 2026, for the app screenshots and the K2.6 fallback behavior

FAQ

Is Kimi K3 free to use?

Partly. kimi.com includes K3 on the free tier under a usage quota. When the quota runs out, the app keeps accepting prompts with K3 selected but serves them with K2.6 until the quota refreshes or you upgrade. The API is pay as you go at $3 input and $15 output per million tokens, with no subscription required.

Is Kimi K3 better than Claude?

Close, cheaper, and slower rather than better. Independent testing has K3 three points behind Claude Fable 5 on intelligence, and Moonshot's own chart shows Fable 5 winning FrontierSWE and Moonshot's internal benchmark. K3 wins Program Bench, SWE Marathon, and the early Code Arena voting, at about a third of Fable 5's measured cost per task.

When do the Kimi K3 weights come out?

Moonshot committed to releasing the full weights by July 27, 2026. Until then K3 is open weight in promise only: you can use it through the app and API, and the 2.8 trillion parameter checkpoint itself is not yet downloadable.

What hardware do you need to run Kimi K3 locally?

Realistically, more than any individual has. A 2.8 trillion parameter model is data-center scale even heavily quantized. The practical local story is downstream: distillations and fine-tunes built from the released weights, in the 30B range the community is already asking for.

Paras Tiwari
Written by
Paras Tiwari
Founder, Spectrum AI Labs

Founder of Spectrum AI Labs — testing AI tools and models, and writing up what actually ships.

More about Paras →

Not Sure Which Model Fits Your Stack?

We compare, test, and price the current AI models so you do not have to. Get a free 15-minute call and we will map your workload to the right one.