Kimi AI Statistics 2026
By Axis Intelligence Research
Co-author: Sarah Mitchell | Last updated: August 7, 2026 | License: CC BY 4.0
Kimi K3 processed 1.41 trillion tokens on OpenRouter in its first 22 days of availability — 76% of every token the entire Kimi model family has ever routed through the platform. Moonshot AI’s 2.8-trillion-parameter open-weight flagship now scores 57 on the Artificial Analysis Intelligence Index, the highest of any downloadable model, at $3.00 per million input tokens.
Quick Answer
Kimi is the model family built by Beijing-based Moonshot AI, and 2026 is the year it stopped being a cheap alternative and became a frontier competitor. Kimi K3, released July 16, 2026, is a 2.8T-parameter Mixture-of-Experts model activating 104B parameters across 16 of 896 experts, with a 1,048,576-token context window. It scores 57 on the Artificial Analysis Intelligence Index — ahead of every other open-weight model — and has been downloaded 1,258,043 times from Hugging Face in the trailing month. According to Axis Intelligence Research, K3 accumulated OpenRouter production traffic at 64.1 billion tokens per day, roughly 40× the daily rate of its predecessor Kimi K2.6. Moonshot’s annualized recurring revenue reached $300 million in June 2026 against a reported $30 billion valuation.
Key Findings
- According to Axis Intelligence Research’s compilation of OpenRouter author-page data retrieved August 7, 2026, Kimi K3 has processed 1.41 trillion tokens since its July 16 listing — 76.0% of the 1.86 trillion tokens processed across all seven Kimi models ever listed on the platform.
- Kimi K3 scores 57 on the Artificial Analysis Intelligence Index, ahead of GLM-5.2 (51) and DeepSeek V4 Pro (44), per Artificial Analysis’s July 17, 2026 evaluation.
- According to Axis Intelligence Research, Kimi K3 registers 76.9 on the Axis Open-Weight Diffusion Index (OWDI™) at its August 7, 2026 baseline reading — second behind GLM-5.2 at 95.9, because K3’s 1.4-terabyte weight footprint suppresses self-hosting and derivative creation despite category-leading capability.
- Moonshot AI prices Kimi K3 at $3.00 per million cache-miss input tokens, $0.30 per million cache-hit input tokens and $15.00 per million output tokens, flat across the full 1M-token context window, per the company’s official pricing page.
- According to Axis Intelligence Research’s cross-source calculation, Kimi K3 buys 1.08 percentage points of factual accuracy for every additional percentage point of hallucination it introduces relative to Kimi K2.6 — accuracy rose from 33% to 46% while hallucination rate rose from 39% to 51% on Artificial Analysis’s AA-Omniscience evaluation.
What Is Kimi K3 and How Is It Built?
Kimi K3 is Moonshot AI’s fourth-generation flagship and, per the company’s Kimi K3 technical blog, the first open model to reach 2.8 trillion parameters. The company states that for nine of the twelve months preceding launch, Kimi models set the upper bound of open-model sizes.
The architecture is the story. K3 replaces uniform residual accumulation with Attention Residuals (AttnRes) and layers Kimi Delta Attention (KDA) underneath — 69 KDA layers interleaved with 24 Gated MLA layers across 93 total layers. Sparsity was pushed hard: 16 active experts out of 896, under a Stable LatentMoE framework that Moonshot says delivers roughly 2.5× better scaling efficiency than Kimi K2.
Kimi K3 specification table
| Attribute | Value | Source |
|---|---|---|
| Total parameters | 2.8T | Moonshot AI model card, Jul 2026 |
| Activated parameters | 104B | Moonshot AI model card, Jul 2026 |
| Layers | 93 (1 dense) | Moonshot AI model card, Jul 2026 |
| Experts / selected per token | 896 / 16 (+2 shared) | Moonshot AI model card, Jul 2026 |
| Attention composition | 69 KDA + 24 Gated MLA | Moonshot AI model card, Jul 2026 |
| Attention heads | 96 | Moonshot AI model card, Jul 2026 |
| Context window | 1,048,576 tokens | Moonshot AI model card, Jul 2026 |
| Vocabulary | 160K | Moonshot AI model card, Jul 2026 |
| Vision encoder | MoonViT-V2, 401M params | Moonshot AI model card, Jul 2026 |
| Quantization | MXFP4 weights / MXFP8 activations (QAT) | Moonshot AI model card, Jul 2026 |
| License | Kimi K3 License | Hugging Face model card, Jul 2026 |
The quantization line matters more than it looks. K3 applies quantization-aware training from the supervised fine-tuning stage onward, which is why the published weights on Hugging Face ship at 8-bit precision rather than as a post-hoc compression of a BF16 checkpoint. Moonshot recommends serving on supernode configurations of 64 or more accelerators.
How Fast Is Kimi Being Adopted in Production?
Benchmark scores measure capability. Token counts measure whether anyone is actually paying to run the thing. OpenRouter publishes cumulative tokens processed per model, which makes it the cleanest public proxy for third-party production demand — and, because listing dates are published alongside, it can be normalized into a rate.
Kimi router velocity by model generation
Axis Intelligence Research compiled the following series from OpenRouter author-page data retrieved August 7, 2026. Velocity = cumulative tokens ÷ days since the model’s OpenRouter listing date.
| Model | OpenRouter listing | Days live | Cumulative tokens | Tokens/day | Source |
|---|---|---|---|---|---|
| Kimi K3 | Jul 16, 2026 | 22 | 1.41T | 64.09B | OpenRouter (Axis calculation) |
| Kimi K2.7 Code | Jun 12, 2026 | 56 | 91.8B | 1.64B | OpenRouter (Axis calculation) |
| Kimi K2.6 | Apr 20, 2026 | 109 | 175B | 1.61B | OpenRouter (Axis calculation) |
| Kimi K2.5 | Jan 27, 2026 | 192 | 155B | 0.81B | OpenRouter (Axis calculation) |
| Kimi K2 0905 | Sep 4, 2025 | 337 | 12.8B | 0.038B | OpenRouter (Axis calculation) |
| Kimi K2 Thinking | Nov 6, 2025 | 274 | 8.37B | 0.031B | OpenRouter (Axis calculation) |
| Kimi K2 0711 | Jul 11, 2025 | 392 | 3.43B | 0.009B | OpenRouter (Axis calculation) |
The discontinuity is not incremental. According to Axis Intelligence Research, Kimi K3’s daily token rate runs 39.9× that of Kimi K2.6, the fastest of the prior generations, and K3 alone accounts for 76.0% of the 1.856 trillion tokens the Kimi family has cumulatively processed on OpenRouter since July 2025.
Two caveats belong on that number rather than buried later. K3’s 22-day window contains its launch spike, and cumulative totals for older models were accumulated while those models were current — a fair comparison of steady-state demand will require a second reading in October. That is exactly why this series is dated and will be re-measured.
Hugging Face weight downloads by Kimi model
| Model | Params | Downloads (trailing 30 days) | Likes | Source |
|---|---|---|---|---|
| Kimi-K3 | 2.8T | 1,258,043 | 10,200 | Hugging Face |
| Kimi-K2.5 | 1.1T | 929,000 | 2,860 | Hugging Face |
| Kimi-K2.6 | 1.1T | 856,000 | 1,590 | Hugging Face |
| Kimi-K2.7-Code | 1T | 675,000 | 1,350 | Hugging Face |
| Kimi-VL-A3B-Thinking | 16B | 182,000 | 450 | Hugging Face |
| Kimi-K2-Instruct | 1T | 181,000 | 2,370 | Hugging Face |
| Moonlight-16B-A3B-Instruct | 16B | 44,000 | 202 | Hugging Face |
| Moonlight-16B-A3B | 16B | 13,800 | 116 | Hugging Face |
| Kimi-VL-A3B-Thinking-2506 | 16B | 8,900 | 373 | Hugging Face |
| Kimi-K2-Base | 1T | 7,530 | 306 | Hugging Face |
According to Axis Intelligence Research’s summation of Moonshot’s ten most-downloaded Hugging Face repositories, the Kimi family drew 4,155,273 weight downloads in the trailing 30 days to August 7, 2026, with K3 accounting for 30.3% of that total despite being available for only eleven of those days.
Sarah Mitchell: The download table is the one people will misread. K3 is not winning self-hosting — it is winning curiosity. A 2.8T model at 8-bit is roughly 1.4 TB of weights; almost nobody downloading it is going to serve it, and the community quantization repos openly say llama.cpp cannot yet load the
kimi_k3architecture at all. Compare that to K2.5, an 1.1T model still pulling 929,000 downloads five months after release from people who can actually run it on rented eight-GPU nodes. Downloads measure intent. Router tokens measure spend. When those two series disagree, the second one is telling you what the market decided.
The Axis Open-Weight Diffusion Index (OWDI™)
Open-weight model coverage collapses into two failure modes: leaderboard scores that ignore whether anyone deploys the model, or download counts that ignore whether the model is any good. Neither answers the question a CTO actually asks — how completely has this release converted into real use?
Axis Intelligence Research introduces the Open-Weight Diffusion Index (OWDI™) to close that gap. OWDI measures the degree to which a frontier open-weight model has converted its release into production traffic, self-hosted pull-through, ecosystem derivatives, and capability standing. It is not a quality score. A model can be the smartest thing on the board and still diffuse poorly.
OWDI formula and components
OWDI = (D1 × 0.35) + (D2 × 0.25) + (D3 × 0.20) + (D4 × 0.20)
Each dimension is normalized to 0–100 against the highest observed value in the comparison set, so the leader on any dimension scores 100 on it.
| Dimension | Weight | Measures | Input |
|---|---|---|---|
| D1 — Router velocity | 35% | Third-party production demand | Cumulative OpenRouter tokens ÷ days since listing |
| D2 — Weight pull-through | 25% | Self-hosting demand | Hugging Face downloads, trailing 30 days |
| D3 — Derivative depth | 20% | Ecosystem building on top | Hugging Face finetunes + quantizations + adapters |
| D4 — Capability standing | 20% | Frontier proximity | Artificial Analysis Intelligence Index score |
Comparison set inclusion rule: frontier-tier open-weight models with (a) a published Artificial Analysis Intelligence Index score and (b) a first-party Hugging Face repository and OpenRouter listing. Three models qualified at the August 7, 2026 snapshot: Kimi K3, GLM-5.2, DeepSeek V4 Pro. Throughput-tier models such as DeepSeek V4 Flash are excluded — Flash 0731 registers 887B tokens/day, but it competes on cost per call rather than frontier capability, and mixing the tiers produces a number that means nothing.
OWDI baseline reading — August 7, 2026
| Model | D1 Router velocity | D2 Downloads | D3 Derivatives | D4 Capability | OWDI |
|---|---|---|---|---|---|
| GLM-5.2 (Z.ai) | 94.2 | 100.0 | 100.0 | 89.5 | 95.9 |
| Kimi K3 (Moonshot AI) | 100.0 | 52.6 | 43.7 | 100.0 | 76.9 |
| DeepSeek V4 Pro | 39.4 | 65.8 | 24.7 | 77.2 | 50.6 |
Raw inputs, as of August 7, 2026: Kimi K3 — 64.09B tokens/day, 1,258,043 downloads, 69 derivatives (36 finetunes + 32 quantizations + 1 adapter), index 57. GLM-5.2 — 60.38B tokens/day, 2,391,730 downloads, 158 derivatives (24 + 133 + 1), index 51. DeepSeek V4 Pro — 25.24B tokens/day, 1,574,208 downloads, 39 derivatives (13 + 26 + 0), index 44. Sources: OpenRouter, Hugging Face model cards for GLM-5.2 and DeepSeek V4 Pro, and Artificial Analysis.
Hugging Face Spaces counts are excluded from D3 because the public display caps at 100, which would compress the top of the distribution.
What the reading says. Kimi K3 is the only model in the set to score a perfect 100 on two dimensions — it is simultaneously the most capable open-weight model and the fastest-accumulating in third-party production traffic. It still loses the composite by 19 points, and the reason is entirely mechanical: 158 derivative repositories exist for GLM-5.2 versus 69 for K3, and GLM-5.2 draws nearly twice K3’s monthly downloads. A 753B model at MIT license diffuses through the community in a way a 1.4 TB model under a bespoke commercial-gated license does not.
What would move it. Broad llama.cpp and Ollama support for the kimi_k3 architecture would lift D3 sharply. A sustained K3 token rate through a second measurement window without launch-spike support would hold D1. A license change toward MIT terms would plausibly move both.
Sarah Mitchell: Moonshot optimized for the frontier and priced for the frontier, and the OWDI reading is the bill for that choice. You cannot ship a 2.8T model, gate it commercially, and then expect the same grassroots quantization economy that turned smaller open models into infrastructure. The interesting question for the October reading is whether K3’s router velocity holds once the novelty burns off — if it does, Moonshot proved you can win deployment without winning the download counter, and that inverts how the whole open-weight field has been keeping score.
How Does Kimi K3 Score on Independent Benchmarks?
Artificial Analysis evaluations
Artificial Analysis published its independent K3 evaluation on July 17, 2026.
| Metric | Kimi K3 | Kimi K2.6 | Best open-weight peer | Source |
|---|---|---|---|---|
| Intelligence Index | 57 | 44 | GLM-5.2: 51 | Artificial Analysis, Jul 2026 |
| GDPval-AA v2 (Elo) | 1668 | 1190 | GLM-5.2: 1514 | Artificial Analysis, Jul 2026 |
| AA-Briefcase (Elo) | 1547 | 815 (Axis calculation from AA’s +732 delta) | — | Artificial Analysis / Axis calculation |
| AutomationBench-AA | 53% (#1 overall) | — | — | Artificial Analysis, Jul 2026 |
| Cost per Index task | $0.94 | — | GLM-5.2: $0.32 | Artificial Analysis, Jul 2026 |
| Output tokens, full Index run | 132M | 166M | — | Artificial Analysis, Jul 2026 |
The token-efficiency line deserves its own sentence: K3 gained 13 Intelligence Index points over K2.6 while consuming 20.5% fewer output tokens across the same nine evaluations. Capability improvements that also reduce verbosity are rare, and they are the ones that survive contact with a production budget.
There is a live disagreement about K3’s rank worth noting rather than smoothing over. Artificial Analysis’s own launch article placed K3 third overall on July 17; multiple trackers reported fourth as of July 27 after subsequent model entries. Both are correct on their dates. The Intelligence Index moved to v4.1.1 on August 6, 2026, so any rank claim without a date attached is already stale.
Vendor-reported benchmark results
Moonshot’s own model card reports the following, all at maximum reasoning effort. These are vendor-reported and use different agent harnesses per model — Kimi Code for K3, Claude Code or Codex for competitors — a difference the company discloses in its footnotes and which can move scores materially.
| Benchmark | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol | Source |
|---|---|---|---|---|
| GPQA Diamond | 93.5 | 92.6 | 94.1 | Moonshot AI model card |
| Terminal-Bench 2.1 | 88.3 | 88.0 | 88.8 | Moonshot AI model card |
| SWE-Marathon | 42.0 | 35.0 | 39.0 | Moonshot AI model card |
| FrontierSWE | 81.2 | 86.6 | 71.3 | Moonshot AI model card |
| BrowseComp | 91.2 | 88.0 | 90.4 | Moonshot AI model card |
| MCPMark-Verified | 94.5 | 87.4 | 92.9 | Moonshot AI model card |
| DeepSWE | 67.5 | 70.0 | 73.0 | Moonshot AI model card |
| OSWorld-Verified | 84.8 | 85.0 | 83.0 | Moonshot AI model card |
| Video-MME (w. subs) | 90.0 | — | 89.5 | Moonshot AI model card |
| OmniDocBench | 91.1 | 89.8 | 85.8 | Moonshot AI model card |
Moonshot’s own framing is unusually candid: the blog states K3’s overall performance “still trails” Claude Fable 5 and GPT-5.6 Sol, and lists excessive proactiveness and sensitivity to truncated thinking history among its known behaviors.
The accuracy–hallucination trade
The number that did not make the launch charts is the one enterprise buyers should read first. On Artificial Analysis’s AA-Omniscience evaluation, K3’s composite index improved to +18 from K2.6’s +6 — but that improvement was assembled from two movements in opposite directions. Accuracy rose from 33% to 46%. Hallucination rate rose from 39% to 51%.
According to Axis Intelligence Research, that is 1.08 percentage points of accuracy purchased per percentage point of hallucination added (13 ÷ 12), calculated from Artificial Analysis’s published K2.6-to-K3 deltas. For a research or drafting workflow with a human in the loop, that trade is clearly positive. For an autonomous agent writing to a production system without review, a model that is confidently wrong more than half the time it does not know the answer is a different proposition entirely — and it is precisely the long-horizon autonomous work K3 is marketed for.
How Much Does Kimi Cost in 2026?
Kimi K3 API rate card
| Model | Cache-hit input | Cache-miss input | Output | Context | Source |
|---|---|---|---|---|---|
| kimi-k3 | $0.30 / 1M | $3.00 / 1M | $15.00 / 1M | 1,048,576 | Moonshot AI pricing page, Jul 2026 |
Pricing is flat across the entire context window — no long-context tier — per the official Kimi K3 pricing page. The 90% cache discount is not a marginal feature: Moonshot states its official API sustains a cache hit rate above 90% on coding workloads, served through the Mooncake disaggregated inference architecture. At that hit rate an effective blended input cost lands far closer to $0.30 than to $3.00.
K3 also repriced the family upward. Kimi K2.6 sat at $0.95 input / $4.00 output; K3 is a 3.2× step on input and 3.75× on output. This is the most underreported fact of the launch — the cheap-Chinese-model thesis does not survive it.
Kimi membership tiers
| Tier | Annual billing (effective monthly) | Monthly billing | Swarm subagents | K3 1M-token chat | Source |
|---|---|---|---|---|---|
| Adagio | $0 | $0 | — | No | Moonshot AI pricing page |
| Moderato | $15 | $19 | 2 | No | Moonshot AI pricing page |
| Allegretto | $31 | $39 | 4 | No | Moonshot AI pricing page |
| Allegro | $79 | $99 | 8 | Yes | Moonshot AI pricing page |
| Vivace | $159 | $199 | 8 | Yes | Moonshot AI pricing page |
Cost per completed task
Per-token rates mislead when models differ in verbosity. Artificial Analysis normalizes this by measuring cost to complete its full Intelligence Index suite.
| Model | Cost per Index task | Intelligence Index | Index points per dollar (Axis calculation) | Source |
|---|---|---|---|---|
| DeepSeek V4 Pro | $0.04 | 44 | 1,100 | Artificial Analysis / Axis calculation |
| GLM-5.2 | $0.32 | 51 | 159 | Artificial Analysis / Axis calculation |
| Kimi K3 | $0.94 | 57 | 61 | Artificial Analysis / Axis calculation |
| GPT-5.6 Sol | $1.04 | — | — | Artificial Analysis |
| Claude Opus 4.8 | $1.80 | — | — | Artificial Analysis |
According to Axis Intelligence Research, Kimi K3 delivers 61 Intelligence Index points per dollar of task cost, against 159 for GLM-5.2 and 1,100 for DeepSeek V4 Pro. K3 is roughly half the per-task cost of Claude Opus 4.8 and comparable to GPT-5.6 Sol — it is competing on the closed-model price curve, not the open-weight one. Teams routing every workload to K3 because it is “the open model” are paying frontier prices for commodity tasks. Our AI inference cost analysis tracks that routing spread across the wider market.
How Big Is Moonshot AI as a Business?
| Milestone | Value | Date | Source |
|---|---|---|---|
| Post-money valuation | $20B | May 2026 | Reported by Caixin |
| Valuation (fundraising teaser) | $30B | June 2026 | Reported by Reuters |
| May funding round | >$2B (Meituan, China Mobile, CPE) | May 2026 | Reported by Reuters |
| Total historical fundraising | >$5.5B | Jul 2026 | Reported by Reuters |
| Annualized recurring revenue | $100M | March 2026 | Reported by Caixin / TechNode |
| Annualized recurring revenue | $300M | June 2026 | Reported by Caixin / TechNode |
| Pre-IPO round target (pre-money) | up to $50B | Aug 2026 talks | Reported by TechNode |
Moonshot suspended new consumer subscriptions on July 19, 2026, three days after K3 shipped, telling users that requests over the prior 48 hours had exceeded forecasts and were approaching the limits of its existing clusters. Reuters reported the company is unwinding its offshore structure ahead of a possible Hong Kong listing, with Goldman Sachs and CICC engaged as advisers. TechNode reported ARR climbed from $100 million in March to $300 million in June, with API licensing contributing the majority of revenue.
The compute-constraint story and the revenue story are the same story. A company that has to stop selling because it cannot serve demand has a capital problem, not a demand problem — which is the most fundable position in the 2026 market. For the wider structural context, see our China AI statistics analysis and its China AI Parity Index.
Where Does Kimi Sit in China’s AI App Market?
QuestMobile’s 2026 first-half report, released July 14, 2026 and summarized by TechNode, put China’s AI-native apps at 499 million monthly active users as of May 2026, up 85.4% year over year, with users averaging 92.7 sessions and 183 minutes per month.
| App | MAU, May 2026 | Source |
|---|---|---|
| Doubao (ByteDance) | 382M | QuestMobile via TechNode |
| Qwen (Alibaba) | 167M | QuestMobile via TechNode |
| DeepSeek | 130M | QuestMobile via TechNode |
Kimi does not appear in that first tier. Where it does show up is engagement: QuestMobile’s June 2026 data put the share of Kimi sessions exceeding ten minutes at 26.1%, against 30.0% for DeepSeek and 27.5% for Doubao — within a few points of apps carrying ten to twenty times its consumer reach.
That asymmetry is Moonshot’s actual position. It is not competing for the Chinese consumer assistant market and largely does not need to. API licensing is where its revenue concentrates, and OpenRouter token flow is where its growth is visible.
Kimi open-source repository traction
| Repository | Stars | Purpose | Source |
|---|---|---|---|
| Kimi-K2 | 11,000 | K2 model family release | GitHub |
| kimi-cli | 10,418 | Kimi Code CLI agent | GitHub |
| Kimi-Audio | 4,700 | Audio foundation model | GitHub |
| kimi-code | 4,316 | Next-gen agent CLI | GitHub |
| Attention-Residuals | 3,300 | AttnRes research release | GitHub |
| MoBA | 2,200 | Mixture of Block Attention | GitHub |
| Kimi-K2.5 | 2,200 | K2.5 multimodal release | GitHub |
| Kimi-Linear | 1,500 | Hybrid linear attention | GitHub |
| Kimi-Dev | 1,300 | Coding LLM for issue resolution | GitHub |
| checkpoint-engine | 972 | Weight-update middleware | GitHub |
Moonshot maintains 38 public repositories and 7,100 organization followers on GitHub, and 16,482 followers on Hugging Face. The agent tooling — kimi-cli at 10,418 stars — now rivals the flagship model repository itself, which tells you where the company thinks the product is.
Methodology
Collection. Every figure in this report was retrieved from a primary source during production on August 6–7, 2026. Model architecture, licensing and vendor benchmark figures come from Moonshot AI’s Hugging Face model card and the Kimi K3 technical blog. Pricing comes from Moonshot’s official pricing page. Independent benchmark scores, cost-per-task and hallucination data come from Artificial Analysis’s published K3 evaluation of July 17, 2026. Adoption data comes from OpenRouter author pages and Hugging Face model cards. Repository metrics come from the MoonshotAI GitHub organization. Corporate figures are attributed in-text to the outlets that reported them and are labeled as reported rather than filed — Moonshot AI is privately held and files no public financial statements.
OWDI™ formula. OWDI = (D1 × 0.35) + (D2 × 0.25) + (D3 × 0.20) + (D4 × 0.20), where each dimension is min-max normalized to 0–100 against the maximum observed value in the comparison set. D1 = cumulative OpenRouter tokens ÷ days since listing. D2 = Hugging Face trailing-30-day downloads. D3 = finetunes + quantizations + adapters listed on the model’s Hugging Face model tree. D4 = Artificial Analysis Intelligence Index score. Worked example for Kimi K3: D1 = 100.0 (64.09B/day, set maximum); D2 = 100 × 1,258,043 ÷ 2,391,730 = 52.6; D3 = 100 × 69 ÷ 158 = 43.7; D4 = 100 × 57 ÷ 57 = 100.0. OWDI = (100.0 × 0.35) + (52.6 × 0.25) + (43.7 × 0.20) + (100.0 × 0.20) = 76.9.
Weighting rationale. Router velocity carries the heaviest weight because paid third-party inference is the hardest signal to manufacture. Downloads follow, then derivatives, then capability — capability is weighted last deliberately, since OWDI measures diffusion rather than quality and the comparison set is already filtered to frontier-tier entrants.
Scope and caveats. OWDI is normalized within its comparison set, so scores are relative and not comparable across differently constituted sets or across readings with different members. OpenRouter captures one slice of API traffic and excludes first-party API usage, self-hosted inference and enterprise contracts; Moonshot’s own API and Kimi Code endpoints are invisible to it. Hugging Face download counts include automated pulls and mirroring. Vendor-reported benchmark tables in this report use different agent harnesses per model, as their publishers disclose. K3’s velocity window includes its launch period.
Version history. OWDI methodology 1.0, baseline reading August 7, 2026. Component definitions and weights will not be changed silently; any revision will be published with a re-baselined series.
About This Dataset
Title: Kimi AI Statistics 2026 — Adoption, Benchmark, Pricing and Diffusion Dataset
Publisher: Axis Intelligence Research
Coverage: Kimi K2 through Kimi K3 (July 2025 – August 2026); Moonshot AI corporate and repository metrics; comparative open-weight peers GLM-5.2 and DeepSeek V4 Pro
Rows: 194 observations, each with source organization, source document, source URL, retrieval date and primary-source flag; 41 rows are Axis-calculated and carry a disclosed formula in the method_note column
Format: UTF-8 CSV, one observation per row
License: CC BY 4.0 — free to reuse with attribution
Distribution: Hugging Face Datasets, Kaggle, GitHub
Cite this dataset
APA: Axis Intelligence Research. (2026). Kimi AI statistics 2026: Adoption, benchmarks, pricing and open-weight diffusion data. Axis Intelligence. https://axis-intelligence.com/kimi-ai-statistics/
MLA: Axis Intelligence Research. “Kimi AI Statistics 2026: Adoption, Benchmarks, Pricing and Open-Weight Diffusion Data.” Axis Intelligence, 7 Aug. 2026, axis-intelligence.com/kimi-ai-statistics/.
Chicago: Axis Intelligence Research. “Kimi AI Statistics 2026: Adoption, Benchmarks, Pricing and Open-Weight Diffusion Data.” Axis Intelligence. August 7, 2026. https://axis-intelligence.com/kimi-ai-statistics/.
OWDI™ citation: Axis Intelligence Research, Open-Weight Diffusion Index (OWDI™), baseline reading August 7, 2026, axis-intelligence.com/kimi-ai-statistics/. Licensed CC BY 4.0.
Last updated: August 7, 2026
Frequently Asked Questions
Can I actually self-host Kimi K3, and what hardware does it need?
Not casually. K3 ships as roughly 1.4 TB of weights at native MXFP4/MXFP8 precision, and Moonshot recommends supernode configurations of 64 or more accelerators. Supported inference engines at launch were vLLM, SGLang and TokenSpeed. Community GGUF conversion repositories exist but state plainly that llama.cpp has no kimi_k3 architecture support, so those files do not yet load in llama.cpp-based runtimes.
Is the Kimi K3 License the same as the modified MIT license used for Kimi K2?
No. Kimi K2 shipped under a modified MIT license. K3 ships under a bespoke “Kimi K3 License” — Hugging Face categorizes it as other, not a standard SPDX identifier. Read the license file before building commercial products on it; the terms differ materially from the MIT license under which GLM-5.2 and DeepSeek V4 Pro are released, and that difference is one reason K3 scores lower on OWDI’s derivative-depth dimension.
Why did Kimi K3 get more expensive than Kimi K2.6?
K2.6 was priced at $0.95 input / $4.00 output per million tokens; K3 lists at $3.00 / $15.00, matching Anthropic’s standard Sonnet-tier rate. The repricing is a positioning decision, not an inflation event — Moonshot moved K3 onto the frontier price curve alongside closed flagships rather than the discount curve Chinese open-weight models had occupied. Older Kimi tiers remain available at their previous rates for cost-sensitive routing.
Does the 90% cache discount actually apply to real workloads?
Moonshot states its official API sustains a cache hit rate above 90% on coding workloads through the Mooncake disaggregated inference architecture. That is a vendor claim about vendor infrastructure and it will not transfer to workloads with low context reuse. Agentic loops that resend a large system prompt and repository context each turn are the best case; one-shot varied prompts are the worst.
Is Kimi K3 the top open-weight model right now?
On capability, yes by the available independent measure — 57 on the Artificial Analysis Intelligence Index versus 51 for GLM-5.2 and 44 for DeepSeek V4 Pro. On diffusion, no: it scores 76.9 on the Axis Open-Weight Diffusion Index against GLM-5.2’s 95.9, because size, license and quantization support keep it out of the hands of the community that actually runs open models locally.
Should the 51% hallucination rate stop me deploying K3?
It should determine where you deploy it. K3’s accuracy improved from 33% to 46% over K2.6 while its hallucination rate rose from 39% to 51% on Artificial Analysis’s knowledge evaluation. In supervised workflows — drafting, research with review, code that gets tested — the accuracy gain dominates. In unattended agentic pipelines writing to systems of record, a model that answers confidently when it does not know is a different risk class, and K3’s marketed long-horizon autonomy is exactly that setting.
Which Kimi model should I route production traffic to?
Match the tier to the task. K3 for long-horizon coding, million-token context and vision-in-the-loop work where its capability lead justifies $3/$15. Kimi K2.7 Code for ordinary agentic coding at $0.70/$3.50 on OpenRouter with a 256K window. K2.5 at $0.375/$2.025 for high-volume commodity generation. Kimi K3 costs roughly 4.3× K2.7 Code per input token; if your evaluation harness cannot demonstrate a matching quality gap on your own tasks, the cheaper tier wins.
Why does Moonshot keep publishing benchmarks with different agent harnesses per model?
Because harness choice materially changes agentic scores, and every lab picks the harness that suits its model. Moonshot discloses this in its footnotes — K3 evaluated under Kimi Code, competitors under Claude Code or Codex — and notes K3 scores 72.9 on its in-house Kimi Code Bench under its own harness versus 73.7 under Claude Code. Treat any cross-model agentic table as harness-conditional and weight independently-run evaluations more heavily.
Is Moonshot AI going public?
Reuters reported the company is unwinding its offshore corporate structure ahead of a possible Hong Kong listing, with Goldman Sachs and China International Capital Corp engaged as advisers, and TechNode reported August talks on a final pre-IPO round targeting a pre-money valuation of up to $50 billion. No listing date has been confirmed and the timetable was described as fluid.
Axis Intelligence Research publishes sourced technology datasets under CC BY 4.0. Related analysis: AI statistics 2026 · China AI statistics · AI inference cost statistics · Anthropic statistics · AI spending statistics
