Luxsynthesis Blog
All articles
Models

Kimi K3 Released: 2.8 Trillion Parameters, the World's First 3-Trillion-Class Open-Source Model

Moonshot AI has launched Kimi K3: a 2.8-trillion-parameter MoE built on KDA hybrid linear attention and Attention Residuals, with native visual understanding and a 1M-token context window — the world's first open-source model in the 3-trillion class. This post rounds up the official benchmarks, long-horizon coding cases, architecture notes, access channels, and pricing.

Author
Luxsynthesis
Published
Reading time
12 min read
Kimi K3 Released: 2.8 Trillion Parameters, the World's First 3-Trillion-Class Open-Source Model

Moonshot AI has officially launched Kimi K3, positioned as its most capable model to date.

Kimi K3 is a 2.8-trillion-parameter model built on the Kimi Delta Attention (KDA) hybrid linear attention mechanism and Attention Residuals, with native visual understanding and a 1M-token context window. It is the world's first open-source model in the 3-trillion-parameter class, targeting frontier-intelligence scenarios such as long-horizon coding, knowledge work, and reasoning.

According to Moonshot AI, Kimi K3's overall performance still trails the strongest closed-source models — Claude Fable 5 and GPT-5.6 Sol — but it demonstrates frontier-level capability across the entire evaluation suite and consistently outperforms every other model.

Starting today (2026-07-17), Kimi K3 is available via kimi.com, the latest Kimi mobile app, the latest Kimi Work desktop client, Kimi Code, and the Kimi API. The default thinking effort is currently max, with low and high modes coming in a future update. Full model weights will be released by July 27, 2026, and further details on the architecture, training, and evaluation will be published alongside the technical report.

Kimi K3 core specs

Moonshot AI's most capable model to date: 2.8T parameters, 1M-token context, available across all channels since 2026-07-17.

Total parameters2.8TMoE, weights by 2026-07-27World's first 3T-class open model
Context window1Mwith native visual understanding
Active experts16 / 896Stable LatentMoE, ~1.8% sparsity
Scaling efficiency vs K2≈2.5×from architecture + data recipe changes
Architecture:KDA hybrid linear attentionAttention ResidualsThinking effort: max (low / high coming)

1. Benchmark Results

The official evaluations cover three categories — Coding, General Agents, and Visual Agents — with all models run at maximum thinking effort (max or xhigh). A note at the bottom of the official charts states: all Fable 5 scores include potential fallbacks, and all GPT-5.6 Sol scores include potential cyberguards.

Coding

Kimi K3 vs. Fable 5, GPT-5.6 Sol, Opus 4.8, GPT-5.5, and GLM-5.2 across six coding benchmarks

Coding benchmarkKimi K3Claude Fable 5GPT-5.6 SolClaude Opus 4.8GPT-5.5GLM-5.2
DeepSWE67.570.073.059.067.046.2
Terminal Bench 2.188.384.688.884.683.482.7
FrontierSWE81.286.671.366.764.967.3
Program Bench77.876.877.671.970.863.7
Kimi Code Bench 2.0 (internal)72.976.964.871.769.064.2
SWE Marathon42.035.039.040.014.013.0

General Agents and Visual Agents

Kimi K3 on general-agent benchmarks (GDPval, JobBench, BrowseComp) and visual-agent benchmarks (CharXiv, Zerobench)

General-agent benchmarkKimi K3Claude Fable 5GPT-5.6 SolClaude Opus 4.8GPT-5.5GLM-5.2
GDPval-AA v2 Elo1668.01760.01748.01600.01494.01514.0
JobBench52.957.446.548.438.343.4
AA-Briefcase Elo1548.01583.01495.01354.01158.01260.0
SpreadsheetBench 234.834.732.431.629.128.1
Automation Bench30.829.129.727.222.712.9
BrowseComp91.288.090.484.384.4
Visual-agent benchmarkKimi K3Claude Fable 5GPT-5.6 SolClaude Opus 4.8GPT-5.5
CharXiv (RQ) w/ tools91.393.589.189.989.0
Zerobench w/ tools (Pass@5)41.046.035.034.041.0

Scores and Cost

Moonshot AI also published score-vs.-per-task-cost comparisons for Kimi Code Bench V2 and BrowseComp: Kimi K3 (max) reaches scores close to top closed-source models at a significantly lower per-task cost. Approximate values read from the charts:

  • Kimi Code Bench V2: K3 (max) costs about $4 per task (72.9%), roughly 40% of what Claude Fable 5 costs (max, ~$10.6, 76.9%); Claude Opus 4.8 (max) is about $6.8 (71.7%). K3's low and high modes cost roughly $1.3 and $2.5 respectively (66.5% / 70.8%).
  • BrowseComp: K3 (max) costs about $4.3 (91.2%); in the same chart, GPT-5.6 Sol (max) is about $6.3, Claude Sonnet 5 (max) about $20.7, Claude Opus 4.8 (max) about $25, and Claude Mythos 5 (max, 10M tokens) about $26.
Score vs. cost: near-frontier scores at ~40% of the cost

Official score-vs.-per-task-cost comparisons. Values are approximate readings from the official charts.

Kimi Code Bench V2: score vs. per-task cost
$0$3$6$9$12657075K3 lowK3 highK3 maxOpus 4.8 maxFable 5 max

Y axis: score (%); X axis: per-task cost. K3 max costs roughly 40% of Fable 5 max.

BrowseComp: per-task cost (max effort)
Kimi K3 max$4.3
GPT-5.6 Sol max$6.3
Claude Sonnet 5 max$20.7
Claude Opus 4.8 max$25.0
Claude Mythos 5 max$26.0

K3 max costs ~$4.3 per task; other models run up to $25+.

Score vs. per-task cost on Kimi Code Bench V2: Kimi K3 max reaches 72.9% at ~$4; Fable 5 max costs ~$10.6

Score vs. per-task cost on BrowseComp: Kimi K3 max reaches 91.2% at ~$4.3; other models cost up to $25+

Full Benchmark Table

The official full benchmark table also includes PostTrain Bench, MLS Bench, Toolathlon, MCP Atlas, GPQA, HLE, MMMU, and more:

Full benchmark comparison

37 benchmarks × 6 models. Chart view normalizes within each row for relative comparison; data view lets you inspect and copy the raw table.

Kimi K3Fable 5GPT-5.6 SolOpus 4.8GPT-5.5GLM-5.2
DeepSWE
67.5#3
Program Bench
77.8#1
Terminal Bench 2.1
88.3#2
FrontierSWE
81.2#2
SWE Marathon
42.0#1
PostTrain Bench
36.6#2
MLS Bench
48.3#2
Kimi Code Bench 2.0 (internal)
72.9#2

Scales differ across benchmarks (Elo vs. percentages), so dot positions are normalized within each row for relative comparison only. The halo marks the row's best score; the right column shows K3's score and rank.

* Marks carried over from the official table (the official footnotes do not explain their meaning). "—" means the official table does not report a score.

Evaluation methodology (key points from the official footnotes):

  • All Kimi K3 results were obtained at max thinking effort with temperature = 1.0 and top-p = 1.0. Depending on the benchmark, models were evaluated with one of three agentic harnesses: KimiCode, Claude Code, or Codex.
  • Claude Fable 5 was evaluated by a third party; under the Claude Code harness, requests refused by Fable 5 due to its usage policy automatically fall back to Claude Opus 4.8.
  • DeepSWE: K3's 67.5 was achieved with the KimiCode harness; on the official leaderboard, its score with the mini-SWE-agent harness is 67.3.
  • BrowseComp: uses the context-compaction strategy from the Claude model card (triggered at 300K tokens); K3 scores 90.4 with a 1M context and no context management.
  • MCP Atlas: a public subset of 500 tasks with a 100-turn cap, judged by Gemini 3.1 Pro. AutomationBench: a public subset of 600 tasks.
  • GDPval-AA v2 and AA-Briefcase results are sourced from artificialanalysis.ai; FrontierSWE dominance scores were recomputed using the official script, with data as of 2026-07-16.
  • All multimodal benchmarks are averaged over 3 runs, except ZeroBench (official setting, 5 runs). PerceptionBench is an official internal benchmark focused on atomic visual perception.

2. A 3-Trillion-Class Open-Source Model

Kimi K3 is the first open-source model to reach 2.8 trillion parameters. According to Moonshot AI, Kimi models have held the parameter-scale record among open-source models for 9 of the past 12 months.

Parameter-scale timeline of flagship open-source models (July 2025 – July 2026): Kimi K3 leads at 2.8T, followed by DeepSeek 1.6T, Xiaomi 1.02T, Thinking Machines 975B, Z.AI 744B, MiniMax 428B, and Alibaba 397B

Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) — two architectural updates designed to let information flow more smoothly through longer sequences and deeper models. MoE sparsity is pushed further: with the Stable LatentMoE framework, the model efficiently activates 16 out of 896 experts. Combined with improvements to the training methodology and data recipe, these structural changes give Kimi K3 roughly 2.5x the overall scaling efficiency of K2.

Stable LatentMoE: 896 experts, only 16 active per token

Highlighted cells are the experts active for the current token; a new batch lights up every ~2s to simulate routing of incoming tokens.

Total experts
896
Active / token
16
Activation rate
1.8%

At this sparsity, routing and optimization become the key challenges — Quantile Balancing assigns experts directly by routing-score quantiles.

3. Long-Horizon Coding

Moonshot AI says Kimi K3 has strong long-horizon coding capabilities: with minimal human supervision, it can sustain long engineering tasks, understand and work through large codebases, and orchestrate terminal tools. It also excels at tasks that combine software engineering with visual reasoning, using screenshots and visual feedback in scenarios such as game development, front-end work, and CAD.

The Kernel Optimization Arena

Moonshot AI built a kernel optimization arena: each model runs in an isolated GPU sandbox with identical tasks, test harnesses, and production workloads, and has up to 24 hours to analyze existing implementations, rewrite kernels, and repeatedly benchmark and verify the results. The arena spans two hardware platforms, three kernel types, and four tasks: Attention Residuals and KDA linear attention on NVIDIA H200, an MLA kernel with a 512 head dimension implemented from scratch, and a KDA task on a domestic Chinese GPU.

At maximum thinking effort, Kimi K3 performs close to Fable 5 (with fallbacks) and clearly leads Opus 4.8, GPT-5.6 Sol, and GPT 5.5. On the AttnRes task, for example, speedups over the FLA Triton baseline are: Kimi K3 +59.7%, Claude Fable 5 +57.1%, GPT 5.5 +30.8%, and GPT 5.6 Sol +17.3%.

AttnRes task in the kernel optimization arena: speedup curves over runtime, with Kimi K3 leading at +59.7%

Moonshot AI shared further details on the AttnRes task:

  • The task provides the FLA Triton implementation of the AttnRes kernel along with its production shape (96 layers, model dimension 8192, 8192 tokens), and asks for the training-side operator to be made as fast as possible without changing numerical results.
  • Over 15 consecutive hours of iteration, K3 designed a new two-phase kernel algorithm that fuses kernels while preserving numerical equivalence, cutting forward + backward time from 283.6 ms to 114.4 ms. K3 and Fable 5 (with potential fallbacks) ended up with similar final performance, and K3 optimized faster per iteration.
  • Moonshot AI also noted that in the later stages of Kimi K3's development, an early version of K3 was already handling most of the team's kernel optimization work.
  • Evaluation notes: Claude Fable 5 was evaluated by a third party and its results may include fallback behavior; some trajectories from most models contain minor precision shortcuts within numerical tolerance.
Kernel optimization arena: AttnRes speedup

Speedup over the FLA Triton baseline at max thinking effort. Each model runs in an isolated GPU sandbox with up to 24 hours to rewrite and verify kernels.

Kimi K3
+59.7%
Claude Fable 5*
+57.1%
GPT 5.5
+30.8%
GPT 5.6 Sol
+17.3%
K3's two-phase kernel: forward + backward time
Before
283.6 ms
After 15h of iteration
114.4 ms

Kernels were fused while preserving numerical equivalence, cutting the time to 40.3% of the original.

* Fable 5 was evaluated by a third party; results may include fallback behavior.

GPU Compiler Development: MiniTriton

Moonshot AI further evaluated whether Kimi K3 could build a complete GPU programming system. From scratch, Kimi K3 developed MiniTriton: a lightweight, Triton-like compiler that defines its own tile-level intermediate representation on top of MLIR (without relying on Triton IR) and implements the full pipeline from optimization passes to PTX code generation.

In roofline benchmarks, MiniTriton matches or beats Triton and torch.compile on the operators it currently supports — outperforming Triton on some workloads. It currently supports FP32 and FP64 computation, with TF32 and BF16 still in development. According to Moonshot AI, without taking benchmark-specific shortcuts or obvious hard-coded optimizations, MiniTriton already generates competitive GPU code, and its from-scratch Tensor Core path rivals the heavily optimized Triton stack.

CUDA-core roofline benchmark (fp32) on NVIDIA L20: MiniTriton vs. Triton, torch.compile, and other implementations

As an end-to-end validation, Moonshot AI used TensorLite to train nanoGPT and observed stable convergence, with the loss curve closely tracking the reference implementation and deviating only slightly — evidence that K3 built a genuinely usable compiler stack: from the DSL frontend and IR passes to PTX code generation and runtime, plus the ability to handle real training workloads.

Creating Digital Works

Kimi K3 combines 3D reasoning, programming, and visual capabilities to turn concepts, images, and videos into playable interactive experiences, iterating between code and live screenshots to achieve vision in the loop.

Moonshot AI released four demo videos: a 3D simulation of the Long March 10 rocket launch and recovery, a 3D open-world game, a 3D GBA emulator, and a recreation of Gargantua, the black hole from Interstellar.

There are also nine interactive game/creation demos: the 3D open world, a GBA emulator, a cyberpunk rope-swinging game, a typewriter, a voxel colosseum, a fighting game, a wuxia (martial-arts fantasy) RPG, an FPS arena, and the Gargantua black hole.

The flagship case, 3D Open World: K3 used Three.js WebGPU and GPU compute to build a fully procedural 3D exploration game in the browser — the environment is procedurally generated, the rider and horse models were created with 3D asset-generation tools, and the world includes forests, wooden-cabin villages, snow-capped mountains, and dynamic weather (external assets: animated cowboy and horse models, terrain data).

K3-built 3D open-world game: procedurally generated forests, log-cabin villages, and snow-capped mountainsK3-built 3D GBA emulator: a handheld-style console UI with built-in arcade mini-gamesK3-built cyberpunk rope-swinging game: swinging between city towers at duskK3's recreation of the Gargantua black hole: accretion disk and gravitational lensing in an interactive demo

Chip Design

As an early proof of concept, Kimi K3 designed a chip for running a nano model built on K3's own architecture. In a single continuous 48-hour autonomous agent run, K3 independently completed the chip's construction, optimization, and verification using open-source EDA tools and the Nangate 45nm process design kit:

  • 4 mm² die area integrating 1.46 million standard cells
  • 0.277 MB of SRAM, with an INT4 MAC array featuring fused dequantization
  • Timing closure achieved at 100 MHz
  • Simulated decoding throughput sustained above 8,700 tokens per second

(The demo video for this case runs 00:22 and was generated by Kimi K3 connected to Blender MCP.)

A chip in 48 hours: key metrics

Using open-source EDA tools and the Nangate 45nm PDK, K3 independently built, optimized, and verified a chip that runs a nano model based on its own architecture.

Autonomous agent run0hourssingle uninterrupted session
Die area0mm²Nangate 45nm PDK
Standard cells0.00MINT4 MAC array, fused dequant
SRAM0.000MBon-chip memory
Timing closure0MHzclock frequency achieved
Decoding throughput0+tok/ssustained in simulation

Scientific Programming

In a case reproducing the I-Love-Q universal relations from computational astrophysics, Kimi K3 completed in about two hours work that would typically take an experienced researcher one to two weeks: reading and cross-validating more than 20 papers, implementing the full numerical pipeline, evaluating over 300 equations of state, spotting inconsistencies in published formulas, generating more than 3,000 lines of Python code, and producing an interactive HTML dashboard for exploring the results (interactive version: https://coding-science.ok.kimi.link/).

Kimi K3's interactive I-Love-Q universal-relations dashboard: three chart groups covering pressure–energy density, stellar sequences, and universal relations with deviations

4. Knowledge Work

Beyond public benchmarks, Kimi K3 (max) also shows consistent gains in Moonshot AI's internal evaluations, which are drawn from task patterns and challenges that recur in real user-agent collaboration workflows.

Internal Knowledge Work Bench: Kimi K3 leads GPT 5.5 and Claude Opus 4.8 on all three internal benchmarks — Online Exp Bench (75.5), DECK-Bench (73.5), and Finance-Bench (62.6)

Internal knowledge-work benchmarkKimi K3GPT 5.5Claude Opus 4.8
Online Exp Bench75.570.665.9
DECK-Bench73.568.266.9
Finance-Bench62.658.460.7

Research and Visualization

Moonshot AI showcased three K3 cases in Kimi Work across financial consulting and scientific research scenarios:

Inference-chip industry research: covering 42 years of ASIC industry history, generated by K3 over more than 120 rounds of recursive self-improvement. The process involved more than 2,800 web search-and-scrape operations, 1,100+ terminal data pulls, and the processing of 87 quarterly reports and 99 raw PDFs — over 11,000 pages of material in total. K3 turned the evidence into custom charts, animated diagrams, and interactive visual narratives. Full interactive report: https://asic42cn.ok.kimi.link/ (Chinese), https://asic42.ok.kimi.link/ (English)

Kimi K3's interactive ASIC industry research report: lifelines of three generations of custom-computing players (1975–2026) and a consolidation flow map

Controlled nuclear fusion industry research: a consulting-style industry study featuring timelines, tree diagrams, waterfall charts, and Gantt charts, plus a publication-quality slide deck.

Kimi K3's controlled-fusion industry research report: analysis pages on the two major bottlenecks, HTS tape and tritium fuel

GWTC-5 gravitational-wave analysis: an analysis of 391 gravitational-wave events using more than 20 concurrent subagents, producing 7 scientific visualizations and 2 tables while synthesizing content from 10+ papers.

Kimi K3's GWTC-5 gravitational-wave analysis report: a page explaining the principles of gravitational-wave detection

Moonshot AI also noted that Kimi K3 is particularly good at producing infographic-style presentations, and showed two fully editable examples: a heatmap and an annual report.

Widgets and Dashboard

Kimi Work is introducing two new features:

  • Widgets: generate interactive components directly in conversation, connected to local data or external plugins for continuous updates
  • Dashboard: gather the widgets you care about most into a persistent personalized view, organized around a topic, project, or goal

Video Editing

Moonshot AI says Kimi K3 excels at motion design, animation, and video editing because its natively multimodal architecture understands text, images, and video within a single model. Two cases:

  • K3 produced a 3Blue1Brown-style motion-graphics explainer video about its own architecture, turning technical concepts into animated diagrams and transitions, running four and a half minutes
  • K3 edited its own brand video from 56 raw clips, handling footage selection, match cuts, frame-accurate beat syncing, audio processing, and multiple rounds of revision. According to Moonshot AI, such information-dense short videos typically take an experienced editor 1–2 working days, and a novice 3–5 days

5. Architecture and Infrastructure

Kimi K3's architectural backbone consists of KDA and AttnRes: KDA provides an efficient foundation for scaling attention, while AttnRes selectively retrieves representations across depth instead of simply accumulating them uniformly across layers. Together they support scaling the model beyond a trillion parameters.

Kimi K3 architecture diagram: Stable LatentMoE and KDA modules on the left, the AttnRes operation at upper right, and the Block Attention Residuals backbone on the right

Key technical details:

  • Stable LatentMoE: activates 16 of 896 experts. At this level of sparsity, routing and optimization become the key challenges
  • Quantile Balancing: assigns experts directly based on quantiles of the routing scores, avoiding heuristic updates and sensitive balancing hyperparameters
  • Per-Head Muon: extends Muon to per-attention-head optimization, making learning in large-scale training more adaptive
  • Sigmoid Tanh Unit (SiTU) and Gated MLA: strengthen activation control and attention selectivity, respectively
  • Quantization-aware training: starting from the SFT stage, uses MXFP4 weights and MXFP8 activations to accommodate a broader range of hardware
  • Perfectly balanced expert-parallel training: prevents uneven expert loads from hurting throughput in large-scale expert parallelism; uses static shapes with no host synchronization on the critical path
  • Deployment recommendation: inference efficiency benefits from larger high-bandwidth communication domains; Moonshot AI recommends deploying on supernode configurations of 64 or more accelerators
  • vLLM contribution: KDA poses new challenges for traditional prefix caching; Moonshot AI has contributed its implementation to the vLLM community, to be released alongside the model. With KDA plus prefill cache, the model can be served at a competitive token price even at large scale and long context

Further technical details will be published with the upcoming technical report.

6. Access and Pricing

How to use:

  • Kimi Agents: download or upgrade to the latest Kimi app from your mobile app store (iOS, Android, and HarmonyOS supported), or visit kimi.com directly
  • Kimi Work: download the latest desktop client (version 3.1.0 or later), available for Windows and Apple-silicon Macs
  • Kimi Code: runs in your computer's terminal; select the K3 model with the /model command
  • Kimi API: select kimi-k3 on the open platform

API pricing (per million tokens):

ItemRMBUSD
Input (cache hit)¥2$0.30
Input (cache miss)¥20$3.00
Output¥100$15.00

According to Moonshot AI, thanks to the Mooncake disaggregated inference architecture, the official Kimi API achieves a cache rate above 90% in programming scenarios, making the effective input price just 1/4 of the standard input price. A top-up bonus promotion of up to 30% is running concurrently.

API pricing (per million tokens)

Model name kimi-k3 on the open platform; a high cache rate brings the effective input cost well below the list price.

90%+cache rate

90%+ cache rate in coding scenarios (Mooncake disaggregated inference)

Input (cache hit)¥2$0.30
Input (cache miss)¥20$3.00
Output¥100$15.00

Thanks to the high cache rate, the effective input price in coding scenarios is just 1/4 of the list input price.

For enterprises and teams:

  • Kimi Enterprise: enterprise-grade data privacy protection and member management, with complete data isolation between personal and enterprise accounts
  • Kimi Hosted Agent (coming soon): a managed agent platform for enterprises, offering an agent harness, isolated sandboxes, and a long-task runtime environment; the waitlist is now open

Open-source plan: Moonshot AI is working with inference partners and open-source maintainers to align on technical details; full model weights will be released by July 27, 2026.

7. Officially Stated Limitations

  1. Sensitive to historical thinking content: Kimi K3 uses thinking-history preservation mode throughout post-training. If an agent framework fails to pass back the full thinking history as required, or if you switch to Kimi K3 mid-session from another model, context interference may occur and generation quality may become unstable. Moonshot AI recommends using compatibility-verified agent frameworks such as Kimi Code, and avoiding switching to Kimi K3 mid-session.
  2. Overly proactive: Kimi K3's training focused on long-horizon, high-difficulty tasks; when it encounters minor issues or ambiguous user intent during execution, it may make unexpected decisions on the user's behalf. If you want the agent to stay within tighter boundaries, set clearer behavioral constraints in the system prompt or AGENTS.md.
  3. Still behind top closed-source models: while Kimi K3 is overall a highly competitive model, it still lags Claude Fable 5 and GPT-5.6 Sol in user experience.

References

All articles