Kimi K3: Open Frontier IntelligenceKimi K3:开放前沿智能 Today, we are introducing Kimi K3—our most capable model. Kimi K3 isa 2>8T-parameter model built on our Kimi Delta Attention and AttentionResiduals, with native vision capabilities and a 1-millionEtoken context window. 今天,我们正式推出Kimi K3⸺这是我们迄今为⽌能⼒最强的模型。Kimi K3是⼀个拥有2'8万亿参数的模型,基于我们的Kimi DeltaAttention和Attention Residuals技术构建,具备原⽣视觉能⼒,并⽀持100万token的上下⽂窗⼝。 It is the world's first open 3T-class model, designed for frontierintelligence across longEhorizon coding, knowledge work, and reasoning. 它是全球⾸个开源的3T级模型,专为⻓周期编程、知识⼯作和推理等前沿智能场景⽽设计。 While its overall performance still trails the most powerful proprietarymodels, Claude Fable 5 and GPT 5>6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperformingother tested models. 尽管其整体性能仍落后于最强⼤的专有模型⸺Claude Fable 5和GPT 5'6 Sol,但Kimi K3在我们的评估套件中展现了前沿⽔平的表现,持续优于其他参与测试的模型。 Kimi K3 is available today onKimi.com,Kimi Work,Kimi Code, and theKimi API. At launch, Kimi K3 will use max thinking effort by default, withlow- and highEeffort modes to be introduced in subsequent updates. Weare currently working closely with inference partners and openEsourcemaintainers to align technical details and ensure a reliable rollout acrossthe ecosystem. Kimi K3现已登陆Kimi.com、Kimi Work、Kimi Code和KimiAPI。上线初期,Kimi K3默认采⽤最⼤思考⼒度模式,低⼒度和⾼⼒度模式将在后续更新中引⼊。我们⽬前正与推理合作伙伴及开源维护者紧密协作,以统⼀技术细节,确保在整个⽣态系统中实现可靠的部署。 The full model weights will be released by July 27, 2026. Further detailson the architecture, training, and evaluations will be released alongsidethe Kimi K3 technical report. 完整模型权重将于2026年7⽉27⽇前发布。关于架构、训练和评估的更多细节,将与Kimi K3技术报告⼀同公布。 An Open 3T-Class Model ⼀款开放的3T级模型 Kimi K3 is the first open model to reach 2>8 trillion parameters. It marksthe latest step in Kimi's sustained push at the scaling frontier: for nine ofthe past twelve months, Kimi models have set the upper bound of open-model sizes. Kimi K3是⾸个达到2'8万亿参数的开源模型。这标志着Kimi在规模前沿持续突破的最新⼀步:在过去⼗⼆个⽉中,有九个⽉Kimi模型都刷新了开源模型规模的最⾼纪录。 Total parameters of each company's flagship model, Jul 2025 - Jul 2026solid = released as of Jul 16, 2026 · dotted = frontier held, no new release since · right label = company · latest size Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals(AttnRes), two architectural updates designed to improve how informationflows across sequence length and model depth. We have also scaled upMixture of Experts (MoE) sparsity, effectively activating 16 out of 896experts when paired with a Stable LatentMoE framework. Kimi K3基于Kimi Delta Attention(KDA)和AttentionResiduals(AttnRes)两⼤架构升级构建,这两项改进旨在优化信息在序列⻓度和模型深度间的流动。我们还提升了混合专家模型(MoE)的稀疏性,在稳定潜在MoE框架下,有效激活了896个专家中的16个。 Together with refined training and data recipes, these structural changes yield an approximate 2>5× improvement in overall scaling efficiencycompared to Kimi K2, allowing the model to convert compute intointelligence more effectively. 结合优化的训练策略与数据配⽅,这些结构性变⾰使整体扩展效率相⽐Kimi K2提升了约2'5倍,让模型能更⾼效地将算⼒转化为智能。 WUXL3 4 0 8 Kimi K3 architecture: the Stable LatentMoE and KDA modules (left), the AttnResoperationα(top right), and the Block Attention Residuals backbone (right). Kimi K3架构:Stable LatentMoE与KDA模块(左图)、AttnRes操作α(右上图)以及Block Attention Residuals⻣⼲⽹络(右图)。 Coding Kimi K3 has strong longEhorizon coding performance. Operating withminimal human oversight, it can sustain long engineering sessions,navigate massive repositories, and orchestrate terminal tools. Kimi K3在⻓周期编程任务中表现强劲。在极少⼈⼯⼲预下,它能持续进⾏⻓时间⼯程会话,驾驭⼤型代码仓库,并协调终端⼯具。 Kimi K3 also excels in tasks blending software engineering with visualreasoning—it leverages screenshots and visuals to optimize game dev,frontend, and CAD. Kimi K3还擅⻓融合软件⼯程与视觉推理的任务⸺它通过截图和视觉信息优化游戏开发、前端⼯程和CAD设计。 The case studies below show how Kimi K3's coding capability translatesinto openEended software creation and scientific research. 以下案例展示了Kimi K3的编程能⼒如何转化为开放式软件创作与科学研究。 Kernel Optimization内核优化 We tested the models' capability to optimize GPU kernels. Each modelworks independently in an identical sandbox, with up to 24 hours toprofile, rewrite, and benchmark four tasks spanning AttnRes, KDA, and a512-headEdimension MLA kernel across NVIDIA H200 and GPGPU from analternative vendor. 我们测试了模型优化GPU内核的能⼒。每个模型在相同的沙盒环境中独⽴运⾏,拥有最多24⼩时的时间来对四个任务进⾏性能分析、重写和基准测试,这些任务涵盖AttnRes、KDA以及⼀个512头维度的MLA内核,分别在NVIDIA H200和另⼀供应商的GPGPU上执⾏。 Kimi K3 performed competitively with Fable 5 (with fallback) andsubstantially outperformed Opus 4>8, GPT 5>6 Sol, and GPT 5>5. Kimi K3在与Fable 5(含回退机制)的竞争中表现出⾊,并⼤幅优于Opus 4'8、GPT 5'6 Sol和GPT 5'5。 AttnRes Kernel Optimization AttnRes内核优化 Given the FLA Triton implementation of AttnRes at its production shape (96 layers, modeldim 8192, 8192 tokens), the task is to make the trainingEside operation as fast as possiblewithout changing the numerics. 给定AttnRes在其⽣产规模(96层,模型维度8192,8192个token)下的FLATriton实现,任务是在不改变数值结果的前提下,尽可能加快训练侧的操作速度。 Across 15 hours of iterations nonstop, K3 designed a novel twoEphase kernel algorithm,fused kernels while