Kimi K3 是一款开源权重的原生多模态智能体模型,也是我们目前能力最强的模型。它是基于 Kimi Delta Attention(KDA)和 Attention Residuals(AttnRes)技术构建的 2.8 万亿参数模型,具备原生视觉理解能力,并支持 100 万 token 的上下文窗口。作为全球首个开源的 3T 级模型,Kimi K3 旨在为长程编码、知识工作和推理等前沿智能场景提供支持。
| 架构 | 混合专家模型(MoE) |
| 总参数 | 2.8T |
| 激活参数 | 104B |
| 层数 | 93 |
| 密集层数 | 1 |
| 注意力层构成 | 69 个 KDA + 24 个 Gated MLA |
| 注意力隐藏维度 | 7168 |
| 注意力头数 | 96 |
| 潜在 MoE 维度 | 3584 |
| MoE 隐藏维度(每专家) | 3072 |
| 专家数量 | 896 |
| 每 token 选中专家数 | 16 |
| 共享专家数量 | 2 |
| 词汇量 | 160K |
| 上下文长度 | 1048576 |
| 注意力机制 | KDA & Gated MLA |
| 激活函数 | SiTU-GLU |
| 视觉编码器 | MoonViT-V2 |
| 视觉编码器参数 | 401M |
| 量化方式 | MXFP4 权重 / MXFP8 激活 (量化感知训练) |
| 模态 | 文本、图像 |
| 基准测试 | Kimi K3 (max) | Claude Fable 5 (max, w/ fallback) | GPT-5.6 Sol (max) | Claude Opus 4.8 (max) | GPT-5.5 (xhigh) | GLM-5.2 (max) |
|---|---|---|---|---|---|---|
| 推理与知识 | ||||||
| GPQA Diamond | 93.5 | 92.6 | 94.1 | 91.0 | 93.5 | 91.2 |
| CritPt | 23.4 | 28.6 | 32.3 | 20.9 | 27.1 | 20.9 |
| AA-LCR | 74.7 | 70.0 | 73.7 | 67.7 | 74.3 | 71.3 |
| HLE-Full | 43.5 / 56.0 | 53.3 / 63.0 | 44.5 / 58.0 | 49.8 / 57.9 | 41.4 / 52.2 | — |
| 编码 | ||||||
| DeepSWE | 67.5 | 70.0 | 73.0 | 59.0 | 67.0 | 46.2 |
| ProgramBench | 77.8 | 76.8 | 77.6 | 71.9 | 70.8 | 63.7 |
| Terminal-Bench 2.1 | 88.3 | 88.0 | 88.8 | 84.6 | 83.4 | 82.7 |
| FrontierSWE | 81.2 | 86.6 | 71.3 | 66.7 | 64.9 | 67.3 |
| SWE-Marathon | 42.0 | 35.0 | 39.0 | 40.0 | 14.0 | 13.0 |
| PostTrainBench | 36.6 | 41.4 | 34.6 | 34.1 | 28.4 | 34.3 |
| MLS-Bench-Lite | 48.3 | 49.9 | 46.2 | 42.8 | 35.5 | 40.4 |
| SciCode | 58.7 | 60.2 | 56.1 | 53.5 | 56.1 | 50.5 |
| Kimi Code Bench 2.0 | 72.9 | 76.9 | 64.8 | 71.7 | 69.0 | 64.2 |
| 智能体能力 | ||||||
| BrowseComp | 91.2 | 88.0 | 90.4 | 84.3 | 84.4 | — |
| DeepSearchQA (F1) | 95.0 | 94.2 | — | 93.1 | — | — |
| ResearchRubrics | 76.2 | — | 73.8 | 73.5 | 64.0 | 71.1 |
| GDPval-AA v2 (Elo) | 1686 | 1747 | 1736 | 1593 | 1491 | 1510 |
| Toolathlon-Verified | 76.5 | 77.9 | 74.9 | 76.2 | 73.5 | 59.9 |
| MCPMark-Verified | 94.5 | 87.4 | 92.9 | 76.4 | 92.9 | — |
| MCP-Atlas | 84.2 | 84.7 | 83.6 | 83.6 | 82.8 | 82.6 |
| AutomationBench | 30.8 | 29.1 | 29.7 | 27.2 | 22.7 | 12.9 |
| JobBench | 54.3 | 57.4 | 45.4 | 48.4 | 38.3 | 43.4 |
| AA-Briefcase (Elo) | 1548 | 1583 | 1495 | 1354 | 1158 | 1260 |
| Agents' Last Exam | 28.3 | 25.7† | 29.6 | 27.0 | 26.6 | 20.4 |
| APEX-Agents | 41.0 | 43.3 | 39.9 | 39.4 | 38.5 | 35.6 |
| OfficeQA Pro | 63.3 | 69.9 | 63.2 | 63.9 | 60.9 | 41.4 |
| SpreadsheetBench 2 | 34.8 | 34.7 | 32.4 | 31.6 | 29.1 | 28.1 |
| OSWorld-Verified | 84.8 | 85.0 | 83.0 | 83.4 | 79.0 | — |
| OSWorld 2.0 | 58.3 | 66.1 | 62.6 | 55.7 | 49.5 | — |
| SaaS-Bench | 60.1 | — | 61.4 | 56.1 | 43.8 | — |
| τ³-Banking | 33.4 | 26.8 | 33.0 | 27.6 | 31.3 | 26.8 |
| Harvey Lab-AA | 94.6 | 93.6 | 87.2 | 91.1 | 86.3 | 91.0 |
| CorpFin v2 | 71.6 | 71.8 | 64.4 | |||
所有 Kimi K3 的结果均在推理力度设为“max”且温度值为 1.0 的条件下获得。对于单步任务,如 GPQA Diamond、HLE-Full 以及无需工具的视觉基准测试,我们将 top-p 设置为 0.95;对于智能体任务,top-p 设置为 1.0。对于 HLE-Full、MMMU-Pro、CharXiv (RQ)、MathVision 和 ZeroBench,每个单元格均按顺序报告未使用工具增强和使用工具增强(HLE-Full 使用通用工具,视觉基准测试使用 Python)情况下的分数。
Kimi K3 从 SFT 阶段开始应用量化感知训练,采用 MXFP4 权重与 MXFP8 激活值,以实现广泛的硬件兼容性。
[!Note] 您可以通过访问 https://platform.kimi.ai 并选择
kimi-k3来使用 Kimi K3 的 API,我们提供与 OpenAI/Anthropic 兼容的 API。目前,建议在以下推理引擎上运行 Kimi K3:
Kimi K3 始终启用思考功能,并会返回 reasoning_content。思考力度通过顶层 reasoning_effort 请求字段进行配置,支持 "low"、"high" 和 "max"(默认值为 "max")。
Kimi K3 是在保留思考历史的模式下训练的。对于多轮对话和工具调用,Kimi K3 要求将 API 返回的完整助手消息原样传递回 messages 中——包括 reasoning_content 和 tool_calls,而不仅仅是 content:
import openai
def chat_with_preserved_thinking(client: openai.OpenAI, model_name: str):
messages = [
{
"role": "user",
"content": "Tell me three random numbers."
},
{
"role": "assistant",
"reasoning_content": "I'll start by listing five numbers: 473, 921, 235, 215, 222, and I'll tell you the first three.",
"content": "473, 921, 235"
},
{
"role": "user",
"content": "What are the other two numbers you have in mind?"
}
]
response = client.chat.completions.create(
model=model_name,
messages=messages,
stream=False,
max_tokens=4096,
reasoning_effort="max",
)
# the assistant should mention 215 and 222 that appear in the prior reasoning content
print(f"response: {response.choices[0].message.reasoning}")
return response.choices[0].message.content如需完整指南和示例(视觉输入、结构化输出、部分模式、工具选择、动态工具加载、上下文缓存),请参阅 Kimi K3 快速入门 和 思考深度。
Kimi K3 与 Kimi Code CLI 作为其代理框架配合使用时效果最佳。我们诚挚邀请您试用 — 在终端中运行 Kimi Code,并使用 /model 命令选择 Kimi K3。希望您喜欢使用 Kimi K3 进行构建,我们也非常期待听到您的反馈!
代码仓库和模型权重均根据 Kimi K3 许可协议 发布。
如有任何问题,请通过 support@moonshot.ai 与我们联系。