Head-to-head正面对比
Kimi K3 vs DeepSeek V4 Pro 0423
Moonshot AI meets DeepSeek: sourced benchmark scores side by side, dimension by dimension, with prices and specs. Every conclusion below is derived from the published dataset — no opinions, no sponsorship.
Moonshot AI 对 DeepSeek:带来源的 benchmark 成绩逐项并排,附价格与规格。下面每条结论都从公开数据集推导——没有主观意见,没有厂商赞助。
◈
The verdict, derived from data结论(全部由数据推导)
- On overall score they are nearly tied (85 vs 87.4) — decide on price, context and workload mix.总分几乎打平(85 对 87.4)——按价格、上下文和任务结构来选。
- Kimi K3 is stronger in coding, preference.Kimi K3 在代码、偏好上更强。
- DeepSeek V4 Pro 0423 is stronger in math, agent.DeepSeek V4 Pro 0423 在数学、智能体上更强。
- DeepSeek V4 Pro 0423 is 8.1× cheaper on blended price ($0.633 vs $5.10 per 1M tokens, in/out average).混合价(输入输出均值)上 DeepSeek V4 Pro 0423 便宜 8.1 倍:$0.633 对 $5.10/百万 token。
- Need self-hosting or fine-tuning: DeepSeek V4 Pro 0423 ships open weights.需要自托管或微调:DeepSeek V4 Pro 0423 开放权重。
● Kimi K3 ● DeepSeek V4 Pro 0423 · missing dimensions count as 0 in the chart缺数据的维度在图上按 0 计
≡
Spec by spec逐项规格
| Kimi K3 | DeepSeek V4 Pro 0423 | |
|---|---|---|
| Provider厂商 | Moonshot AI | DeepSeek |
| Released发布 | 2026-07-16 | 2026-04-24 |
| Intelligence Score智能评分 | 85 | 87.4 |
| Benchmark coveragebenchmark 覆盖 | 95% | 95% |
| Reasoning推理 | 87.5 | 89.2 |
| Coding代码 | 97.7 | 85.9 |
| Knowledge知识 | 95 | 94.5 |
| Math数学 | 70.7 | 85.5 |
| Agent智能体 | 37.7 | 97.7 |
| Preference偏好 | 66.5 | 58.8 |
| API input $/1M输入价 $/百万 | $1.70 | $0.422 |
| API output $/1M输出价 $/百万 | $8.50 | $0.845 |
| Blended $/1M混合价 $/百万 | $5.10 | $0.633 |
| Value (score per $)性价比(分数/美元) | 16.7 | 138 |
| Context window上下文 | 1M | 1M |
| Max output最大输出 | 944K | 384K |
| Reasoning tiers推理档位 | always-on thinking | thinking off / high effort / max effort |
| Open weights开放权重 | No否 | Yes支持 |
| CN-direct国内直连 | Yes支持 | Yes支持 |
| Free tier免费层 | No否 | No否 |
⚔
Benchmark by benchmark逐项 benchmark
| Benchmark基准 | Kimi K3 | DeepSeek V4 Pro 0423 |
|---|---|---|
| Humanity's Last Exam | 44.3 | 48.2⚡max |
| GPQA Diamond | 93.5 | 90.1 |
| MMLU-Pro | 88 | 87.5 |
| SWE-bench Verified | 93.4 | 80.6 |
| AIME 2025 | 96.2 | 96.67⚡max |
| LiveCodeBench | 93.5 | 93.5⚡max |
| FrontierMath | 39⚡max | 68.3 |
| Terminal-Bench | 88.3 | 67.9⚡max |
| τ²-bench | 37.1⚡max | 96.2⚡max |
| LMArena (Chatbot Arena) | 1489⚡max | 1458 |
Kimi K3 dossier → 档案 → · DeepSeek V4 Pro 0423 dossier → 档案 → · How we score评分方法