Model dossier模型档案
Grok 4.20
Grok 4.20 is a high-tier model from xAI (Intelligence Score 85.1/100). Strongest in agent (98) and knowledge (93.2); best fit: tool-use and multi-step agents. Mid-priced: $1.25 in / $2.50 out per 1M tokens. Context window 2M tokens — long-document friendly.
Grok 4.20 是 xAI 的高水平梯队模型(智能评分 85.1/100)。 最强项是智能体(98分),其次是知识(93.2分);适合工具调用与多步 Agent。 定价属中档定价:每百万 token 输入 $1.25 / 输出 $2.50。 上下文 2M token,长文档友好。
85.1
Intelligence Score · benchmark coverage 智能评分 · benchmark 覆盖 95%
◈
Benchmark scoresBenchmark 成绩
10 entries 条| Benchmark基准 | Score成绩 | Index指数 | Dated日期 | Measured by测评方 | |
|---|---|---|---|---|---|
| Humanity's Last ExamReasoning | 50.7 | 88 | — | 3rd-party第三方 | source ↗×2⚠±18.5 |
| GPQA DiamondReasoning | 91 | 95 | — | 3rd-party第三方 | source ↗×4⚠±3.5 |
| MMLU-ProKnowledge | 86.3 | 93 | — | 3rd-party第三方 | source ↗×2⚠±8.7 |
| SWE-bench VerifiedCoding | 76.7⚡high | 80 | 2026-04-11 | 3rd-party第三方 | source ↗×2 |
| AIME 2025Math | 91.7 | 92 | — | 3rd-party第三方 | source ↗×3 |
| LiveCodeBenchCoding | 84.27 | 90 | — | 3rd-party第三方 | source ↗ |
| FrontierMathMath | 67.1 | 71 | — | 3rd-party第三方 | source ↗ |
| Terminal-BenchCoding | 47.1⚡high | 51 | — | 3rd-party第三方 | source ↗ |
| τ²-benchAgent | 96.5⚡high | 98 | 2026-06-10 | 3rd-party第三方 | source ↗⚠±26.9 |
| ↳ ⚡off | 69.6 | — | 2026-06-10 | 3rd-party第三方 | source ↗ |
| LMArena (Chatbot Arena)Preference | 1475 | 63 | 2026-08-12 | official官方榜 | source ↗×3⚠±38 |
≡
Specs & pricing规格与价格
| Provider厂商 | xAI |
|---|---|
| Released发布 | 2026-03-31 |
| Context window上下文 | 2M tokens |
| Max output最大输出 | 1.8M tokens |
| Modality模态 | text+image+file->text |
| Reasoning tiers推理档位 | ⚡low ⚡medium ⚡high default默认 high · low / medium / high (cannot disable) |
| API input price输入价格 | $1.25 / 1M tokens |
| API output price输出价格 | $2.50 / 1M tokens |
⛓
Tools using this model使用它的工具
How we score →评分方法 → · Catalog & pricing via the official OpenRouter API.模型目录与价格来自 OpenRouter 官方 API。