The AI tools index that doesn't waste your time.不浪费你时间的 AI 工具索引。
Coding · weight 6%代码 · 权重 6%

Terminal-Bench

official site官网

What it tests: Real command-line tasks inside Docker containers: compiling, debugging, configuring systems. The model acts as an agent in an actual terminal.

考什么:在真实 Docker 容器里完成命令行任务:编译、调试、配环境。模型以 agent 身份在真终端里操作。

Why it counts: The fastest-rising agentic-coding benchmark of 2026 — closest proxy for 'can this model drive my dev environment'.

为什么算数:2026 年上升最快的代码 benchmark——最接近'这个模型能不能直接开动我的开发环境'的真实体验。

Maintainer维护方Harbor Framework 团队
License许可MIT
Reference point (prior-gen SOTA)参考点(上一代 SOTA)60

Standings实时排名

41 models with sourced scores 个模型有溯源成绩
#Model模型Score成绩Index指数Dated日期Measured by测评方
1GPT-5.6 SolOpenAI91.9100official官方榜source ↗×10⚠±57.3
2Kimi K3Moonshot AI88.396official官方榜source ↗×6
3Qwen3.8 Max (0902)Qwen86.6943rd-party第三方source ↗×3
4Claude Fable 5Anthropic83.891official官方榜source ↗×7⚠±53.9
5Grok 4.5xAI83.391official官方榜source ↗×4
6Muse Spark 1.2meta82.9903rd-party第三方source ↗×2
7DeepSeek V4 Flash 0731DeepSeek82.7903rd-party第三方source ↗×4
8DeepSeek V4 Flash 0423DeepSeek82.7902026-07-313rd-party第三方source ↗×3⚠±25.8
9GPT-5.5OpenAI82.790official官方榜source ↗×6
10GLM 5.2Z.ai81883rd-party第三方source ↗×5
11Claude Sonnet 5Anthropic80.487official官方榜source ↗×7
12GPT-5.6 TerraOpenAI78.485official官方榜source ↗×6⚠±9.6
13Gemini 3.6 FlashGoogle78852026-073rd-party第三方source ↗×3⚠±4.22
14Gemini 3.5 FlashGoogle76.2832026-073rd-party第三方source ↗×2
15Muse Spark 1.1meta76.2833rd-party第三方source ↗
16GPT-5.6 LunaOpenAI75.782official官方榜source ↗×6⚠±9
17Claude Opus 4.8Anthropic74.681official官方榜source ↗×5⚠±10.4
18Qwen3.7 PlusQwen70.3773rd-party第三方source ↗
19Qwen3.7 MaxQwen69.7762026-063rd-party第三方source ↗×4⚠±4.8
20Gemini 3.1 Pro PreviewGoogle68.575official官方榜source ↗×6⚠±41.4
21DeepSeek V4 Pro 0423DeepSeek67.9⚡max743rd-party第三方source ↗×5⚠±36.4
22Kimi K2.7 CodeMoonshot AI66.7733rd-party第三方source ↗
23MiniMax M3MiniMax6672official官方榜source ↗×5
24Mistral Medium 3.5Mistral AI66723rd-party第三方source ↗
25Mistral LargeMistral AI66722026-08-013rd-party第三方source ↗×2⚠±42.25
26GLM 5.1Z.ai63.5692026-073rd-party第三方source ↗×5⚠±7.3
27Step 3.7 FlashStepFun59.5653rd-party第三方source ↗
28MiniMax M2.7MiniMax5762official官方榜source ↗×3⚠±17.6
29Gemini 3.5 Flash LiteGoogle54593rd-party第三方source ↗×2
30Grok 4.20xAI47.1⚡high513rd-party第三方source ↗
31DeepSeek V3.2DeepSeek46.450official官方榜source ↗×2
32Seed-2.0-LiteByteDance Seed45493rd-party第三方source ↗
33Claude Opus 5Anthropic43.5⚡max47official官方榜source ↗×7⚠±46.4
34Grok 4.3xAI37.9413rd-party第三方source ↗×2⚠±4.05
35Kimi K2 ThinkingMoonshot AI35.7⚡high39official官方榜source ↗×2
36Claude Haiku 4.5Anthropic27.530official官方榜source ↗×3⚠±13.5
37Mercury 2Inception26.5293rd-party第三方source ↗
38gpt-oss-120bOpenAI18.7203rd-party第三方source ↗
39Gemini 3.1 Flash LiteGoogle17.5193rd-party第三方source ↗
40Llama 4 MaverickMeta8.8103rd-party第三方source ↗
41Llama 4 ScoutMeta8.8103rd-party第三方source ↗

How scores become the Intelligence Score →成绩如何合成智能评分 →