The AI tools index that doesn't waste your time.不浪费你时间的 AI 工具索引。
Preference · weight 5%偏好 · 权重 5%

LMArena (Chatbot Arena)

official site官网

What it tests: Millions of blind side-by-side human votes: users compare two anonymous models and pick the better answer. Scores are Bradley-Terry ratings (Elo-like).

考什么:数百万次真人盲测投票:用户并排看两个匿名模型的回答,选更好的那个。分数是 Bradley-Terry 评分(类 Elo)。

Why it counts: The only large-scale measure of what humans actually prefer — but 2025's 'Leaderboard Illusion' paper showed big labs can game it, so we cap its weight at 5%.

为什么算数:唯一大规模反映'真人真实偏好'的数据——但 2025 年《排行榜幻觉》论文实锤了大厂能刷榜,所以权重压到 5% 并保留争议说明。

Maintainer维护方Arena AI, Inc.(原 UC Berkeley LMSYS)
License许可CC-BY-4.0(官方数据集)
Reference point (prior-gen SOTA)参考点(上一代 SOTA)1450

Standings实时排名

48 models with sourced scores 个模型有溯源成绩
#Model模型Score成绩Index指数Dated日期Measured by测评方
1GPT-5.6 SolOpenAI1623⚡max1002026-07-31official官方榜source ↗×10⚠±142.2
2GPT-5.6 TerraOpenAI1522⚡max752026-07-31official官方榜source ↗×5⚠±57
3Claude Fable 5Anthropic1506712026-08-12official官方榜source ↗×8
4Muse Spark 1.2meta1498692026-08-12official官方榜source ↗
5Qwen3.8 Max (0902)Qwen1491672026-08-12official官方榜source ↗×3⚠±367
6Kimi K3Moonshot AI1489⚡max672026-08-12official官方榜source ↗×8⚠±193
7Claude Opus 5Anthropic1489⚡max672026-08-12official官方榜source ↗×10⚠±210
8Muse Spark 1.1meta1489672026-08-12official官方榜source ↗
9Gemini 3.1 Pro PreviewGoogle1486662026-08-12official官方榜source ↗×8
10Gemini 3.6 FlashGoogle1484⚡high652026-08-12official官方榜source ↗×6⚠±64
11GPT-5.5OpenAI1482⚡high652026-08-12official官方榜source ↗×4
12GPT-5.6 Sol ProOpenAI1481⚡max652026-08-12official官方榜source ↗
13Gemini 3.5 FlashGoogle1477642026-08-12official官方榜source ↗×6
14Grok 4.20xAI1475632026-08-12official官方榜source ↗×3⚠±38
15Qwen3.7 FlashQwen1475633rd-party第三方source ↗
16Qwen3.7 MaxQwen1474632026-08-12official官方榜source ↗×5
17Claude Opus 4.8Anthropic1474⚡high63official官方榜source ↗×5⚠±38
18GLM 5.2Z.ai1471622026-08-12official官方榜source ↗×4
19Grok 4.5xAI1469622026-08-12official官方榜source ↗×5⚠±50
20MiMo-V2.5-ProXiaomi1468612026-08-12official官方榜source ↗
21GLM 5.1Z.ai1467612026-08-12official官方榜source ↗×3
22GPT-5.6 Terra ProOpenAI1464⚡max602026-08-12official官方榜source ↗×2⚠±58
23Claude Sonnet 5Anthropic1462⚡high602026-08-12official官方榜source ↗×7
24Kimi K2.7 CodeMoonshot AI146160official官方榜source ↗
25DeepSeek V4 Pro 0423DeepSeek1458592026-08-12official官方榜source ↗×5⚠±39.95
26Qwen3.7 PlusQwen1458592026-08-12official官方榜source ↗×2
27Gemini 3.5 Flash LiteGoogle1458592026-08-12official官方榜source ↗×3
28Hy3Tencent1457592026-08-12official官方榜source ↗
29GPT-5.6 LunaOpenAI145257official官方榜source ↗
30GPT-5.6 Luna ProOpenAI1450⚡max572026-08-12official官方榜source ↗
31MiniMax M3MiniMax1444552026-08-12official官方榜source ↗×4
32Grok 4.3xAI1442552026-08-12official官方榜source ↗×3⚠±95
33DeepSeek V4 Flash 0423DeepSeek1435532026-08-12official官方榜source ↗×4⚠±41
34Gemini 3.1 Flash LiteGoogle1432522026-08-12official官方榜source ↗×4⚠±41
35Kimi K2 ThinkingMoonshot AI1430⚡high522026-08-12official官方榜source ↗×3
36Nemotron 3 UltraNVIDIA1427512026-08-12official官方榜source ↗
37Mistral Medium 3.5Mistral AI1427512026-08-12official官方榜source ↗×2
38DeepSeek V3.2DeepSeek1425512026-08-12official官方榜source ↗×4⚠±43
39Llama 4 MaverickMeta141749official官方榜source ↗×9
40MiniMax M2.7MiniMax1416482026-08-12official官方榜source ↗×2
41Mistral LargeMistral AI1415482026-08-12official官方榜source ↗×3⚠±60
42Claude Haiku 4.5Anthropic1413482026-08-12official官方榜source ↗×6⚠±90
43Step 3.7 FlashStepFun139543official官方榜source ↗×2
44Llama 4 ScoutMeta1390423rd-party第三方source ↗×2⚠±109
45gpt-oss-120bOpenAI1365363rd-party第三方source ↗×2
46Command ACohere135433official官方榜source ↗×2
47Seed-2.0-LiteByteDance Seed135232official官方榜source ↗
48Phi 4Microsoft125683rd-party第三方source ↗×2

How scores become the Intelligence Score →成绩如何合成智能评分 →