Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #10 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 52
- Percentile
- 98.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
98.4% percentile inside its fair comparison set52Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #37 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 53%
- Percentile
- 92.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
92.1% percentile inside its fair comparison set53%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #135 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 37.1%
- Percentile
- 70.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
70.4% percentile inside its fair comparison set37.1%Raw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #282 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $2 /1M input tokens
- Percentile
- 26.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
26.8% percentile inside its fair comparison set$2 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #285 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $10 /1M output tokens
- Percentile
- 23.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
23.7% percentile inside its fair comparison set$10 /1M output tokensRaw benchmark value
Output Speed
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #109 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 104.9 tokens/s
- Percentile
- 58.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianOutputTokensPerSecond`.
58.8% percentile inside its fair comparison set104.9 tokens/sRaw benchmark value
Time to first token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #227 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 32.46s
- Percentile
- 14.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstTokenSeconds`.
14.4% percentile inside its fair comparison set32.46sRaw benchmark value
Time to first answer token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #182 · Source label: Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 32.46s
- Percentile
- 30.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstAnswerTokenSeconds`.
30.9% percentile inside its fair comparison set32.46sRaw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #42 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,471
- Percentile
- 89.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: overall. Source rank: #45. Votes: 3145. Organization: anthropic. License: Proprietary.
89.1% percentile inside its fair comparison set1,471Raw benchmark valueCI 1,461 - 1,481
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #40 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,447
- Percentile
- 89.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: creative_writing. Source rank: #42. Votes: 658. Organization: anthropic. License: Proprietary.
89.6% percentile inside its fair comparison set1,447Raw benchmark valueCI 1,424 - 1,471
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #17 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,492
- Percentile
- 95.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: english. Source rank: #17. Votes: 1196. Organization: anthropic. License: Proprietary.
95.7% percentile inside its fair comparison set1,492Raw benchmark valueCI 1,475 - 1,509
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #43 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,478
- Percentile
- 88.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: exclude_ties. Source rank: #46. Votes: 2287. Organization: anthropic. License: Proprietary.
88.8% percentile inside its fair comparison set1,478Raw benchmark valueCI 1,464 - 1,493
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #24 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,504
- Percentile
- 93.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: hard_prompts. Source rank: #24. Votes: 1985. Organization: anthropic. License: Proprietary.
93.9% percentile inside its fair comparison set1,504Raw benchmark valueCI 1,491 - 1,517
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #10 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,518
- Percentile
- 97.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: hard_prompts_english. Source rank: #10. Votes: 764. Organization: anthropic. License: Proprietary.
97.6% percentile inside its fair comparison set1,518Raw benchmark valueCI 1,496 - 1,539
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #18 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,484
- Percentile
- 95.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: instruction_following. Source rank: #18. Votes: 1079. Organization: anthropic. License: Proprietary.
95.5% percentile inside its fair comparison set1,484Raw benchmark valueCI 1,466 - 1,502
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #24 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,493
- Percentile
- 93.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: longer_query. Source rank: #24. Votes: 1329. Organization: anthropic. License: Proprietary.
93.5% percentile inside its fair comparison set1,493Raw benchmark valueCI 1,476 - 1,509
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #45 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,473
- Percentile
- 88.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: multi_turn. Source rank: #51. Votes: 438. Organization: anthropic. License: Proprietary.
88.2% percentile inside its fair comparison set1,473Raw benchmark valueCI 1,446 - 1,501
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #32 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,467
- Percentile
- 91.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: overall. Source rank: #34. Votes: 3145. Organization: anthropic. License: Proprietary.
91.8% percentile inside its fair comparison set1,467Raw benchmark valueCI 1,456 - 1,477
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #31 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,450
- Percentile
- 92%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: creative_writing. Source rank: #33. Votes: 658. Organization: anthropic. License: Proprietary.
92% percentile inside its fair comparison set1,450Raw benchmark valueCI 1,427 - 1,474
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #16 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,487
- Percentile
- 96%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: english. Source rank: #16. Votes: 1196. Organization: anthropic. License: Proprietary.
96% percentile inside its fair comparison set1,487Raw benchmark valueCI 1,470 - 1,504
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #33 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,471
- Percentile
- 91.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: exclude_ties. Source rank: #34. Votes: 2287. Organization: anthropic. License: Proprietary.
91.5% percentile inside its fair comparison set1,471Raw benchmark valueCI 1,456 - 1,485
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #17 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,493
- Percentile
- 95.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: hard_prompts. Source rank: #17. Votes: 1985. Organization: anthropic. License: Proprietary.
95.7% percentile inside its fair comparison set1,493Raw benchmark valueCI 1,480 - 1,506
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #11 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,507
- Percentile
- 97.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: hard_prompts_english. Source rank: #11. Votes: 764. Organization: anthropic. License: Proprietary.
97.3% percentile inside its fair comparison set1,507Raw benchmark valueCI 1,486 - 1,528
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #11 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,490
- Percentile
- 97.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: instruction_following. Source rank: #11. Votes: 1079. Organization: anthropic. License: Proprietary.
97.3% percentile inside its fair comparison set1,490Raw benchmark valueCI 1,472 - 1,507
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #16 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,488
- Percentile
- 95.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: longer_query. Source rank: #16. Votes: 1329. Organization: anthropic. License: Proprietary.
95.8% percentile inside its fair comparison set1,488Raw benchmark valueCI 1,472 - 1,505
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #41 · Source label: claude-sonnet-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,463
- Percentile
- 89.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-sonnet-5.5-xhigh`. Category: multi_turn. Source rank: #44. Votes: 438. Organization: anthropic. License: Proprietary.
89.3% percentile inside its fair comparison set1,463Raw benchmark valueCI 1,435 - 1,490