Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 54
- Percentile
- 99.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
99.5% percentile inside its fair comparison set54Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #7 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 64.6%
- Percentile
- 98.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
98.7% percentile inside its fair comparison set64.6%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #158 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 32.4%
- Percentile
- 65.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
65.3% percentile inside its fair comparison set32.4%Raw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #321 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $4 /1M input tokens
- Percentile
- 11.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
11.6% percentile inside its fair comparison set$4 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #322 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $20 /1M output tokens
- Percentile
- 11%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
11% percentile inside its fair comparison set$20 /1M output tokensRaw benchmark value
Output Speed
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #155 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 74.4 tokens/s
- Percentile
- 41.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianOutputTokensPerSecond`.
41.2% percentile inside its fair comparison set74.4 tokens/sRaw benchmark value
Time to first token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #230 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 37.37s
- Percentile
- 13.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstTokenSeconds`.
13.3% percentile inside its fair comparison set37.37sRaw benchmark value
Time to first answer token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #193 · Source label: Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 37.37s
- Percentile
- 26.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstAnswerTokenSeconds`.
26.7% percentile inside its fair comparison set37.37sRaw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,504
- Percentile
- 99.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: overall. Source rank: #4. Votes: 4552. Organization: anthropic. License: Proprietary.
99.2% percentile inside its fair comparison set1,504Raw benchmark valueCI 1,495 - 1,513
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,516
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: creative_writing. Source rank: #2. Votes: 1058. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,516Raw benchmark valueCI 1,497 - 1,536
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #6 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,506
- Percentile
- 98.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: english. Source rank: #6. Votes: 1719. Organization: anthropic. License: Proprietary.
98.7% percentile inside its fair comparison set1,506Raw benchmark valueCI 1,492 - 1,520
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,523
- Percentile
- 99.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: exclude_ties. Source rank: #4. Votes: 3283. Organization: anthropic. License: Proprietary.
99.2% percentile inside its fair comparison set1,523Raw benchmark valueCI 1,510 - 1,536
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,534
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: hard_prompts. Source rank: #2. Votes: 2837. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,534Raw benchmark valueCI 1,523 - 1,545
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,536
- Percentile
- 99.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: hard_prompts_english. Source rank: #3. Votes: 1083. Organization: anthropic. License: Proprietary.
99.5% percentile inside its fair comparison set1,536Raw benchmark valueCI 1,518 - 1,553
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,517
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: instruction_following. Source rank: #2. Votes: 1550. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,517Raw benchmark valueCI 1,502 - 1,533
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,526
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: longer_query. Source rank: #2. Votes: 1929. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,526Raw benchmark valueCI 1,512 - 1,540
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #9 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,500
- Percentile
- 97.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: multi_turn. Source rank: #9. Votes: 567. Organization: anthropic. License: Proprietary.
97.9% percentile inside its fair comparison set1,500Raw benchmark valueCI 1,476 - 1,525
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,512
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: overall. Source rank: #2. Votes: 4552. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,512Raw benchmark valueCI 1,503 - 1,521
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #1 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,532
- Percentile
- 100%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: creative_writing. Source rank: #1. Votes: 1058. Organization: anthropic. License: Proprietary.
100% percentile inside its fair comparison set1,532Raw benchmark valueCI 1,512 - 1,551
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,513
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: english. Source rank: #2. Votes: 1719. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,513Raw benchmark valueCI 1,498 - 1,527
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,532
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: exclude_ties. Source rank: #2. Votes: 3283. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,532Raw benchmark valueCI 1,519 - 1,545
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,535
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: hard_prompts. Source rank: #2. Votes: 2837. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,535Raw benchmark valueCI 1,524 - 1,547
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,537
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: hard_prompts_english. Source rank: #2. Votes: 1083. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,537Raw benchmark valueCI 1,519 - 1,555
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,534
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: instruction_following. Source rank: #2. Votes: 1550. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,534Raw benchmark valueCI 1,519 - 1,549
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,534
- Percentile
- 99.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: longer_query. Source rank: #2. Votes: 1929. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,534Raw benchmark valueCI 1,520 - 1,548
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #8 · Source label: claude-opus-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,501
- Percentile
- 98.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5.5-high`. Category: multi_turn. Source rank: #8. Votes: 567. Organization: anthropic. License: Proprietary.
98.1% percentile inside its fair comparison set1,501Raw benchmark valueCI 1,477 - 1,526