Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #29 · Source label: Claude Opus 4.7 (Non-reasoning, High Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 44
- Percentile
- 94%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
94% percentile inside its fair comparison set44Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #26 · Source label: Claude Opus 4.7 (Non-reasoning, High Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 44.7%
- Percentile
- 93.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
93.2% percentile inside its fair comparison set44.7%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #63 · Source label: Claude Opus 4.7 (Non-reasoning, High Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 45.9%
- Percentile
- 83.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
83.1% percentile inside its fair comparison set45.9%Raw benchmark value
IFBench
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #151 · Source label: Claude Opus 4.7 (Non-reasoning, High Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 43.6%
- Percentile
- 56.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `ifbench`.
56.5% percentile inside its fair comparison set43.6%Raw benchmark value
Blended price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #282 · Source label: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $10 /1M tokens
- Percentile
- 7.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mBlended0To3To1`.
7.6% percentile inside its fair comparison set$10 /1M tokensRaw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #284 · Source label: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $5 /1M input tokens
- Percentile
- 8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
8% percentile inside its fair comparison set$5 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #282 · Source label: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $25 /1M output tokens
- Percentile
- 7.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
7.6% percentile inside its fair comparison set$25 /1M output tokensRaw benchmark value
Output Speed
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #196 · Source label: Claude Opus 4.7 (Non-reasoning, High Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 43.6 tokens/s
- Percentile
- 12.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianOutputTokensPerSecond`.
12.9% percentile inside its fair comparison set43.6 tokens/sRaw benchmark value
Time to first token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #191 · Source label: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 15.61s
- Percentile
- 15.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstTokenSeconds`.
15.2% percentile inside its fair comparison set15.61sRaw benchmark value
Time to first answer token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #108 · Source label: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 15.61s
- Percentile
- 52.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstAnswerTokenSeconds`.
52.2% percentile inside its fair comparison set15.61sRaw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,502
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: overall. Source rank: #3. Votes: 32629. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,502Raw benchmark valueCI 1,498 - 1,507
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,486
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: creative_writing. Source rank: #3. Votes: 5500. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,486Raw benchmark valueCI 1,477 - 1,495
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,509
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: english. Source rank: #3. Votes: 15799. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,509Raw benchmark valueCI 1,503 - 1,514
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,520
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: exclude_ties. Source rank: #3. Votes: 25085. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,520Raw benchmark valueCI 1,514 - 1,526
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,524
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: hard_prompts. Source rank: #4. Votes: 21557. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,524Raw benchmark valueCI 1,519 - 1,530
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,527
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: hard_prompts_english. Source rank: #4. Votes: 10929. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,527Raw benchmark valueCI 1,520 - 1,534
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,504
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: instruction_following. Source rank: #3. Votes: 11075. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,504Raw benchmark valueCI 1,497 - 1,511
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,514
- Percentile
- 99.3%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: longer_query. Source rank: #4. Votes: 14253. Organization: anthropic. License: Proprietary.
99.3% percentile inside its fair comparison set1,514Raw benchmark valueCI 1,508 - 1,521
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,517
- Percentile
- 99.7%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: multi_turn. Source rank: #2. Votes: 5605. Organization: anthropic. License: Proprietary.
99.7% percentile inside its fair comparison set1,517Raw benchmark valueCI 1,508 - 1,526
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,489
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: overall. Source rank: #4. Votes: 32629. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,489Raw benchmark valueCI 1,485 - 1,494
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #5 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,481
- Percentile
- 98.8%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: creative_writing. Source rank: #6. Votes: 5500. Organization: anthropic. License: Proprietary.
98.8% percentile inside its fair comparison set1,481Raw benchmark valueCI 1,472 - 1,490
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,493
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: english. Source rank: #4. Votes: 15799. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,493Raw benchmark valueCI 1,488 - 1,499
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,500
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: exclude_ties. Source rank: #4. Votes: 25085. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,500Raw benchmark valueCI 1,494 - 1,506
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,503
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: hard_prompts. Source rank: #4. Votes: 21557. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,503Raw benchmark valueCI 1,497 - 1,508
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,505
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: hard_prompts_english. Source rank: #4. Votes: 10929. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,505Raw benchmark valueCI 1,498 - 1,512
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,497
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: instruction_following. Source rank: #4. Votes: 11075. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,497Raw benchmark valueCI 1,490 - 1,504
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3 · Source label: claude-opus-4-7-thinking
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,502
- Percentile
- 99.3%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7-thinking`. Category: longer_query. Source rank: #4. Votes: 14253. Organization: anthropic. License: Proprietary.
99.3% percentile inside its fair comparison set1,502Raw benchmark valueCI 1,496 - 1,509
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,505
- Percentile
- 99.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-4-7`. Category: multi_turn. Source rank: #3. Votes: 5978. Organization: anthropic. License: Proprietary.
99.4% percentile inside its fair comparison set1,505Raw benchmark valueCI 1,496 - 1,514