Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #14
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,489
- Percentile
- 96.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: overall. Source rank: #14. Votes: 28937. Organization: anthropic. License: Proprietary.
96.5% percentile inside its fair comparison set1,489Raw benchmark valueCI 1,485 - 1,494
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #15
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,471
- Percentile
- 96.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: creative_writing. Source rank: #15. Votes: 6526. Organization: anthropic. License: Proprietary.
96.3% percentile inside its fair comparison set1,471Raw benchmark valueCI 1,462 - 1,479
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #20
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,491
- Percentile
- 94.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: english. Source rank: #20. Votes: 11702. Organization: anthropic. License: Proprietary.
94.9% percentile inside its fair comparison set1,491Raw benchmark valueCI 1,484 - 1,497
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #14
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,503
- Percentile
- 96.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: exclude_ties. Source rank: #14. Votes: 21814. Organization: anthropic. License: Proprietary.
96.5% percentile inside its fair comparison set1,503Raw benchmark valueCI 1,497 - 1,510
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #13
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,514
- Percentile
- 96.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: hard_prompts. Source rank: #13. Votes: 18823. Organization: anthropic. License: Proprietary.
96.8% percentile inside its fair comparison set1,514Raw benchmark valueCI 1,509 - 1,520
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #15
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,512
- Percentile
- 96.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: hard_prompts_english. Source rank: #15. Votes: 7522. Organization: anthropic. License: Proprietary.
96.3% percentile inside its fair comparison set1,512Raw benchmark valueCI 1,505 - 1,520
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #10
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,493
- Percentile
- 97.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: instruction_following. Source rank: #10. Votes: 10126. Organization: anthropic. License: Proprietary.
97.6% percentile inside its fair comparison set1,493Raw benchmark valueCI 1,486 - 1,500
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #14
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,500
- Percentile
- 96.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: longer_query. Source rank: #14. Votes: 13751. Organization: anthropic. License: Proprietary.
96.3% percentile inside its fair comparison set1,500Raw benchmark valueCI 1,493 - 1,506
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #28
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,487
- Percentile
- 92.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: multi_turn. Source rank: #28. Votes: 4395. Organization: anthropic. License: Proprietary.
92.8% percentile inside its fair comparison set1,487Raw benchmark valueCI 1,477 - 1,496
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,507
- Percentile
- 99.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: overall. Source rank: #4. Votes: 28937. Organization: anthropic. License: Proprietary.
99.2% percentile inside its fair comparison set1,507Raw benchmark valueCI 1,502 - 1,512
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #6
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,491
- Percentile
- 98.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: creative_writing. Source rank: #6. Votes: 6526. Organization: anthropic. License: Proprietary.
98.7% percentile inside its fair comparison set1,491Raw benchmark valueCI 1,482 - 1,499
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,505
- Percentile
- 98.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: english. Source rank: #5. Votes: 11702. Organization: anthropic. License: Proprietary.
98.9% percentile inside its fair comparison set1,505Raw benchmark valueCI 1,498 - 1,511
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,524
- Percentile
- 99.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: exclude_ties. Source rank: #4. Votes: 21814. Organization: anthropic. License: Proprietary.
99.2% percentile inside its fair comparison set1,524Raw benchmark valueCI 1,518 - 1,531
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,526
- Percentile
- 99.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: hard_prompts. Source rank: #4. Votes: 18823. Organization: anthropic. License: Proprietary.
99.2% percentile inside its fair comparison set1,526Raw benchmark valueCI 1,521 - 1,532
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,523
- Percentile
- 98.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: hard_prompts_english. Source rank: #5. Votes: 7522. Organization: anthropic. License: Proprietary.
98.9% percentile inside its fair comparison set1,523Raw benchmark valueCI 1,515 - 1,531
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,516
- Percentile
- 98.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: instruction_following. Source rank: #5. Votes: 10126. Organization: anthropic. License: Proprietary.
98.9% percentile inside its fair comparison set1,516Raw benchmark valueCI 1,509 - 1,523
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #6
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,514
- Percentile
- 98.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: longer_query. Source rank: #6. Votes: 13751. Organization: anthropic. License: Proprietary.
98.6% percentile inside its fair comparison set1,514Raw benchmark valueCI 1,508 - 1,521
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #9
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,500
- Percentile
- 97.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `claude-opus-5-max`. Category: multi_turn. Source rank: #9. Votes: 4395. Organization: anthropic. License: Proprietary.
97.9% percentile inside its fair comparison set1,500Raw benchmark valueCI 1,490 - 1,509