Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #258 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 12
- Percentile
- 53.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
53.7% percentile inside its fair comparison set12Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #246 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 18.7%
- Percentile
- 45.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
45.9% percentile inside its fair comparison set18.7%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #202 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 22.8%
- Percentile
- 55.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
55.6% percentile inside its fair comparison set22.8%Raw benchmark value
IFBench
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #131 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 46.1%
- Percentile
- 62.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `ifbench`.
62.6% percentile inside its fair comparison set46.1%Raw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #114 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $0.2 /1M input tokens
- Percentile
- 68.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
68.1% percentile inside its fair comparison set$0.2 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #125 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $0.9 /1M output tokens
- Percentile
- 65%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
65% percentile inside its fair comparison set$0.9 /1M output tokensRaw benchmark value
Output Speed
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #189 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 60.5 tokens/s
- Percentile
- 28.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianOutputTokensPerSecond`.
28.2% percentile inside its fair comparison set60.5 tokens/sRaw benchmark value
Time to first token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #120 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 2.46s
- Percentile
- 54.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstTokenSeconds`.
54.9% percentile inside its fair comparison set2.46sRaw benchmark value
Time to first answer token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #56 · Source label: Qwen3 235B A22B 2507 Instruct
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 2.46s
- Percentile
- 79%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstAnswerTokenSeconds`.
79% percentile inside its fair comparison set2.46sRaw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #112
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,422
- Percentile
- 70.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: overall. Source rank: #126. Votes: 97788. Organization: alibaba. License: Apache 2.0.
70.5% percentile inside its fair comparison set1,422Raw benchmark valueCI 1,420 - 1,425
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #127
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,378
- Percentile
- 66.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: creative_writing. Source rank: #144. Votes: 13820. Organization: alibaba. License: Apache 2.0.
66.3% percentile inside its fair comparison set1,378Raw benchmark valueCI 1,373 - 1,384
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #118
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,429
- Percentile
- 68.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: english. Source rank: #132. Votes: 40895. Organization: alibaba. License: Apache 2.0.
68.9% percentile inside its fair comparison set1,429Raw benchmark valueCI 1,426 - 1,433
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #112
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,414
- Percentile
- 70.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: exclude_ties. Source rank: #125. Votes: 69623. Organization: alibaba. License: Apache 2.0.
70.5% percentile inside its fair comparison set1,414Raw benchmark valueCI 1,411 - 1,418
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #101
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,447
- Percentile
- 73.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: hard_prompts. Source rank: #114. Votes: 53398. Organization: alibaba. License: Apache 2.0.
73.4% percentile inside its fair comparison set1,447Raw benchmark valueCI 1,444 - 1,450
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #110
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,450
- Percentile
- 70.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: hard_prompts_english. Source rank: #123. Votes: 22661. Organization: alibaba. License: Apache 2.0.
70.9% percentile inside its fair comparison set1,450Raw benchmark valueCI 1,445 - 1,454
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #110
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,414
- Percentile
- 71%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: instruction_following. Source rank: #123. Votes: 27734. Organization: alibaba. License: Apache 2.0.
71% percentile inside its fair comparison set1,414Raw benchmark valueCI 1,410 - 1,418
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #110
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,433
- Percentile
- 69.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: longer_query. Source rank: #123. Votes: 27037. Organization: alibaba. License: Apache 2.0.
69.3% percentile inside its fair comparison set1,433Raw benchmark valueCI 1,429 - 1,437
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #96
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,436
- Percentile
- 74.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: multi_turn. Source rank: #109. Votes: 17057. Organization: alibaba. License: Apache 2.0.
74.6% percentile inside its fair comparison set1,436Raw benchmark valueCI 1,431 - 1,442
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #110
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,419
- Percentile
- 71%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: overall. Source rank: #118. Votes: 97788. Organization: alibaba. License: Apache 2.0.
71% percentile inside its fair comparison set1,419Raw benchmark valueCI 1,417 - 1,422
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #122
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,375
- Percentile
- 67.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: creative_writing. Source rank: #136. Votes: 13820. Organization: alibaba. License: Apache 2.0.
67.6% percentile inside its fair comparison set1,375Raw benchmark valueCI 1,369 - 1,380
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #118
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,424
- Percentile
- 68.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: english. Source rank: #127. Votes: 40895. Organization: alibaba. License: Apache 2.0.
68.9% percentile inside its fair comparison set1,424Raw benchmark valueCI 1,420 - 1,427
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #107
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,408
- Percentile
- 71.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: exclude_ties. Source rank: #115. Votes: 69623. Organization: alibaba. License: Apache 2.0.
71.8% percentile inside its fair comparison set1,408Raw benchmark valueCI 1,404 - 1,411
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #95
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,433
- Percentile
- 75%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: hard_prompts. Source rank: #103. Votes: 53398. Organization: alibaba. License: Apache 2.0.
75% percentile inside its fair comparison set1,433Raw benchmark valueCI 1,430 - 1,437
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #105
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,436
- Percentile
- 72.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: hard_prompts_english. Source rank: #112. Votes: 22661. Organization: alibaba. License: Apache 2.0.
72.2% percentile inside its fair comparison set1,436Raw benchmark valueCI 1,432 - 1,441
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #99
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,408
- Percentile
- 73.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: instruction_following. Source rank: #108. Votes: 27734. Organization: alibaba. License: Apache 2.0.
73.9% percentile inside its fair comparison set1,408Raw benchmark valueCI 1,404 - 1,412
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #91
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,426
- Percentile
- 74.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: longer_query. Source rank: #99. Votes: 27037. Organization: alibaba. License: Apache 2.0.
74.6% percentile inside its fair comparison set1,426Raw benchmark valueCI 1,422 - 1,430
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #88
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,432
- Percentile
- 76.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-instruct-2507`. Category: multi_turn. Source rank: #97. Votes: 17057. Organization: alibaba. License: Apache 2.0.
76.7% percentile inside its fair comparison set1,432Raw benchmark valueCI 1,427 - 1,437