Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #138
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,403
- Percentile
- 63.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: overall. Source rank: #155. Votes: 37279. Organization: alibaba. License: Apache 2.0.
63.6% percentile inside its fair comparison set1,403Raw benchmark valueCI 1,398 - 1,407
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #138
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,366
- Percentile
- 63.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: creative_writing. Source rank: #157. Votes: 4524. Organization: alibaba. License: Apache 2.0.
63.4% percentile inside its fair comparison set1,366Raw benchmark valueCI 1,357 - 1,375
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #154
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,406
- Percentile
- 59.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: english. Source rank: #171. Votes: 16515. Organization: alibaba. License: Apache 2.0.
59.3% percentile inside its fair comparison set1,406Raw benchmark valueCI 1,401 - 1,412
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #140
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,385
- Percentile
- 63%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: exclude_ties. Source rank: #156. Votes: 26521. Organization: alibaba. License: Apache 2.0.
63% percentile inside its fair comparison set1,385Raw benchmark valueCI 1,379 - 1,391
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #140
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,420
- Percentile
- 63%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: hard_prompts. Source rank: #155. Votes: 16329. Organization: alibaba. License: Apache 2.0.
63% percentile inside its fair comparison set1,420Raw benchmark valueCI 1,415 - 1,426
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #157
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,419
- Percentile
- 58.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: hard_prompts_english. Source rank: #176. Votes: 7349. Organization: alibaba. License: Apache 2.0.
58.3% percentile inside its fair comparison set1,419Raw benchmark valueCI 1,412 - 1,427
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #154
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,381
- Percentile
- 59.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: instruction_following. Source rank: #173. Votes: 9162. Organization: alibaba. License: Apache 2.0.
59.3% percentile inside its fair comparison set1,381Raw benchmark valueCI 1,374 - 1,388
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #136
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,410
- Percentile
- 62%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: longer_query. Source rank: #155. Votes: 7120. Organization: alibaba. License: Apache 2.0.
62% percentile inside its fair comparison set1,410Raw benchmark valueCI 1,402 - 1,418
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #133
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,410
- Percentile
- 64.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: multi_turn. Source rank: #149. Votes: 6536. Organization: alibaba. License: Apache 2.0.
64.7% percentile inside its fair comparison set1,410Raw benchmark valueCI 1,402 - 1,418
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #144
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,394
- Percentile
- 62%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: overall. Source rank: #160. Votes: 37279. Organization: alibaba. License: Apache 2.0.
62% percentile inside its fair comparison set1,394Raw benchmark valueCI 1,389 - 1,398
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #140
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,356
- Percentile
- 62.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: creative_writing. Source rank: #157. Votes: 4524. Organization: alibaba. License: Apache 2.0.
62.8% percentile inside its fair comparison set1,356Raw benchmark valueCI 1,347 - 1,366
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #148
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,397
- Percentile
- 60.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: english. Source rank: #164. Votes: 16515. Organization: alibaba. License: Apache 2.0.
60.9% percentile inside its fair comparison set1,397Raw benchmark valueCI 1,391 - 1,402
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #145
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,371
- Percentile
- 61.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: exclude_ties. Source rank: #161. Votes: 26521. Organization: alibaba. License: Apache 2.0.
61.7% percentile inside its fair comparison set1,371Raw benchmark valueCI 1,365 - 1,377
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #151
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,393
- Percentile
- 60.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: hard_prompts. Source rank: #168. Votes: 16329. Organization: alibaba. License: Apache 2.0.
60.1% percentile inside its fair comparison set1,393Raw benchmark valueCI 1,388 - 1,399
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #153
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,393
- Percentile
- 59.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: hard_prompts_english. Source rank: #170. Votes: 7349. Organization: alibaba. License: Apache 2.0.
59.4% percentile inside its fair comparison set1,393Raw benchmark valueCI 1,385 - 1,400
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #156
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,362
- Percentile
- 58.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: instruction_following. Source rank: #174. Votes: 9162. Organization: alibaba. License: Apache 2.0.
58.8% percentile inside its fair comparison set1,362Raw benchmark valueCI 1,355 - 1,368
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #146
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,389
- Percentile
- 59.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: longer_query. Source rank: #163. Votes: 7120. Organization: alibaba. License: Apache 2.0.
59.2% percentile inside its fair comparison set1,389Raw benchmark valueCI 1,382 - 1,397
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #137
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,398
- Percentile
- 63.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-235b-a22b-no-thinking`. Category: multi_turn. Source rank: #152. Votes: 6536. Organization: alibaba. License: Apache 2.0.
63.6% percentile inside its fair comparison set1,398Raw benchmark valueCI 1,390 - 1,406