Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #148
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,395
- Percentile
- 60.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: overall. Source rank: #165. Votes: 7805. Organization: alibaba. License: Apache 2.0.
60.9% percentile inside its fair comparison set1,395Raw benchmark valueCI 1,388 - 1,402
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #168
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,340
- Percentile
- 55.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: creative_writing. Source rank: #188. Votes: 1006. Organization: alibaba. License: Apache 2.0.
55.3% percentile inside its fair comparison set1,340Raw benchmark valueCI 1,322 - 1,358
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #157
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,403
- Percentile
- 58.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: english. Source rank: #175. Votes: 3465. Organization: alibaba. License: Apache 2.0.
58.5% percentile inside its fair comparison set1,403Raw benchmark valueCI 1,393 - 1,413
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #151
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,372
- Percentile
- 60.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: exclude_ties. Source rank: #168. Votes: 5499. Organization: alibaba. License: Apache 2.0.
60.1% percentile inside its fair comparison set1,372Raw benchmark valueCI 1,362 - 1,382
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #149
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,417
- Percentile
- 60.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: hard_prompts. Source rank: #166. Votes: 4014. Organization: alibaba. License: Apache 2.0.
60.6% percentile inside its fair comparison set1,417Raw benchmark valueCI 1,408 - 1,426
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #160
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,418
- Percentile
- 57.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: hard_prompts_english. Source rank: #179. Votes: 1813. Organization: alibaba. License: Apache 2.0.
57.5% percentile inside its fair comparison set1,418Raw benchmark valueCI 1,404 - 1,432
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #149
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,382
- Percentile
- 60.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: instruction_following. Source rank: #168. Votes: 2137. Organization: alibaba. License: Apache 2.0.
60.6% percentile inside its fair comparison set1,382Raw benchmark valueCI 1,370 - 1,394
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #144
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,404
- Percentile
- 59.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: longer_query. Source rank: #163. Votes: 1758. Organization: alibaba. License: Apache 2.0.
59.7% percentile inside its fair comparison set1,404Raw benchmark valueCI 1,390 - 1,418
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #164
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,384
- Percentile
- 56.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: multi_turn. Source rank: #183. Votes: 1260. Organization: alibaba. License: Apache 2.0.
56.4% percentile inside its fair comparison set1,384Raw benchmark valueCI 1,368 - 1,401
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #135
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,401
- Percentile
- 64.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: overall. Source rank: #149. Votes: 7805. Organization: alibaba. License: Apache 2.0.
64.4% percentile inside its fair comparison set1,401Raw benchmark valueCI 1,394 - 1,408
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #150
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,347
- Percentile
- 60.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: creative_writing. Source rank: #169. Votes: 1006. Organization: alibaba. License: Apache 2.0.
60.2% percentile inside its fair comparison set1,347Raw benchmark valueCI 1,328 - 1,365
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #123
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,420
- Percentile
- 67.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: english. Source rank: #133. Votes: 3465. Organization: alibaba. License: Apache 2.0.
67.6% percentile inside its fair comparison set1,420Raw benchmark valueCI 1,410 - 1,430
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #137
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,379
- Percentile
- 63.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: exclude_ties. Source rank: #152. Votes: 5499. Organization: alibaba. License: Apache 2.0.
63.8% percentile inside its fair comparison set1,379Raw benchmark valueCI 1,370 - 1,389
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #138
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,406
- Percentile
- 63.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: hard_prompts. Source rank: #151. Votes: 4014. Organization: alibaba. License: Apache 2.0.
63.6% percentile inside its fair comparison set1,406Raw benchmark valueCI 1,397 - 1,416
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #133
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,416
- Percentile
- 64.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: hard_prompts_english. Source rank: #146. Votes: 1813. Organization: alibaba. License: Apache 2.0.
64.7% percentile inside its fair comparison set1,416Raw benchmark valueCI 1,402 - 1,430
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #146
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,374
- Percentile
- 61.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: instruction_following. Source rank: #162. Votes: 2137. Organization: alibaba. License: Apache 2.0.
61.4% percentile inside its fair comparison set1,374Raw benchmark valueCI 1,362 - 1,387
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #142
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,391
- Percentile
- 60.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: longer_query. Source rank: #159. Votes: 1758. Organization: alibaba. License: Apache 2.0.
60.3% percentile inside its fair comparison set1,391Raw benchmark valueCI 1,377 - 1,405
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #150
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,387
- Percentile
- 60.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `qwen3-vl-235b-a22b-thinking`. Category: multi_turn. Source rank: #167. Votes: 1260. Organization: alibaba. License: Apache 2.0.
60.2% percentile inside its fair comparison set1,387Raw benchmark valueCI 1,370 - 1,404