Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #201 · Source label: DeepSeek V3.2 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 16
- Percentile
- 64%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
64% percentile inside its fair comparison set16Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #177 · Source label: DeepSeek V3.2 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 24%
- Percentile
- 61.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
61.1% percentile inside its fair comparison set24%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #395 · Source label: DeepSeek V3.2 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 6.7%
- Percentile
- 13%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
13% percentile inside its fair comparison set6.7%Raw benchmark value
IFBench
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #114 · Source label: DeepSeek V3.2 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 49%
- Percentile
- 67.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `ifbench`.
67.5% percentile inside its fair comparison set49%Raw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #127 · Source label: DeepSeek V3.2 (Reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $0.3 /1M input tokens
- Percentile
- 64.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
64.7% percentile inside its fair comparison set$0.3 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #75 · Source label: DeepSeek V3.2 (Reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $0.4 /1M output tokens
- Percentile
- 79.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
79.4% percentile inside its fair comparison set$0.4 /1M output tokensRaw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #109 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,425
- Percentile
- 71.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: overall. Source rank: #122. Votes: 47999. Organization: deepseek. License: MIT.
71.3% percentile inside its fair comparison set1,425Raw benchmark valueCI 1,421 - 1,428
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #98 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,401
- Percentile
- 74.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: creative_writing. Source rank: #108. Votes: 6858. Organization: deepseek. License: MIT.
74.1% percentile inside its fair comparison set1,401Raw benchmark valueCI 1,393 - 1,408
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #107 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,438
- Percentile
- 71.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: english. Source rank: #119. Votes: 19559. Organization: deepseek. License: MIT.
71.8% percentile inside its fair comparison set1,438Raw benchmark valueCI 1,433 - 1,443
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #108 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,416
- Percentile
- 71.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: exclude_ties. Source rank: #121. Votes: 33948. Organization: deepseek. License: MIT.
71.5% percentile inside its fair comparison set1,416Raw benchmark valueCI 1,411 - 1,421
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #102 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,447
- Percentile
- 73.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: hard_prompts. Source rank: #115. Votes: 27043. Organization: deepseek. License: MIT.
73.1% percentile inside its fair comparison set1,447Raw benchmark valueCI 1,443 - 1,451
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #99 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,459
- Percentile
- 73.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: hard_prompts_english. Source rank: #109. Votes: 11084. Organization: deepseek. License: MIT.
73.8% percentile inside its fair comparison set1,459Raw benchmark valueCI 1,452 - 1,465
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #99 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,420
- Percentile
- 73.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: instruction_following. Source rank: #110. Votes: 13529. Organization: deepseek. License: MIT.
73.9% percentile inside its fair comparison set1,420Raw benchmark valueCI 1,414 - 1,426
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #100 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,440
- Percentile
- 72.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: longer_query. Source rank: #111. Votes: 13703. Organization: deepseek. License: MIT.
72.1% percentile inside its fair comparison set1,440Raw benchmark valueCI 1,434 - 1,446
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #105 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,429
- Percentile
- 72.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: multi_turn. Source rank: #119. Votes: 8337. Organization: deepseek. License: MIT.
72.2% percentile inside its fair comparison set1,429Raw benchmark valueCI 1,422 - 1,436
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #96 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,424
- Percentile
- 74.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: overall. Source rank: #103. Votes: 47999. Organization: deepseek. License: MIT.
74.7% percentile inside its fair comparison set1,424Raw benchmark valueCI 1,421 - 1,428
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #93 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,400
- Percentile
- 75.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: creative_writing. Source rank: #102. Votes: 6858. Organization: deepseek. License: MIT.
75.4% percentile inside its fair comparison set1,400Raw benchmark valueCI 1,392 - 1,407
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #99 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,437
- Percentile
- 73.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: english. Source rank: #106. Votes: 19559. Organization: deepseek. License: MIT.
73.9% percentile inside its fair comparison set1,437Raw benchmark valueCI 1,432 - 1,442
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #96 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,415
- Percentile
- 74.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: exclude_ties. Source rank: #103. Votes: 33948. Organization: deepseek. License: MIT.
74.7% percentile inside its fair comparison set1,415Raw benchmark valueCI 1,410 - 1,419
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #93 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,433
- Percentile
- 75.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: hard_prompts. Source rank: #101. Votes: 27043. Organization: deepseek. License: MIT.
75.5% percentile inside its fair comparison set1,433Raw benchmark valueCI 1,429 - 1,438
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #89 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,445
- Percentile
- 76.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: hard_prompts_english. Source rank: #94. Votes: 11084. Organization: deepseek. License: MIT.
76.5% percentile inside its fair comparison set1,445Raw benchmark valueCI 1,439 - 1,451
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #92 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,412
- Percentile
- 75.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: instruction_following. Source rank: #100. Votes: 13529. Organization: deepseek. License: MIT.
75.8% percentile inside its fair comparison set1,412Raw benchmark valueCI 1,407 - 1,418
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #87 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,428
- Percentile
- 75.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: longer_query. Source rank: #94. Votes: 13703. Organization: deepseek. License: MIT.
75.8% percentile inside its fair comparison set1,428Raw benchmark valueCI 1,422 - 1,434
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #95 · Source label: deepseek-v3.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,427
- Percentile
- 74.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `deepseek-v3.2`. Category: multi_turn. Source rank: #105. Votes: 8337. Organization: deepseek. License: MIT.
74.9% percentile inside its fair comparison set1,427Raw benchmark valueCI 1,420 - 1,434