Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #145 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 22
- Percentile
- 69.2%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
69.2% percentile inside its fair comparison set22Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #217 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 16.6%
- Percentile
- 41%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
41% percentile inside its fair comparison set16.6%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #186 · Source label: Gemma 4 31B (Reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 15%
- Percentile
- 49.5%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
49.5% percentile inside its fair comparison set15%Raw benchmark value
IFBench
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #93 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 53.5%
- Percentile
- 73.3%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `ifbench`.
73.3% percentile inside its fair comparison set53.5%Raw benchmark value
Blended price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #70 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $0.2 /1M tokens
- Percentile
- 77.1%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `price1mBlended0To3To1`.
77.1% percentile inside its fair comparison set$0.2 /1M tokensRaw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #71 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $0.2 /1M input tokens
- Percentile
- 76.7%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
76.7% percentile inside its fair comparison set$0.2 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #66 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $0.4 /1M output tokens
- Percentile
- 78.4%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
78.4% percentile inside its fair comparison set$0.4 /1M output tokensRaw benchmark value
Output Speed
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #212 · Source label: Gemma 4 31B (Reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 35.9 tokens/s
- Percentile
- 5.8%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `medianOutputTokensPerSecond`.
5.8% percentile inside its fair comparison set35.9 tokens/sRaw benchmark value
Time to first token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #62 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 1.41s
- Percentile
- 72.8%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstTokenSeconds`.
72.8% percentile inside its fair comparison set1.41sRaw benchmark value
Time to first answer token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #195 · Source label: Gemma 4 31B (Reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 49.48s
- Percentile
- 13.4%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstAnswerTokenSeconds`.
13.4% percentile inside its fair comparison set49.48sRaw benchmark value
Openness Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #174 · Source label: Gemma 4 31B (Non-reasoning)
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 39
- Percentile
- 52.1%
- Last updated
- recent
- Eligibility
- benchmark_derived_model
Parsed from Artificial Analysis public leaderboard field `opennessBreakdown.opennessIndex`.
52.1% percentile inside its fair comparison set39Raw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #32
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,451
- Percentile
- 90.5%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: overall. Source rank: #43. Votes: 5884. Organization: google. License: Apache 2.0.
90.5% percentile inside its fair comparison set1,451Raw benchmark valueCI 1,443 - 1,458
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #40
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,422
- Percentile
- 87.9%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: creative_writing. Source rank: #52. Votes: 940. Organization: google. License: Apache 2.0.
87.9% percentile inside its fair comparison set1,422Raw benchmark valueCI 1,402 - 1,441
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #38
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,459
- Percentile
- 88.6%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: english. Source rank: #49. Votes: 2603. Organization: google. License: Apache 2.0.
88.6% percentile inside its fair comparison set1,459Raw benchmark valueCI 1,448 - 1,470
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #32
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,454
- Percentile
- 90.5%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: exclude_ties. Source rank: #43. Votes: 4043. Organization: google. License: Apache 2.0.
90.5% percentile inside its fair comparison set1,454Raw benchmark valueCI 1,442 - 1,465
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #34
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,473
- Percentile
- 89.8%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: hard_prompts. Source rank: #46. Votes: 3367. Organization: google. License: Apache 2.0.
89.8% percentile inside its fair comparison set1,473Raw benchmark valueCI 1,463 - 1,483
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #30
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,482
- Percentile
- 91%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: hard_prompts_english. Source rank: #39. Votes: 1533. Organization: google. License: Apache 2.0.
91% percentile inside its fair comparison set1,482Raw benchmark valueCI 1,468 - 1,497
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #25
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,452
- Percentile
- 92.6%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: instruction_following. Source rank: #33. Votes: 1661. Organization: google. License: Apache 2.0.
92.6% percentile inside its fair comparison set1,452Raw benchmark valueCI 1,438 - 1,466
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #28
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,467
- Percentile
- 91.1%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: longer_query. Source rank: #36. Votes: 1657. Organization: google. License: Apache 2.0.
91.1% percentile inside its fair comparison set1,467Raw benchmark valueCI 1,453 - 1,481
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #33
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,464
- Percentile
- 90.1%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: multi_turn. Source rank: #44. Votes: 1066. Organization: google. License: Apache 2.0.
90.1% percentile inside its fair comparison set1,464Raw benchmark valueCI 1,446 - 1,482
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #34
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,441
- Percentile
- 89.8%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: overall. Source rank: #43. Votes: 5884. Organization: google. License: Apache 2.0.
89.8% percentile inside its fair comparison set1,441Raw benchmark valueCI 1,434 - 1,449
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #33
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,417
- Percentile
- 90.1%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: creative_writing. Source rank: #43. Votes: 940. Organization: google. License: Apache 2.0.
90.1% percentile inside its fair comparison set1,417Raw benchmark valueCI 1,398 - 1,437
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #39
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,447
- Percentile
- 88.3%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: english. Source rank: #49. Votes: 2603. Organization: google. License: Apache 2.0.
88.3% percentile inside its fair comparison set1,447Raw benchmark valueCI 1,436 - 1,459
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #32
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,439
- Percentile
- 90.5%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: exclude_ties. Source rank: #41. Votes: 4043. Organization: google. License: Apache 2.0.
90.5% percentile inside its fair comparison set1,439Raw benchmark valueCI 1,428 - 1,450
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #39
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,445
- Percentile
- 88.3%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: hard_prompts. Source rank: #47. Votes: 3367. Organization: google. License: Apache 2.0.
88.3% percentile inside its fair comparison set1,445Raw benchmark valueCI 1,435 - 1,455
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #38
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,454
- Percentile
- 88.6%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: hard_prompts_english. Source rank: #45. Votes: 1533. Organization: google. License: Apache 2.0.
88.6% percentile inside its fair comparison set1,454Raw benchmark valueCI 1,439 - 1,468
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #32
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,431
- Percentile
- 90.5%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: instruction_following. Source rank: #40. Votes: 1661. Organization: google. License: Apache 2.0.
90.5% percentile inside its fair comparison set1,431Raw benchmark valueCI 1,417 - 1,445
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #33
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,443
- Percentile
- 89.5%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: longer_query. Source rank: #41. Votes: 1657. Organization: google. License: Apache 2.0.
89.5% percentile inside its fair comparison set1,443Raw benchmark valueCI 1,429 - 1,457
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #30
verified runtimeexact aliasBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,451
- Percentile
- 91%
- Last updated
- aging
- Eligibility
- benchmark_derived_model
Parsed from Arena leaderboard dataset row `gemma-4-31b`. Category: multi_turn. Source rank: #38. Votes: 1066. Organization: google. License: Apache 2.0.
91% percentile inside its fair comparison set1,451Raw benchmark valueCI 1,433 - 1,469