Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #62 · Source label: GPT-5.5 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 36
- Percentile
- 87%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
87% percentile inside its fair comparison set36Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #24 · Source label: GPT-5.5 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 45.6%
- Percentile
- 93.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
93.7% percentile inside its fair comparison set45.6%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #295 · Source label: GPT-5.5 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 7.3%
- Percentile
- 19.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
19.7% percentile inside its fair comparison set7.3%Raw benchmark value
IFBench
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #127 · Source label: GPT-5.5 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 46.1%
- Percentile
- 63.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `ifbench`.
63.5% percentile inside its fair comparison set46.1%Raw benchmark value
Blended price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #289 · Source label: GPT-5.5 (low)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $11.3 /1M tokens
- Percentile
- 4.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mBlended0To3To1`.
4.7% percentile inside its fair comparison set$11.3 /1M tokensRaw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #281 · Source label: GPT-5.5 (low)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $5 /1M input tokens
- Percentile
- 8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
8% percentile inside its fair comparison set$5 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #289 · Source label: GPT-5.5 (low)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $30 /1M output tokens
- Percentile
- 4.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
4.7% percentile inside its fair comparison set$30 /1M output tokensRaw benchmark value
Output Speed
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #144 · Source label: GPT-5.5 (Non-reasoning)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 69.5 tokens/s
- Percentile
- 36.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianOutputTokensPerSecond`.
36.2% percentile inside its fair comparison set69.5 tokens/sRaw benchmark value
Time to first token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #214 · Source label: GPT-5.5 (xhigh)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 95.87s
- Percentile
- 4.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstTokenSeconds`.
4.9% percentile inside its fair comparison set95.87sRaw benchmark value
Time to first answer token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #212 · Source label: GPT-5.5 (xhigh)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 95.87s
- Percentile
- 5.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstAnswerTokenSeconds`.
5.8% percentile inside its fair comparison set95.87sRaw benchmark value
Instruction following
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #16 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 70.7%
- Percentile
- 65.1%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: IF. Tasks scored: 4.
65.1% percentile inside its fair comparison set70.7%Raw benchmark value
Language
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 87.4%
- Percentile
- 93%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Language. Tasks scored: 3.
93% percentile inside its fair comparison set87.4%Raw benchmark value
Paraphrase
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #12 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 71.7%
- Percentile
- 74.4%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: paraphrase. Category: IF.
74.4% percentile inside its fair comparison set71.7%Raw benchmark value
Simplify
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #16 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 66.3%
- Percentile
- 65.1%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: simplify. Category: IF.
65.1% percentile inside its fair comparison set66.3%Raw benchmark value
Story generation
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #19 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 72.6%
- Percentile
- 58.1%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: story_generation. Category: IF.
58.1% percentile inside its fair comparison set72.6%Raw benchmark value
Summarize
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #15 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 72.3%
- Percentile
- 67.4%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: summarize. Category: IF.
67.4% percentile inside its fair comparison set72.3%Raw benchmark value
Connections
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: connections. Category: Language.
100% percentile inside its fair comparison set100%Raw benchmark value
Plot unscrambling
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #5 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 74.1%
- Percentile
- 90.7%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: plot_unscrambling. Category: Language.
90.7% percentile inside its fair comparison set74.1%Raw benchmark value
Typos
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #4 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 88%
- Percentile
- 93%
- Last updated
- aging
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: typos. Category: Language.
93% percentile inside its fair comparison set88%Raw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #8 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,481
- Percentile
- 97.8%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: overall. Source rank: #10. Votes: 28268. Organization: openai. License: Proprietary.
97.8% percentile inside its fair comparison set1,481Raw benchmark valueCI 1,476 - 1,486
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #16 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,451
- Percentile
- 95.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: creative_writing. Source rank: #19. Votes: 4904. Organization: openai. License: Proprietary.
95.4% percentile inside its fair comparison set1,451Raw benchmark valueCI 1,441 - 1,460
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #12 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,483
- Percentile
- 96.6%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: english. Source rank: #14. Votes: 13577. Organization: openai. License: Proprietary.
96.6% percentile inside its fair comparison set1,483Raw benchmark valueCI 1,477 - 1,490
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #8 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,490
- Percentile
- 97.8%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: exclude_ties. Source rank: #10. Votes: 21714. Organization: openai. License: Proprietary.
97.8% percentile inside its fair comparison set1,490Raw benchmark valueCI 1,484 - 1,497
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #11 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,500
- Percentile
- 96.9%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: hard_prompts. Source rank: #13. Votes: 18193. Organization: openai. License: Proprietary.
96.9% percentile inside its fair comparison set1,500Raw benchmark valueCI 1,494 - 1,505
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #14 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,499
- Percentile
- 96%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: hard_prompts_english. Source rank: #16. Votes: 9194. Organization: openai. License: Proprietary.
96% percentile inside its fair comparison set1,499Raw benchmark valueCI 1,492 - 1,507
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #9 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,479
- Percentile
- 97.5%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: instruction_following. Source rank: #11. Votes: 9331. Organization: openai. License: Proprietary.
97.5% percentile inside its fair comparison set1,479Raw benchmark valueCI 1,471 - 1,486
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #12 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,486
- Percentile
- 96.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: longer_query. Source rank: #16. Votes: 12143. Organization: openai. License: Proprietary.
96.4% percentile inside its fair comparison set1,486Raw benchmark valueCI 1,479 - 1,493
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #16 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,483
- Percentile
- 95.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: multi_turn. Source rank: #19. Votes: 4810. Organization: openai. License: Proprietary.
95.4% percentile inside its fair comparison set1,483Raw benchmark valueCI 1,474 - 1,493
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #11 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,468
- Percentile
- 96.9%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: overall. Source rank: #14. Votes: 28268. Organization: openai. License: Proprietary.
96.9% percentile inside its fair comparison set1,468Raw benchmark valueCI 1,463 - 1,473
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #14 · Source label: gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,451
- Percentile
- 96%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5`. Category: creative_writing. Source rank: #16. Votes: 4977. Organization: openai. License: Proprietary.
96% percentile inside its fair comparison set1,451Raw benchmark valueCI 1,442 - 1,461
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #15 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,467
- Percentile
- 95.7%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: english. Source rank: #18. Votes: 13577. Organization: openai. License: Proprietary.
95.7% percentile inside its fair comparison set1,467Raw benchmark valueCI 1,461 - 1,473
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #13 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,471
- Percentile
- 96.3%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: exclude_ties. Source rank: #16. Votes: 21714. Organization: openai. License: Proprietary.
96.3% percentile inside its fair comparison set1,471Raw benchmark valueCI 1,465 - 1,477
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #7 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,485
- Percentile
- 98.2%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: hard_prompts. Source rank: #9. Votes: 18193. Organization: openai. License: Proprietary.
98.2% percentile inside its fair comparison set1,485Raw benchmark valueCI 1,480 - 1,491
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #11 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,482
- Percentile
- 96.9%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: hard_prompts_english. Source rank: #13. Votes: 9194. Organization: openai. License: Proprietary.
96.9% percentile inside its fair comparison set1,482Raw benchmark valueCI 1,474 - 1,489
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #8 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,477
- Percentile
- 97.8%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: instruction_following. Source rank: #10. Votes: 9331. Organization: openai. License: Proprietary.
97.8% percentile inside its fair comparison set1,477Raw benchmark valueCI 1,469 - 1,484
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #9 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,481
- Percentile
- 97.4%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: longer_query. Source rank: #12. Votes: 12143. Organization: openai. License: Proprietary.
97.4% percentile inside its fair comparison set1,481Raw benchmark valueCI 1,474 - 1,488
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #15 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,471
- Percentile
- 95.7%
- Last updated
- aging
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: multi_turn. Source rank: #18. Votes: 4810. Organization: openai. License: Proprietary.
95.7% percentile inside its fair comparison set1,471Raw benchmark valueCI 1,462 - 1,481