Intelligence Index
AA · Chat / text · Combined
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #11 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 52
- Percentile
- 98.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `intelligenceIndex`.
98.2% percentile inside its fair comparison set52Raw benchmark value
AA-Omniscience accuracy
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #12 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 62.1%
- Percentile
- 97.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceAccuracy`.
97.6% percentile inside its fair comparison set62.1%Raw benchmark value
AA-Omniscience non-hallucination
AA · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #109 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 45.7%
- Percentile
- 76.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `omniscienceNonHallucination`.
76.2% percentile inside its fair comparison set45.7%Raw benchmark value
Input price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #287 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $2 /1M input tokens
- Percentile
- 26.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mInputTokens`.
26.8% percentile inside its fair comparison set$2 /1M input tokensRaw benchmark value
Output price
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #290 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- $10 /1M output tokens
- Percentile
- 23.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `price1mOutputTokens`.
23.7% percentile inside its fair comparison set$10 /1M output tokensRaw benchmark value
Output Speed
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #183 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 63.1 tokens/s
- Percentile
- 30.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianOutputTokensPerSecond`.
30.5% percentile inside its fair comparison set63.1 tokens/sRaw benchmark value
Time to first token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #262 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 291.08s
- Percentile
- 1.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstTokenSeconds`.
1.1% percentile inside its fair comparison set291.08sRaw benchmark value
Time to first answer token
AA · Chat / text · Speed / cost
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #260 · Source label: GPT-6.1 Sol (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 291.08s
- Percentile
- 1.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `medianTimeToFirstAnswerTokenSeconds`.
1.1% percentile inside its fair comparison set291.08sRaw benchmark value
Text Arena
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #21 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,483
- Percentile
- 94.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: overall. Source rank: #21. Votes: 3071. Organization: openai. License: Proprietary.
94.7% percentile inside its fair comparison set1,483Raw benchmark valueCI 1,473 - 1,494
Text Arena · Creative Writing
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #26 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,460
- Percentile
- 93.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: creative_writing. Source rank: #27. Votes: 686. Organization: openai. License: Proprietary.
93.3% percentile inside its fair comparison set1,460Raw benchmark valueCI 1,436 - 1,483
Text Arena · English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #13 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,496
- Percentile
- 96.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: english. Source rank: #13. Votes: 1149. Organization: openai. License: Proprietary.
96.8% percentile inside its fair comparison set1,496Raw benchmark valueCI 1,479 - 1,513
Text Arena · Exclude Ties
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #21 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,496
- Percentile
- 94.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: exclude_ties. Source rank: #21. Votes: 2193. Organization: openai. License: Proprietary.
94.7% percentile inside its fair comparison set1,496Raw benchmark valueCI 1,481 - 1,511
Text Arena · Hard Prompts
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #18 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,509
- Percentile
- 95.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: hard_prompts. Source rank: #18. Votes: 1766. Organization: openai. License: Proprietary.
95.5% percentile inside its fair comparison set1,509Raw benchmark valueCI 1,495 - 1,522
Text Arena · Hard Prompts English
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #8 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,519
- Percentile
- 98.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: hard_prompts_english. Source rank: #8. Votes: 668. Organization: openai. License: Proprietary.
98.1% percentile inside its fair comparison set1,519Raw benchmark valueCI 1,496 - 1,541
Text Arena · Instruction Following
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #13 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,490
- Percentile
- 96.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: instruction_following. Source rank: #13. Votes: 922. Organization: openai. License: Proprietary.
96.8% percentile inside its fair comparison set1,490Raw benchmark valueCI 1,471 - 1,509
Text Arena · Longer Query
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #19 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,496
- Percentile
- 94.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: longer_query. Source rank: #19. Votes: 1180. Organization: openai. License: Proprietary.
94.9% percentile inside its fair comparison set1,496Raw benchmark valueCI 1,479 - 1,513
Text Arena · Multi Turn
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #20 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,493
- Percentile
- 94.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: multi_turn. Source rank: #20. Votes: 383. Organization: openai. License: Proprietary.
94.9% percentile inside its fair comparison set1,493Raw benchmark valueCI 1,463 - 1,522
Text Arena · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #56 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,446
- Percentile
- 85.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: overall. Source rank: #60. Votes: 3071. Organization: openai. License: Proprietary.
85.4% percentile inside its fair comparison set1,446Raw benchmark valueCI 1,435 - 1,456
Text Arena · Creative Writing · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #55 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,429
- Percentile
- 85.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: creative_writing. Source rank: #60. Votes: 686. Organization: openai. License: Proprietary.
85.6% percentile inside its fair comparison set1,429Raw benchmark valueCI 1,406 - 1,452
Text Arena · English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #55 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,454
- Percentile
- 85.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: english. Source rank: #57. Votes: 1149. Organization: openai. License: Proprietary.
85.6% percentile inside its fair comparison set1,454Raw benchmark valueCI 1,437 - 1,471
Text Arena · Exclude Ties · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #58 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,442
- Percentile
- 84.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: exclude_ties. Source rank: #62. Votes: 2193. Organization: openai. License: Proprietary.
84.8% percentile inside its fair comparison set1,442Raw benchmark valueCI 1,427 - 1,457
Text Arena · Hard Prompts · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #46 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,466
- Percentile
- 88%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: hard_prompts. Source rank: #49. Votes: 1766. Organization: openai. License: Proprietary.
88% percentile inside its fair comparison set1,466Raw benchmark valueCI 1,453 - 1,480
Text Arena · Hard Prompts English · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #46 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,471
- Percentile
- 88%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: hard_prompts_english. Source rank: #48. Votes: 668. Organization: openai. License: Proprietary.
88% percentile inside its fair comparison set1,471Raw benchmark valueCI 1,449 - 1,493
Text Arena · Instruction Following · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #29 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,469
- Percentile
- 92.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: instruction_following. Source rank: #30. Votes: 922. Organization: openai. License: Proprietary.
92.6% percentile inside its fair comparison set1,469Raw benchmark valueCI 1,450 - 1,488
Text Arena · Longer Query · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #42 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,465
- Percentile
- 88.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: longer_query. Source rank: #44. Votes: 1180. Organization: openai. License: Proprietary.
88.5% percentile inside its fair comparison set1,465Raw benchmark valueCI 1,448 - 1,482
Text Arena · Multi Turn · No Style Control
AR · Chat / text · Human
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #56 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,453
- Percentile
- 85.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6.1-sol-max`. Category: multi_turn. Source rank: #60. Votes: 383. Organization: openai. License: Proprietary.
85.3% percentile inside its fair comparison set1,453Raw benchmark valueCI 1,424 - 1,483
Instruction following
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #12 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 74.2%
- Percentile
- 83.1%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: IF. Tasks scored: 4.
83.1% percentile inside its fair comparison set74.2%Raw benchmark value
Language
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 90.1%
- Percentile
- 98.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Language. Tasks scored: 3.
98.5% percentile inside its fair comparison set90.1%Raw benchmark value
Paraphrase
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #18 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 72.6%
- Percentile
- 73.8%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: paraphrase. Category: IF.
73.8% percentile inside its fair comparison set72.6%Raw benchmark value
Simplify
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #8 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 72%
- Percentile
- 89.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: simplify. Category: IF.
89.2% percentile inside its fair comparison set72%Raw benchmark value
Story generation
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #36 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 70.8%
- Percentile
- 46.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: story_generation. Category: IF.
46.2% percentile inside its fair comparison set70.8%Raw benchmark value
Summarize
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #8 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 81.3%
- Percentile
- 89.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: summarize. Category: IF.
89.2% percentile inside its fair comparison set81.3%Raw benchmark value
Connections
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #18 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: connections. Category: Language.
100% percentile inside its fair comparison set100%Raw benchmark value
Plot unscrambling
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #2 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 84.4%
- Percentile
- 98.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: plot_unscrambling. Category: Language.
98.5% percentile inside its fair comparison set84.4%Raw benchmark value
Typos
LB · Chat / text · Objective
It tests whether the model is actually useful in normal conversational turns, not just on narrow correctness tasks.
Rank #13 · Source label: gpt-6.1-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 86%
- Percentile
- 89.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: typos. Category: Language.
89.2% percentile inside its fair comparison set86%Raw benchmark value