APEX-Agents-AA
AA · Professional reasoning · Objective
Long-horizon agentic task completion.
Rank #6 · Source label: GPT-5.5 (Xhigh)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 37.7%
- Percentile
- 84.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `apexAgents`.
84.4% percentile inside its fair comparison set37.7%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #25 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,512
- Percentile
- 92.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: expert. Source rank: #25. Votes: 7374. Organization: openai. License: Proprietary.
92.7% percentile inside its fair comparison set1,512Raw benchmark valueCI 1,505 - 1,520
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #10 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,490
- Percentile
- 97.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #10. Votes: 13704. Organization: openai. License: Proprietary.
97.6% percentile inside its fair comparison set1,490Raw benchmark valueCI 1,483 - 1,496
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #34 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,448
- Percentile
- 91.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #35. Votes: 16671. Organization: openai. License: Proprietary.
91.2% percentile inside its fair comparison set1,448Raw benchmark valueCI 1,442 - 1,454
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #19 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,495
- Percentile
- 94.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_legal_and_government. Source rank: #19. Votes: 5607. Organization: openai. License: Proprietary.
94.8% percentile inside its fair comparison set1,495Raw benchmark valueCI 1,486 - 1,504
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #25 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,498
- Percentile
- 93.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_life_and_physical_and_social_science. Source rank: #25. Votes: 11597. Organization: openai. License: Proprietary.
93.6% percentile inside its fair comparison set1,498Raw benchmark valueCI 1,491 - 1,505
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #14 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,506
- Percentile
- 96.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_mathematical. Source rank: #14. Votes: 3985. Organization: openai. License: Proprietary.
96.3% percentile inside its fair comparison set1,506Raw benchmark valueCI 1,496 - 1,516
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #40 · Source label: gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,486
- Percentile
- 88.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5`. Category: industry_medicine_and_healthcare. Source rank: #41. Votes: 5328. Organization: openai. License: Proprietary.
88.7% percentile inside its fair comparison set1,486Raw benchmark valueCI 1,477 - 1,495
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #33 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,510
- Percentile
- 91.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_software_and_it_services. Source rank: #33. Votes: 27383. Organization: openai. License: Proprietary.
91.5% percentile inside its fair comparison set1,510Raw benchmark valueCI 1,505 - 1,515
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #24 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,469
- Percentile
- 93.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_writing_and_literature_and_language. Source rank: #25. Votes: 18108. Organization: openai. License: Proprietary.
93.9% percentile inside its fair comparison set1,469Raw benchmark valueCI 1,463 - 1,475
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #21 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,507
- Percentile
- 93.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: expert. Source rank: #21. Votes: 7374. Organization: openai. License: Proprietary.
93.9% percentile inside its fair comparison set1,507Raw benchmark valueCI 1,500 - 1,515
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #16 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,478
- Percentile
- 95.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #16. Votes: 13704. Organization: openai. License: Proprietary.
95.9% percentile inside its fair comparison set1,478Raw benchmark valueCI 1,472 - 1,484
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #28 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,448
- Percentile
- 92.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #29. Votes: 16671. Organization: openai. License: Proprietary.
92.8% percentile inside its fair comparison set1,448Raw benchmark valueCI 1,442 - 1,454
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #17 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,491
- Percentile
- 95.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_legal_and_government. Source rank: #17. Votes: 5607. Organization: openai. License: Proprietary.
95.4% percentile inside its fair comparison set1,491Raw benchmark valueCI 1,482 - 1,500
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #26 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,486
- Percentile
- 93.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_life_and_physical_and_social_science. Source rank: #27. Votes: 11597. Organization: openai. License: Proprietary.
93.3% percentile inside its fair comparison set1,486Raw benchmark valueCI 1,479 - 1,492
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #21 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,495
- Percentile
- 94.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_mathematical. Source rank: #21. Votes: 3985. Organization: openai. License: Proprietary.
94.4% percentile inside its fair comparison set1,495Raw benchmark valueCI 1,485 - 1,505
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #44 · Source label: gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,465
- Percentile
- 87.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5`. Category: industry_medicine_and_healthcare. Source rank: #45. Votes: 5328. Organization: openai. License: Proprietary.
87.6% percentile inside its fair comparison set1,465Raw benchmark valueCI 1,456 - 1,474
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #27 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,492
- Percentile
- 93.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_software_and_it_services. Source rank: #27. Votes: 27383. Organization: openai. License: Proprietary.
93.1% percentile inside its fair comparison set1,492Raw benchmark valueCI 1,487 - 1,497
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #23 · Source label: gpt-5.5-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,466
- Percentile
- 94.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.5-high`. Category: industry_writing_and_literature_and_language. Source rank: #24. Votes: 18108. Organization: openai. License: Proprietary.
94.1% percentile inside its fair comparison set1,466Raw benchmark valueCI 1,460 - 1,472
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #24 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 40.4%
- Percentile
- 67.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: OpenAI.
67.6% percentile inside its fair comparison set40.4%Raw benchmark valueCI 33.7% - 47.1%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #34 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 3.8%
- Percentile
- 53.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
53.6% percentile inside its fair comparison set3.8%Raw benchmark valueCI 1.4% - 6.1%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #12 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 86.5%
- Percentile
- 91.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: OpenAI.
91.9% percentile inside its fair comparison set86.5%Raw benchmark valueCI 85.7% - 87.3%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #30 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 51.8%
- Percentile
- 58%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: OpenAI.
58% percentile inside its fair comparison set51.8%Raw benchmark valueCI 50.7% - 52.8%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #26 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 49.1%
- Percentile
- 73.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI.
73.7% percentile inside its fair comparison set49.1%Raw benchmark valueCI 44.8% - 53.4%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #18 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 86.9%
- Percentile
- 82.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI.
82.3% percentile inside its fair comparison set86.9%Raw benchmark valueCI 83.1% - 90.7%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #13 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 51.5%
- Percentile
- 85.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
85.2% percentile inside its fair comparison set51.5%Raw benchmark valueCI 43.8% - 59.3%
Public Benefits Bench
VALS-AI · Professional reasoning · Objective
Answering SNAP benefits questions across the public-benefits lifecycle.
Rank #28 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 60.9%
- Percentile
- 40%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: public-benefits-bench; provider: OpenAI.
40% percentile inside its fair comparison set60.9%Raw benchmark valueCI 58.4% - 63.4%
SkillsBench
VALS-AI · Professional reasoning · Objective
Applied professional skills tasks.
Rank #4 · Source label: openai/gpt-5.5-codex
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 62.6%
- Percentile
- 90.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: skillsbench; provider: OpenAI.
90.9% percentile inside its fair comparison set62.6%Raw benchmark valueCI 53.9% - 71.2%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #25 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 75%
- Percentile
- 81.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: OpenAI.
81.4% percentile inside its fair comparison set75%Raw benchmark valueCI 73.3% - 76.7%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #4 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 81.6%
- Percentile
- 95.4%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
95.4% percentile inside its fair comparison set81.6%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #11 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 80.2%
- Percentile
- 84.6%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
84.6% percentile inside its fair comparison set80.2%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #21 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 88.8%
- Percentile
- 69.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
69.2% percentile inside its fair comparison set88.8%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #5 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 55.9%
- Percentile
- 93.8%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
93.8% percentile inside its fair comparison set55.9%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #6 · Source label: gpt-5.5-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
100% percentile inside its fair comparison set100%Raw benchmark value
Public Benefits Bench v1
VALS-AI · Professional reasoning · Objective
Answering public-benefits questions across the benefits lifecycle.
Rank #8 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 57.2%
- Percentile
- 41.7%
- Last updated
- stale
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: public-benefits-bench-v1; provider: OpenAI.
41.7% percentile inside its fair comparison set57.2%Raw benchmark valueCI 54.7% - 59.8%
Hallucination
BB · Professional reasoning · Rubric
Accuracy and fabrication resistance under prompts that invite unsupported claims.
Rank #13 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- BridgeBench
- Raw value
- 73.6%
- Percentile
- 64.7%
- Last updated
- archived
- Eligibility
- headline eligible
Parsed from the BridgeBench page for bridgebench-hallucination.
64.7% percentile inside its fair comparison set73.6%Raw benchmark value
BS pushback
BB · Professional reasoning · Rubric
Resistance to confidently accepting bogus assumptions in expert-style prompts.
Rank #7 · Source label: openai/gpt-5.5
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- BridgeBench
- Raw value
- 88%
- Percentile
- 72.2%
- Last updated
- archived
- Eligibility
- headline eligible
Parsed from the BridgeBench page for bridgebench-pushback.
72.2% percentile inside its fair comparison set88%Raw benchmark value