APEX-Agents-AA
AA · Professional reasoning · Objective
Long-horizon agentic task completion.
Rank #9 · Source label: GPT-5.4 (Xhigh)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 33.3%
- Percentile
- 75%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `apexAgents`.
75% percentile inside its fair comparison set33.3%Raw benchmark value
PRBench Legal
SL · Professional reasoning · Rubric
Applied legal reasoning on professional-domain tasks.
Rank #19 · Source label: gpt-5.4 (High)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Scale Labs
- Raw value
- 44.4%
- Percentile
- 56.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5.4 (High). Collapse policy: highest reported score per canonical model.
56.1% percentile inside its fair comparison set44.4%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #23 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,515
- Percentile
- 93.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: expert. Source rank: #23. Votes: 6333. Organization: openai. License: Proprietary.
93.3% percentile inside its fair comparison set1,515Raw benchmark valueCI 1,507 - 1,524
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #24 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,482
- Percentile
- 93.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #25. Votes: 13075. Organization: openai. License: Proprietary.
93.8% percentile inside its fair comparison set1,482Raw benchmark valueCI 1,475 - 1,488
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #47 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,441
- Percentile
- 87.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #50. Votes: 14124. Organization: openai. License: Proprietary.
87.7% percentile inside its fair comparison set1,441Raw benchmark valueCI 1,435 - 1,448
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #26 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,490
- Percentile
- 92.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_legal_and_government. Source rank: #26. Votes: 5228. Organization: openai. License: Proprietary.
92.8% percentile inside its fair comparison set1,490Raw benchmark valueCI 1,481 - 1,499
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #42 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,488
- Percentile
- 89%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_life_and_physical_and_social_science. Source rank: #44. Votes: 10640. Organization: openai. License: Proprietary.
89% percentile inside its fair comparison set1,488Raw benchmark valueCI 1,481 - 1,495
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #19 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,499
- Percentile
- 94.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_mathematical. Source rank: #19. Votes: 3543. Organization: openai. License: Proprietary.
94.9% percentile inside its fair comparison set1,499Raw benchmark valueCI 1,488 - 1,510
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #60 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,475
- Percentile
- 82.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_medicine_and_healthcare. Source rank: #66. Votes: 4838. Organization: openai. License: Proprietary.
82.9% percentile inside its fair comparison set1,475Raw benchmark valueCI 1,465 - 1,485
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #41 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,507
- Percentile
- 89.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_software_and_it_services. Source rank: #41. Votes: 25424. Organization: openai. License: Proprietary.
89.4% percentile inside its fair comparison set1,507Raw benchmark valueCI 1,502 - 1,512
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #30 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,465
- Percentile
- 92.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_writing_and_literature_and_language. Source rank: #32. Votes: 16324. Organization: openai. License: Proprietary.
92.3% percentile inside its fair comparison set1,465Raw benchmark valueCI 1,459 - 1,471
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #22 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,507
- Percentile
- 93.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: expert. Source rank: #22. Votes: 6333. Organization: openai. License: Proprietary.
93.6% percentile inside its fair comparison set1,507Raw benchmark valueCI 1,499 - 1,515
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #17 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,475
- Percentile
- 95.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #17. Votes: 13075. Organization: openai. License: Proprietary.
95.7% percentile inside its fair comparison set1,475Raw benchmark valueCI 1,468 - 1,481
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #38 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,438
- Percentile
- 90.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #40. Votes: 14124. Organization: openai. License: Proprietary.
90.1% percentile inside its fair comparison set1,438Raw benchmark valueCI 1,432 - 1,444
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #21 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,486
- Percentile
- 94.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_legal_and_government. Source rank: #22. Votes: 5228. Organization: openai. License: Proprietary.
94.3% percentile inside its fair comparison set1,486Raw benchmark valueCI 1,477 - 1,495
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #32 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,478
- Percentile
- 91.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_life_and_physical_and_social_science. Source rank: #34. Votes: 10640. Organization: openai. License: Proprietary.
91.7% percentile inside its fair comparison set1,478Raw benchmark valueCI 1,472 - 1,485
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #16 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,497
- Percentile
- 95.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_mathematical. Source rank: #16. Votes: 3543. Organization: openai. License: Proprietary.
95.8% percentile inside its fair comparison set1,497Raw benchmark valueCI 1,486 - 1,507
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #52 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,459
- Percentile
- 85.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_medicine_and_healthcare. Source rank: #53. Votes: 4838. Organization: openai. License: Proprietary.
85.3% percentile inside its fair comparison set1,459Raw benchmark valueCI 1,449 - 1,468
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #30 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,491
- Percentile
- 92.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_software_and_it_services. Source rank: #30. Votes: 25424. Organization: openai. License: Proprietary.
92.3% percentile inside its fair comparison set1,491Raw benchmark valueCI 1,486 - 1,496
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #31 · Source label: gpt-5.4-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,458
- Percentile
- 92%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-high`. Category: industry_writing_and_literature_and_language. Source rank: #33. Votes: 16324. Organization: openai. License: Proprietary.
92% percentile inside its fair comparison set1,458Raw benchmark valueCI 1,453 - 1,464
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #68 · Source label: openai/gpt-5.4-2026-03-05
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 0%
- Percentile
- 20.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
20.3% percentile inside its fair comparison set0%Raw benchmark valueCI 0% - 0%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #15 · Source label: openai/gpt-5.4-2026-03-05
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 86%
- Percentile
- 89.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: OpenAI.
89.6% percentile inside its fair comparison set86%Raw benchmark valueCI 85.2% - 86.9%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #54 · Source label: openai/gpt-5.4-2026-03-05
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 41.3%
- Percentile
- 44.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI.
44.2% percentile inside its fair comparison set41.3%Raw benchmark valueCI 37.1% - 45.5%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #62 · Source label: openai/gpt-5.4-2026-03-05
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 77.5%
- Percentile
- 36.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI.
36.5% percentile inside its fair comparison set77.5%Raw benchmark valueCI 71% - 84%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #48 · Source label: openai/gpt-5.4-2026-03-05
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 43.3%
- Percentile
- 42%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
42% percentile inside its fair comparison set43.3%Raw benchmark valueCI 37.2% - 49.4%
SkillsBench
VALS-AI · Professional reasoning · Objective
Applied professional skills tasks.
Rank #18 · Source label: openai/gpt-5.4-2026-03-05
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 51.7%
- Percentile
- 48.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: skillsbench; provider: OpenAI.
48.5% percentile inside its fair comparison set51.7%Raw benchmark valueCI 42.5% - 61%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #44 · Source label: openai/gpt-5.4-2026-03-05
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 74%
- Percentile
- 67.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: OpenAI.
67.4% percentile inside its fair comparison set74%Raw benchmark valueCI 72.3% - 75.7%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #15 · Source label: gpt-5.4-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 79.3%
- Percentile
- 78.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
78.5% percentile inside its fair comparison set79.3%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #19 · Source label: gpt-5.4-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 78%
- Percentile
- 72.3%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
72.3% percentile inside its fair comparison set78%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #28 · Source label: gpt-5.4-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 86.2%
- Percentile
- 58.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
58.5% percentile inside its fair comparison set86.2%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #12 · Source label: gpt-5.4-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 51.8%
- Percentile
- 83.1%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
83.1% percentile inside its fair comparison set51.8%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #5 · Source label: gpt-5.4-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
100% percentile inside its fair comparison set100%Raw benchmark value
Hallucination
BB · Professional reasoning · Rubric
Accuracy and fabrication resistance under prompts that invite unsupported claims.
Rank #15 · Source label: openai/gpt-5.4
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- BridgeBench
- Raw value
- 72.8%
- Percentile
- 58.8%
- Last updated
- archived
- Eligibility
- headline eligible
Parsed from the BridgeBench page for bridgebench-hallucination.
58.8% percentile inside its fair comparison set72.8%Raw benchmark value
BS pushback
BB · Professional reasoning · Rubric
Resistance to confidently accepting bogus assumptions in expert-style prompts.
Rank #4 · Source label: openai/gpt-5.4
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- BridgeBench
- Raw value
- 91.5%
- Percentile
- 88.9%
- Last updated
- archived
- Eligibility
- headline eligible
Parsed from the BridgeBench page for bridgebench-pushback.
88.9% percentile inside its fair comparison set91.5%Raw benchmark value
Poker Agent
VALS-AI · Professional reasoning · Objective
Agent profit in poker-style strategic play.
Rank #3 · Source label: openai/gpt-5-2025-08-07
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 1,103.2 score
- Percentile
- 94.7%
- Last updated
- archived
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from Vals AI BenchmarkView overall scores. Vals slug: poker_agent; provider: unknown. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-to-gpt-5.
94.7% percentile inside its fair comparison set1,103.2 scoreRaw benchmark valueCI 1,103.2 score - 1,103.2 score