APEX-Agents-AA
AA · Professional reasoning · Objective
Long-horizon agentic task completion.
Rank #18 · Source label: GPT-5.4 nano (Xhigh)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 24.9%
- Percentile
- 46.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `apexAgents`.
46.9% percentile inside its fair comparison set24.9%Raw benchmark value
PRBench Legal
SL · Professional reasoning · Rubric
Applied legal reasoning on professional-domain tasks.
Rank #10 · Source label: gpt-5-pro
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Scale Labs
- Raw value
- 49.9%
- Percentile
- 82.9%
- Last updated
- recent
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5-pro. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-nano-to-gpt-5.
82.9% percentile inside its fair comparison set49.9%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #121 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,437
- Percentile
- 63.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: expert. Source rank: #134. Votes: 6344. Organization: openai. License: Proprietary.
63.3% percentile inside its fair comparison set1,437Raw benchmark valueCI 1,429 - 1,446
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #137 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,405
- Percentile
- 63.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #151. Votes: 12559. Organization: openai. License: Proprietary.
63.1% percentile inside its fair comparison set1,405Raw benchmark valueCI 1,399 - 1,412
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #153 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,352
- Percentile
- 59.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #172. Votes: 13657. Organization: openai. License: Proprietary.
59.4% percentile inside its fair comparison set1,352Raw benchmark valueCI 1,345 - 1,358
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #150 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,400
- Percentile
- 57.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_legal_and_government. Source rank: #168. Votes: 4954. Organization: openai. License: Proprietary.
57.3% percentile inside its fair comparison set1,400Raw benchmark valueCI 1,391 - 1,409
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #145 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,415
- Percentile
- 61.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_life_and_physical_and_social_science. Source rank: #162. Votes: 10492. Organization: openai. License: Proprietary.
61.5% percentile inside its fair comparison set1,415Raw benchmark valueCI 1,408 - 1,422
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #109 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,431
- Percentile
- 69.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_mathematical. Source rank: #119. Votes: 3478. Organization: openai. License: Proprietary.
69.7% percentile inside its fair comparison set1,431Raw benchmark valueCI 1,420 - 1,442
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #150 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,413
- Percentile
- 56.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_medicine_and_healthcare. Source rank: #168. Votes: 4788. Organization: openai. License: Proprietary.
56.9% percentile inside its fair comparison set1,413Raw benchmark valueCI 1,403 - 1,423
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #125 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,450
- Percentile
- 67%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_software_and_it_services. Source rank: #139. Votes: 24913. Organization: openai. License: Proprietary.
67% percentile inside its fair comparison set1,450Raw benchmark valueCI 1,445 - 1,455
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #159 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,363
- Percentile
- 57.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_writing_and_literature_and_language. Source rank: #179. Votes: 15381. Organization: openai. License: Proprietary.
57.9% percentile inside its fair comparison set1,363Raw benchmark valueCI 1,357 - 1,369
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #142 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,396
- Percentile
- 56.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: expert. Source rank: #159. Votes: 6344. Organization: openai. License: Proprietary.
56.9% percentile inside its fair comparison set1,396Raw benchmark valueCI 1,388 - 1,405
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #158 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,366
- Percentile
- 57.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #175. Votes: 12559. Organization: openai. License: Proprietary.
57.5% percentile inside its fair comparison set1,366Raw benchmark valueCI 1,360 - 1,373
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #172 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,324
- Percentile
- 54.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #193. Votes: 13657. Organization: openai. License: Proprietary.
54.3% percentile inside its fair comparison set1,324Raw benchmark valueCI 1,317 - 1,331
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #167 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,367
- Percentile
- 52.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_legal_and_government. Source rank: #187. Votes: 4954. Organization: openai. License: Proprietary.
52.4% percentile inside its fair comparison set1,367Raw benchmark valueCI 1,358 - 1,377
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #170 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,374
- Percentile
- 54.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_life_and_physical_and_social_science. Source rank: #187. Votes: 10492. Organization: openai. License: Proprietary.
54.8% percentile inside its fair comparison set1,374Raw benchmark valueCI 1,367 - 1,381
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #138 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,408
- Percentile
- 61.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_mathematical. Source rank: #150. Votes: 3478. Organization: openai. License: Proprietary.
61.5% percentile inside its fair comparison set1,408Raw benchmark valueCI 1,398 - 1,419
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #164 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,369
- Percentile
- 52.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_medicine_and_healthcare. Source rank: #183. Votes: 4788. Organization: openai. License: Proprietary.
52.9% percentile inside its fair comparison set1,369Raw benchmark valueCI 1,359 - 1,378
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #154 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,402
- Percentile
- 59.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_software_and_it_services. Source rank: #171. Votes: 24913. Organization: openai. License: Proprietary.
59.3% percentile inside its fair comparison set1,402Raw benchmark valueCI 1,397 - 1,408
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #170 · Source label: gpt-5.4-nano-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,338
- Percentile
- 54.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-nano-high`. Category: industry_writing_and_literature_and_language. Source rank: #192. Votes: 15381. Organization: openai. License: Proprietary.
54.9% percentile inside its fair comparison set1,338Raw benchmark valueCI 1,332 - 1,344
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #64 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 6.3%
- Percentile
- 7.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: OpenAI.
7.4% percentile inside its fair comparison set6.3%Raw benchmark valueCI 3% - 9.5%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #60 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 0%
- Percentile
- 20.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
20.3% percentile inside its fair comparison set0%Raw benchmark valueCI 0% - 0%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #97 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 77.9%
- Percentile
- 28.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: OpenAI.
28.9% percentile inside its fair comparison set77.9%Raw benchmark valueCI 77.1% - 78.8%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #56 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 38.2%
- Percentile
- 20.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: OpenAI.
20.3% percentile inside its fair comparison set38.2%Raw benchmark valueCI 35.9% - 40.5%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #58 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 41%
- Percentile
- 40%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI.
40% percentile inside its fair comparison set41%Raw benchmark valueCI 36.6% - 45.5%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #64 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 77.1%
- Percentile
- 34.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI.
34.4% percentile inside its fair comparison set77.1%Raw benchmark valueCI 73.4% - 80.8%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #60 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 38.1%
- Percentile
- 27.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
27.2% percentile inside its fair comparison set38.1%Raw benchmark valueCI 32% - 44.1%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #100 · Source label: openai/gpt-5.4-nano-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 67.4%
- Percentile
- 23.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: OpenAI.
23.3% percentile inside its fair comparison set67.4%Raw benchmark valueCI 65.6% - 69.2%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #56 · Source label: gpt-5.4-nano-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 67.6%
- Percentile
- 15.4%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
15.4% percentile inside its fair comparison set67.6%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #55 · Source label: gpt-5.4-nano-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 69.6%
- Percentile
- 16.9%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
16.9% percentile inside its fair comparison set69.6%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #56 · Source label: gpt-5.4-nano-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 54.3%
- Percentile
- 15.4%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
15.4% percentile inside its fair comparison set54.3%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #18 · Source label: gpt-5.4-nano-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 50.6%
- Percentile
- 73.8%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
73.8% percentile inside its fair comparison set50.6%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #32 · Source label: gpt-5.4-nano-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 98%
- Percentile
- 61.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
61.5% percentile inside its fair comparison set98%Raw benchmark value
Hallucination
BB · Professional reasoning · Rubric
Accuracy and fabrication resistance under prompts that invite unsupported claims.
Rank #23 · Source label: openai/gpt-5.4-nano
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- BridgeBench
- Raw value
- 69.6%
- Percentile
- 35.3%
- Last updated
- archived
- Eligibility
- headline eligible
Parsed from the BridgeBench page for bridgebench-hallucination.
35.3% percentile inside its fair comparison set69.6%Raw benchmark value
Poker Agent
VALS-AI · Professional reasoning · Objective
Agent profit in poker-style strategic play.
Rank #5 · Source label: openai/gpt-5-2025-08-07
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 1,103.2 score
- Percentile
- 94.7%
- Last updated
- archived
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from Vals AI BenchmarkView overall scores. Vals slug: poker_agent; provider: unknown. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-nano-to-gpt-5.
94.7% percentile inside its fair comparison set1,103.2 scoreRaw benchmark valueCI 1,103.2 score - 1,103.2 score