APEX-Agents-AA
AA · Professional reasoning · Objective
Long-horizon agentic task completion.
Rank #14 · Source label: GPT-5.4 mini (Xhigh)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 28.2%
- Percentile
- 59.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `apexAgents`.
59.4% percentile inside its fair comparison set28.2%Raw benchmark value
PRBench Legal
SL · Professional reasoning · Rubric
Applied legal reasoning on professional-domain tasks.
Rank #9 · Source label: gpt-5-pro
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Scale Labs
- Raw value
- 49.9%
- Percentile
- 82.9%
- Last updated
- recent
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5-pro. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
82.9% percentile inside its fair comparison set49.9%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #69 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,480
- Percentile
- 79.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: expert. Source rank: #72. Votes: 6299. Organization: openai. License: Proprietary.
79.2% percentile inside its fair comparison set1,480Raw benchmark valueCI 1,472 - 1,489
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #57 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,460
- Percentile
- 84.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #61. Votes: 12827. Organization: openai. License: Proprietary.
84.8% percentile inside its fair comparison set1,460Raw benchmark valueCI 1,453 - 1,466
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #84 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,409
- Percentile
- 77.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #92. Votes: 13880. Organization: openai. License: Proprietary.
77.8% percentile inside its fair comparison set1,409Raw benchmark valueCI 1,402 - 1,415
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #72 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,457
- Percentile
- 79.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_legal_and_government. Source rank: #82. Votes: 5040. Organization: openai. License: Proprietary.
79.7% percentile inside its fair comparison set1,457Raw benchmark valueCI 1,448 - 1,466
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #77 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,465
- Percentile
- 79.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_life_and_physical_and_social_science. Source rank: #84. Votes: 10493. Organization: openai. License: Proprietary.
79.7% percentile inside its fair comparison set1,465Raw benchmark valueCI 1,458 - 1,472
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #75 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,452
- Percentile
- 79.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_mathematical. Source rank: #81. Votes: 3420. Organization: openai. License: Proprietary.
79.2% percentile inside its fair comparison set1,452Raw benchmark valueCI 1,441 - 1,463
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #95 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,451
- Percentile
- 72.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_medicine_and_healthcare. Source rank: #108. Votes: 4757. Organization: openai. License: Proprietary.
72.8% percentile inside its fair comparison set1,451Raw benchmark valueCI 1,441 - 1,460
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #77 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,485
- Percentile
- 79.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_software_and_it_services. Source rank: #84. Votes: 24993. Organization: openai. License: Proprietary.
79.8% percentile inside its fair comparison set1,485Raw benchmark valueCI 1,480 - 1,490
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #79 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,423
- Percentile
- 79.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_writing_and_literature_and_language. Source rank: #88. Votes: 15820. Organization: openai. License: Proprietary.
79.2% percentile inside its fair comparison set1,423Raw benchmark valueCI 1,417 - 1,429
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #99 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,435
- Percentile
- 70%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: expert. Source rank: #105. Votes: 6299. Organization: openai. License: Proprietary.
70% percentile inside its fair comparison set1,435Raw benchmark valueCI 1,426 - 1,443
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #96 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,418
- Percentile
- 74.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_business_and_management_and_financial_operations. Source rank: #103. Votes: 12827. Organization: openai. License: Proprietary.
74.3% percentile inside its fair comparison set1,418Raw benchmark valueCI 1,411 - 1,424
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #124 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,373
- Percentile
- 67.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_entertainment_and_sports_and_media. Source rank: #136. Votes: 13880. Organization: openai. License: Proprietary.
67.1% percentile inside its fair comparison set1,373Raw benchmark valueCI 1,366 - 1,379
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #128 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,415
- Percentile
- 63.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_legal_and_government. Source rank: #141. Votes: 5040. Organization: openai. License: Proprietary.
63.6% percentile inside its fair comparison set1,415Raw benchmark valueCI 1,406 - 1,424
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #131 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,417
- Percentile
- 65.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_life_and_physical_and_social_science. Source rank: #142. Votes: 10493. Organization: openai. License: Proprietary.
65.2% percentile inside its fair comparison set1,417Raw benchmark valueCI 1,410 - 1,424
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #110 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,428
- Percentile
- 69.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_mathematical. Source rank: #117. Votes: 3420. Organization: openai. License: Proprietary.
69.4% percentile inside its fair comparison set1,428Raw benchmark valueCI 1,417 - 1,439
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #143 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,395
- Percentile
- 59%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_medicine_and_healthcare. Source rank: #159. Votes: 4757. Organization: openai. License: Proprietary.
59% percentile inside its fair comparison set1,395Raw benchmark valueCI 1,385 - 1,404
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #115 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,436
- Percentile
- 69.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_software_and_it_services. Source rank: #125. Votes: 24993. Organization: openai. License: Proprietary.
69.7% percentile inside its fair comparison set1,436Raw benchmark valueCI 1,430 - 1,441
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #113 · Source label: gpt-5.4-mini-high
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,392
- Percentile
- 70.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.4-mini-high`. Category: industry_writing_and_literature_and_language. Source rank: #123. Votes: 15820. Organization: openai. License: Proprietary.
70.1% percentile inside its fair comparison set1,392Raw benchmark valueCI 1,386 - 1,398
Vals Index
VALS-AI · Professional reasoning · Combined
Weighted model performance across economically relevant Vals tasks.
Rank #38 · Source label: openai/gpt-5.4-mini-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 33
- Percentile
- 7.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: vals_index; provider: OpenAI.
7.5% percentile inside its fair comparison set33Raw benchmark valueCI 31 - 36
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #59 · Source label: openai/gpt-5.4-mini-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 12.5%
- Percentile
- 14.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: OpenAI.
14.7% percentile inside its fair comparison set12.5%Raw benchmark valueCI 8% - 17%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #67 · Source label: openai/gpt-5.4-mini-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 0%
- Percentile
- 20.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
20.3% percentile inside its fair comparison set0%Raw benchmark valueCI 0% - 0%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #17 · Source label: openai/gpt-5-2025-08-07
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 86%
- Percentile
- 88.9%
- Last updated
- recent
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: OpenAI. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
88.9% percentile inside its fair comparison set86%Raw benchmark valueCI 85.3% - 86.8%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #46 · Source label: openai/gpt-5.4-mini-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 45.4%
- Percentile
- 34.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: OpenAI.
34.8% percentile inside its fair comparison set45.4%Raw benchmark valueCI 44.5% - 46.2%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #20 · Source label: openai/gpt-5-2025-08-07
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 49.6%
- Percentile
- 81.1%
- Last updated
- recent
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
81.1% percentile inside its fair comparison set49.6%Raw benchmark valueCI 45.5% - 53.7%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #42 · Source label: openai/gpt-5-2025-08-07
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 83.7%
- Percentile
- 58.3%
- Last updated
- recent
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
58.3% percentile inside its fair comparison set83.7%Raw benchmark valueCI 79.9% - 87.4%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #15 · Source label: openai/gpt-5.4-mini-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 50.8%
- Percentile
- 82.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
82.7% percentile inside its fair comparison set50.8%Raw benchmark valueCI 44.1% - 57.5%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #78 · Source label: openai/gpt-5.4-mini-2026-03-17
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 71.2%
- Percentile
- 40.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: OpenAI.
40.3% percentile inside its fair comparison set71.2%Raw benchmark valueCI 69.5% - 73%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #49 · Source label: gpt-5.4-mini-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 70.8%
- Percentile
- 26.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
26.2% percentile inside its fair comparison set70.8%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #62 · Source label: gpt-5.4-mini-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 66.4%
- Percentile
- 6.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
6.2% percentile inside its fair comparison set66.4%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #53 · Source label: gpt-5.4-mini-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 63%
- Percentile
- 20%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
20% percentile inside its fair comparison set63%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #25 · Source label: gpt-5.4-mini-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 49.4%
- Percentile
- 63.1%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
63.1% percentile inside its fair comparison set49.4%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #4 · Source label: gpt-5.4-mini-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
100% percentile inside its fair comparison set100%Raw benchmark value
Hallucination
BB · Professional reasoning · Rubric
Accuracy and fabrication resistance under prompts that invite unsupported claims.
Rank #17 · Source label: openai/gpt-5.4-mini
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- BridgeBench
- Raw value
- 71.9%
- Percentile
- 52.9%
- Last updated
- archived
- Eligibility
- headline eligible
Parsed from the BridgeBench page for bridgebench-hallucination.
52.9% percentile inside its fair comparison set71.9%Raw benchmark value
BS pushback
BB · Professional reasoning · Rubric
Resistance to confidently accepting bogus assumptions in expert-style prompts.
Rank #9 · Source label: openai/gpt-5.4-mini
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- BridgeBench
- Raw value
- 78.5%
- Percentile
- 55.6%
- Last updated
- archived
- Eligibility
- headline eligible
Parsed from the BridgeBench page for bridgebench-pushback.
55.6% percentile inside its fair comparison set78.5%Raw benchmark value
Poker Agent
VALS-AI · Professional reasoning · Objective
Agent profit in poker-style strategic play.
Rank #4 · Source label: openai/gpt-5-2025-08-07
backfilledproxy backfilledBackground only
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 1,103.2 score
- Percentile
- 94.7%
- Last updated
- archived
- Eligibility
- Fallback benchmark identity is visible for context but excluded from default ranking.
Parsed from Vals AI BenchmarkView overall scores. Vals slug: poker_agent; provider: unknown. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
94.7% percentile inside its fair comparison set1,103.2 scoreRaw benchmark valueCI 1,103.2 score - 1,103.2 score