APEX-Agents-AA
AA · Professional reasoning · Objective
Long-horizon agentic task completion.
Rank #7 · Source label: GPT-5.6 Luna (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 35.8%
- Percentile
- 81.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `apexAgents`.
81.3% percentile inside its fair comparison set35.8%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #56 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,491
- Percentile
- 83.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: expert. Source rank: #58. Votes: 4355. Organization: openai. License: Proprietary.
83.2% percentile inside its fair comparison set1,491Raw benchmark valueCI 1,482 - 1,501
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #51 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,463
- Percentile
- 86.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_business_and_management_and_financial_operations. Source rank: #54. Votes: 7253. Organization: openai. License: Proprietary.
86.4% percentile inside its fair comparison set1,463Raw benchmark valueCI 1,455 - 1,471
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #74 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,418
- Percentile
- 80.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_entertainment_and_sports_and_media. Source rank: #82. Votes: 10020. Organization: openai. License: Proprietary.
80.5% percentile inside its fair comparison set1,418Raw benchmark valueCI 1,411 - 1,425
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #67 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,463
- Percentile
- 81.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_legal_and_government. Source rank: #74. Votes: 3241. Organization: openai. License: Proprietary.
81.1% percentile inside its fair comparison set1,463Raw benchmark valueCI 1,452 - 1,474
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #81 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,462
- Percentile
- 78.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_life_and_physical_and_social_science. Source rank: #89. Votes: 6193. Organization: openai. License: Proprietary.
78.6% percentile inside its fair comparison set1,462Raw benchmark valueCI 1,453 - 1,470
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #38 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,487
- Percentile
- 89.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_mathematical. Source rank: #39. Votes: 2142. Organization: openai. License: Proprietary.
89.6% percentile inside its fair comparison set1,487Raw benchmark valueCI 1,473 - 1,500
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #118 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,439
- Percentile
- 66.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_medicine_and_healthcare. Source rank: #132. Votes: 2828. Organization: openai. License: Proprietary.
66.2% percentile inside its fair comparison set1,439Raw benchmark valueCI 1,427 - 1,451
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #61 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,493
- Percentile
- 84%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_software_and_it_services. Source rank: #66. Votes: 14928. Organization: openai. License: Proprietary.
84% percentile inside its fair comparison set1,493Raw benchmark valueCI 1,487 - 1,499
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #69 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,433
- Percentile
- 81.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_writing_and_literature_and_language. Source rank: #76. Votes: 10472. Organization: openai. License: Proprietary.
81.9% percentile inside its fair comparison set1,433Raw benchmark valueCI 1,426 - 1,440
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #47 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,479
- Percentile
- 85.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: expert. Source rank: #50. Votes: 4355. Organization: openai. License: Proprietary.
85.9% percentile inside its fair comparison set1,479Raw benchmark valueCI 1,469 - 1,488
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #50 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,440
- Percentile
- 86.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_business_and_management_and_financial_operations. Source rank: #53. Votes: 7253. Organization: openai. License: Proprietary.
86.7% percentile inside its fair comparison set1,440Raw benchmark valueCI 1,432 - 1,448
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #83 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,402
- Percentile
- 78.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_entertainment_and_sports_and_media. Source rank: #90. Votes: 10020. Organization: openai. License: Proprietary.
78.1% percentile inside its fair comparison set1,402Raw benchmark valueCI 1,395 - 1,409
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #68 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,447
- Percentile
- 80.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_legal_and_government. Source rank: #72. Votes: 3241. Organization: openai. License: Proprietary.
80.8% percentile inside its fair comparison set1,447Raw benchmark valueCI 1,436 - 1,458
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #107 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,434
- Percentile
- 71.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_life_and_physical_and_social_science. Source rank: #115. Votes: 6193. Organization: openai. License: Proprietary.
71.7% percentile inside its fair comparison set1,434Raw benchmark valueCI 1,426 - 1,442
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #45 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,472
- Percentile
- 87.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_mathematical. Source rank: #47. Votes: 2142. Organization: openai. License: Proprietary.
87.6% percentile inside its fair comparison set1,472Raw benchmark valueCI 1,458 - 1,485
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #142 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,398
- Percentile
- 59.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_medicine_and_healthcare. Source rank: #158. Votes: 2828. Organization: openai. License: Proprietary.
59.2% percentile inside its fair comparison set1,398Raw benchmark valueCI 1,386 - 1,410
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #66 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,464
- Percentile
- 82.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_software_and_it_services. Source rank: #69. Votes: 14928. Organization: openai. License: Proprietary.
82.7% percentile inside its fair comparison set1,464Raw benchmark valueCI 1,458 - 1,470
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #71 · Source label: gpt-5.6-luna-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,419
- Percentile
- 81.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-luna-xhigh`. Category: industry_writing_and_literature_and_language. Source rank: #77. Votes: 10472. Organization: openai. License: Proprietary.
81.3% percentile inside its fair comparison set1,419Raw benchmark valueCI 1,412 - 1,425
Vals Index
VALS-AI · Professional reasoning · Combined
Weighted model performance across economically relevant Vals tasks.
Rank #21 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 52
- Percentile
- 50%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: vals_index; provider: OpenAI.
50% percentile inside its fair comparison set52Raw benchmark valueCI 50 - 54
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #33 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 36.5%
- Percentile
- 52.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: OpenAI.
52.9% percentile inside its fair comparison set36.5%Raw benchmark valueCI 30% - 43.1%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #49 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 1.3%
- Percentile
- 31.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
31.9% percentile inside its fair comparison set1.3%Raw benchmark valueCI 0% - 2.9%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #42 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 84%
- Percentile
- 69.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: OpenAI.
69.6% percentile inside its fair comparison set84%Raw benchmark valueCI 83.2% - 84.9%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #19 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 55%
- Percentile
- 73.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: OpenAI.
73.9% percentile inside its fair comparison set55%Raw benchmark valueCI 54.4% - 55.6%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #49 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 42.4%
- Percentile
- 49.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI.
49.5% percentile inside its fair comparison set42.4%Raw benchmark valueCI 37.9% - 46.8%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #33 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 84.4%
- Percentile
- 66.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI.
66.7% percentile inside its fair comparison set84.4%Raw benchmark valueCI 79.3% - 89.5%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #45 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 44.2%
- Percentile
- 45.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
45.7% percentile inside its fair comparison set44.2%Raw benchmark valueCI 37.7% - 50.7%
Public Benefits Bench
VALS-AI · Professional reasoning · Objective
Answering SNAP benefits questions across the public-benefits lifecycle.
Rank #27 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 61.2%
- Percentile
- 42.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: public-benefits-bench; provider: OpenAI.
42.2% percentile inside its fair comparison set61.2%Raw benchmark valueCI 58.7% - 63.6%
SkillsBench
VALS-AI · Professional reasoning · Objective
Applied professional skills tasks.
Rank #6 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 60.4%
- Percentile
- 84.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: skillsbench; provider: OpenAI.
84.8% percentile inside its fair comparison set60.4%Raw benchmark valueCI 51.1% - 69.8%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #7 · Source label: openai/gpt-5.6-luna
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 76.2%
- Percentile
- 96.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: OpenAI.
96.1% percentile inside its fair comparison set76.2%Raw benchmark valueCI 74.5% - 77.8%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #28 · Source label: gpt-5.6-luna-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 78%
- Percentile
- 58.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
58.5% percentile inside its fair comparison set78%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #45 · Source label: gpt-5.6-luna-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 73.6%
- Percentile
- 32.3%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
32.3% percentile inside its fair comparison set73.6%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #27 · Source label: gpt-5.6-luna-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 86.5%
- Percentile
- 60%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
60% percentile inside its fair comparison set86.5%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #36 · Source label: gpt-5.6-luna-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 47.6%
- Percentile
- 46.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
46.2% percentile inside its fair comparison set47.6%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #11 · Source label: gpt-5.6-luna-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
100% percentile inside its fair comparison set100%Raw benchmark value