APEX-Agents-AA
AA · Professional reasoning · Objective
Long-horizon agentic task completion.
Rank #4 · Source label: GPT-5.6 Terra (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 38.9%
- Percentile
- 90.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `apexAgents`.
90.6% percentile inside its fair comparison set38.9%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #34 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,504
- Percentile
- 89.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: expert. Source rank: #35. Votes: 4391. Organization: openai. License: Proprietary.
89.9% percentile inside its fair comparison set1,504Raw benchmark valueCI 1,495 - 1,514
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #43 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,467
- Percentile
- 88.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_business_and_management_and_financial_operations. Source rank: #45. Votes: 7079. Organization: openai. License: Proprietary.
88.6% percentile inside its fair comparison set1,467Raw benchmark valueCI 1,459 - 1,475
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #61 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,432
- Percentile
- 84%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_entertainment_and_sports_and_media. Source rank: #68. Votes: 9653. Organization: openai. License: Proprietary.
84% percentile inside its fair comparison set1,432Raw benchmark valueCI 1,424 - 1,439
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #41 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,477
- Percentile
- 88.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_legal_and_government. Source rank: #43. Votes: 3229. Organization: openai. License: Proprietary.
88.5% percentile inside its fair comparison set1,477Raw benchmark valueCI 1,466 - 1,488
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #61 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,477
- Percentile
- 84%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_life_and_physical_and_social_science. Source rank: #67. Votes: 6139. Organization: openai. License: Proprietary.
84% percentile inside its fair comparison set1,477Raw benchmark valueCI 1,468 - 1,485
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #32 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,489
- Percentile
- 91.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_mathematical. Source rank: #33. Votes: 2121. Organization: openai. License: Proprietary.
91.3% percentile inside its fair comparison set1,489Raw benchmark valueCI 1,475 - 1,502
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #85 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,459
- Percentile
- 75.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_medicine_and_healthcare. Source rank: #94. Votes: 2718. Organization: openai. License: Proprietary.
75.7% percentile inside its fair comparison set1,459Raw benchmark valueCI 1,447 - 1,472
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #44 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,505
- Percentile
- 88.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_software_and_it_services. Source rank: #45. Votes: 14645. Organization: openai. License: Proprietary.
88.6% percentile inside its fair comparison set1,505Raw benchmark valueCI 1,499 - 1,511
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #59 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,445
- Percentile
- 84.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_writing_and_literature_and_language. Source rank: #64. Votes: 10426. Organization: openai. License: Proprietary.
84.5% percentile inside its fair comparison set1,445Raw benchmark valueCI 1,438 - 1,451
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #33 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,492
- Percentile
- 90.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: expert. Source rank: #34. Votes: 4391. Organization: openai. License: Proprietary.
90.2% percentile inside its fair comparison set1,492Raw benchmark valueCI 1,483 - 1,502
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #44 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,447
- Percentile
- 88.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_business_and_management_and_financial_operations. Source rank: #47. Votes: 7079. Organization: openai. License: Proprietary.
88.3% percentile inside its fair comparison set1,447Raw benchmark valueCI 1,440 - 1,455
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #62 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,416
- Percentile
- 83.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_entertainment_and_sports_and_media. Source rank: #68. Votes: 9653. Organization: openai. License: Proprietary.
83.7% percentile inside its fair comparison set1,416Raw benchmark valueCI 1,409 - 1,423
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #45 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,464
- Percentile
- 87.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_legal_and_government. Source rank: #47. Votes: 3229. Organization: openai. License: Proprietary.
87.4% percentile inside its fair comparison set1,464Raw benchmark valueCI 1,453 - 1,475
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #66 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,454
- Percentile
- 82.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_life_and_physical_and_social_science. Source rank: #70. Votes: 6139. Organization: openai. License: Proprietary.
82.6% percentile inside its fair comparison set1,454Raw benchmark valueCI 1,446 - 1,463
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #41 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,476
- Percentile
- 88.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_mathematical. Source rank: #43. Votes: 2121. Organization: openai. License: Proprietary.
88.8% percentile inside its fair comparison set1,476Raw benchmark valueCI 1,463 - 1,489
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #107 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,429
- Percentile
- 69.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_medicine_and_healthcare. Source rank: #115. Votes: 2718. Organization: openai. License: Proprietary.
69.4% percentile inside its fair comparison set1,429Raw benchmark valueCI 1,417 - 1,441
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #47 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,479
- Percentile
- 87.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_software_and_it_services. Source rank: #49. Votes: 14645. Organization: openai. License: Proprietary.
87.8% percentile inside its fair comparison set1,479Raw benchmark valueCI 1,473 - 1,485
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #61 · Source label: gpt-5.6-terra-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,429
- Percentile
- 84%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-terra-xhigh`. Category: industry_writing_and_literature_and_language. Source rank: #66. Votes: 10426. Organization: openai. License: Proprietary.
84% percentile inside its fair comparison set1,429Raw benchmark valueCI 1,422 - 1,436
Vals Index
VALS-AI · Professional reasoning · Combined
Weighted model performance across economically relevant Vals tasks.
Rank #18 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 53
- Percentile
- 57.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: vals_index; provider: OpenAI.
57.5% percentile inside its fair comparison set53Raw benchmark valueCI 51 - 56
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #21 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 41.3%
- Percentile
- 72.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: OpenAI.
72.1% percentile inside its fair comparison set41.3%Raw benchmark valueCI 34.6% - 48.1%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #52 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 0.8%
- Percentile
- 27.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
27.5% percentile inside its fair comparison set0.8%Raw benchmark valueCI 0% - 2.5%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #24 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 85.1%
- Percentile
- 83%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: OpenAI.
83% percentile inside its fair comparison set85.1%Raw benchmark valueCI 84.2% - 86%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #20 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 54.4%
- Percentile
- 72.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: OpenAI.
72.5% percentile inside its fair comparison set54.4%Raw benchmark valueCI 50.4% - 58.5%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #43 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 43.4%
- Percentile
- 55.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI.
55.8% percentile inside its fair comparison set43.4%Raw benchmark valueCI 39.2% - 47.7%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #48 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 82.9%
- Percentile
- 51%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI.
51% percentile inside its fair comparison set82.9%Raw benchmark valueCI 79.1% - 86.7%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #33 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 47%
- Percentile
- 60.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
60.5% percentile inside its fair comparison set47%Raw benchmark valueCI 40.3% - 53.7%
Public Benefits Bench
VALS-AI · Professional reasoning · Objective
Answering SNAP benefits questions across the public-benefits lifecycle.
Rank #24 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 62.4%
- Percentile
- 48.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: public-benefits-bench; provider: OpenAI.
48.9% percentile inside its fair comparison set62.4%Raw benchmark valueCI 59.9% - 64.9%
SkillsBench
VALS-AI · Professional reasoning · Objective
Applied professional skills tasks.
Rank #10 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 58.9%
- Percentile
- 72.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: skillsbench; provider: OpenAI.
72.7% percentile inside its fair comparison set58.9%Raw benchmark valueCI 50.1% - 67.7%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #6 · Source label: openai/gpt-5.6-terra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 76.2%
- Percentile
- 96.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: OpenAI.
96.1% percentile inside its fair comparison set76.2%Raw benchmark valueCI 74.5% - 77.8%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #16 · Source label: gpt-5.6-terra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 79.3%
- Percentile
- 76.9%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
76.9% percentile inside its fair comparison set79.3%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #21 · Source label: gpt-5.6-terra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 77.9%
- Percentile
- 69.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
69.2% percentile inside its fair comparison set77.9%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #25 · Source label: gpt-5.6-terra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 87.7%
- Percentile
- 63.1%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
63.1% percentile inside its fair comparison set87.7%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #21 · Source label: gpt-5.6-terra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 50.3%
- Percentile
- 69.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
69.2% percentile inside its fair comparison set50.3%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #10 · Source label: gpt-5.6-terra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
100% percentile inside its fair comparison set100%Raw benchmark value