PRBench Legal
SL · Professional reasoning · Rubric
Applied legal reasoning on professional-domain tasks.
Rank #7 · Source label: gpt-5.6-sol (max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Scale Labs
- Raw value
- 50.5%
- Percentile
- 85.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5.6-sol (max). Collapse policy: highest reported score per canonical model.
85.4% percentile inside its fair comparison set50.5%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #8 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,535
- Percentile
- 97.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: expert. Source rank: #8. Votes: 4112. Organization: openai. License: Proprietary.
97.9% percentile inside its fair comparison set1,535Raw benchmark valueCI 1,525 - 1,544
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #18 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,485
- Percentile
- 95.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_business_and_management_and_financial_operations. Source rank: #19. Votes: 6900. Organization: openai. License: Proprietary.
95.4% percentile inside its fair comparison set1,485Raw benchmark valueCI 1,477 - 1,492
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #11 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,471
- Percentile
- 97.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_entertainment_and_sports_and_media. Source rank: #11. Votes: 9287. Organization: openai. License: Proprietary.
97.3% percentile inside its fair comparison set1,471Raw benchmark valueCI 1,464 - 1,478
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #20 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,494
- Percentile
- 94.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_legal_and_government. Source rank: #20. Votes: 3111. Organization: openai. License: Proprietary.
94.6% percentile inside its fair comparison set1,494Raw benchmark valueCI 1,483 - 1,505
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #38 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,489
- Percentile
- 90.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_life_and_physical_and_social_science. Source rank: #40. Votes: 5967. Organization: openai. License: Proprietary.
90.1% percentile inside its fair comparison set1,489Raw benchmark valueCI 1,481 - 1,498
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #10 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,513
- Percentile
- 97.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_mathematical. Source rank: #10. Votes: 2026. Organization: openai. License: Proprietary.
97.5% percentile inside its fair comparison set1,513Raw benchmark valueCI 1,499 - 1,527
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #54 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,480
- Percentile
- 84.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_medicine_and_healthcare. Source rank: #58. Votes: 2672. Organization: openai. License: Proprietary.
84.7% percentile inside its fair comparison set1,480Raw benchmark valueCI 1,468 - 1,492
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #19 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,522
- Percentile
- 95.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_software_and_it_services. Source rank: #19. Votes: 14357. Organization: openai. License: Proprietary.
95.2% percentile inside its fair comparison set1,522Raw benchmark valueCI 1,516 - 1,528
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #11 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,481
- Percentile
- 97.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_writing_and_literature_and_language. Source rank: #11. Votes: 9866. Organization: openai. License: Proprietary.
97.3% percentile inside its fair comparison set1,481Raw benchmark valueCI 1,474 - 1,489
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #16 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,518
- Percentile
- 95.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: expert. Source rank: #16. Votes: 4112. Organization: openai. License: Proprietary.
95.4% percentile inside its fair comparison set1,518Raw benchmark valueCI 1,508 - 1,528
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #37 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,455
- Percentile
- 90.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_business_and_management_and_financial_operations. Source rank: #40. Votes: 6900. Organization: openai. License: Proprietary.
90.2% percentile inside its fair comparison set1,455Raw benchmark valueCI 1,447 - 1,463
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #24 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,450
- Percentile
- 93.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_entertainment_and_sports_and_media. Source rank: #25. Votes: 9287. Organization: openai. License: Proprietary.
93.9% percentile inside its fair comparison set1,450Raw benchmark valueCI 1,442 - 1,457
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #37 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,472
- Percentile
- 89.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_legal_and_government. Source rank: #39. Votes: 3111. Organization: openai. License: Proprietary.
89.7% percentile inside its fair comparison set1,472Raw benchmark valueCI 1,461 - 1,483
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #67 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,454
- Percentile
- 82.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_life_and_physical_and_social_science. Source rank: #71. Votes: 5967. Organization: openai. License: Proprietary.
82.4% percentile inside its fair comparison set1,454Raw benchmark valueCI 1,446 - 1,463
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #19 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,496
- Percentile
- 94.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_mathematical. Source rank: #19. Votes: 2026. Organization: openai. License: Proprietary.
94.9% percentile inside its fair comparison set1,496Raw benchmark valueCI 1,482 - 1,509
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #99 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,432
- Percentile
- 71.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_medicine_and_healthcare. Source rank: #107. Votes: 2672. Organization: openai. License: Proprietary.
71.7% percentile inside its fair comparison set1,432Raw benchmark valueCI 1,420 - 1,444
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #32 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,488
- Percentile
- 91.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_software_and_it_services. Source rank: #32. Votes: 14357. Organization: openai. License: Proprietary.
91.8% percentile inside its fair comparison set1,488Raw benchmark valueCI 1,482 - 1,494
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #27 · Source label: gpt-5.6-sol-xhigh
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,463
- Percentile
- 93.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-5.6-sol-xhigh`. Category: industry_writing_and_literature_and_language. Source rank: #29. Votes: 9866. Organization: openai. License: Proprietary.
93.1% percentile inside its fair comparison set1,463Raw benchmark valueCI 1,456 - 1,470
Vals Index
VALS-AI · Professional reasoning · Combined
Weighted model performance across economically relevant Vals tasks.
Rank #10 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 58
- Percentile
- 77.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: vals_index; provider: OpenAI.
77.5% percentile inside its fair comparison set58Raw benchmark valueCI 56 - 60
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #9 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 48.1%
- Percentile
- 89.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: OpenAI.
89.7% percentile inside its fair comparison set48.1%Raw benchmark valueCI 41.3% - 54.9%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #40 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 2.5%
- Percentile
- 44.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
44.9% percentile inside its fair comparison set2.5%Raw benchmark valueCI 0.9% - 4.1%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #9 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 87%
- Percentile
- 94.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: OpenAI.
94.1% percentile inside its fair comparison set87%Raw benchmark valueCI 86.2% - 87.8%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #23 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 53.8%
- Percentile
- 68.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: OpenAI.
68.1% percentile inside its fair comparison set53.8%Raw benchmark valueCI 52.1% - 55.4%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #41 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 44%
- Percentile
- 57.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI.
57.9% percentile inside its fair comparison set44%Raw benchmark valueCI 39.5% - 48.4%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #28 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 85.2%
- Percentile
- 71.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI.
71.9% percentile inside its fair comparison set85.2%Raw benchmark valueCI 81.4% - 89.1%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #6 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 52.6%
- Percentile
- 93.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
93.8% percentile inside its fair comparison set52.6%Raw benchmark valueCI 45.9% - 59.3%
Public Benefits Bench
VALS-AI · Professional reasoning · Objective
Answering SNAP benefits questions across the public-benefits lifecycle.
Rank #16 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 66.5%
- Percentile
- 66.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: public-benefits-bench; provider: OpenAI.
66.7% percentile inside its fair comparison set66.5%Raw benchmark valueCI 64.1% - 68.9%
SkillsBench
VALS-AI · Professional reasoning · Objective
Applied professional skills tasks.
Rank #14 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 54.1%
- Percentile
- 60.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: skillsbench; provider: OpenAI.
60.6% percentile inside its fair comparison set54.1%Raw benchmark valueCI 44.8% - 63.4%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #30 · Source label: openai/gpt-5.6-sol
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 74.8%
- Percentile
- 77.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: OpenAI.
77.5% percentile inside its fair comparison set74.8%Raw benchmark valueCI 73.1% - 76.5%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #10 · Source label: gpt-5.6-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 79.8%
- Percentile
- 86.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
86.2% percentile inside its fair comparison set79.8%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #10 · Source label: gpt-5.6-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 81.1%
- Percentile
- 86.2%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
86.2% percentile inside its fair comparison set81.1%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #13 · Source label: gpt-5.6-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 90.2%
- Percentile
- 81.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
81.5% percentile inside its fair comparison set90.2%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #26 · Source label: gpt-5.6-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 49.3%
- Percentile
- 61.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
61.5% percentile inside its fair comparison set49.3%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #9 · Source label: gpt-5.6-sol-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
100% percentile inside its fair comparison set100%Raw benchmark value