PRBench Legal
SL · Professional reasoning · Rubric
Applied legal reasoning on professional-domain tasks.
Rank #15 · Source label: GPT 6 Astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Scale Labs
- Raw value
- 48.4%
- Percentile
- 65.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: GPT 6 Astra. Collapse policy: highest reported score per canonical model.
65.9% percentile inside its fair comparison set48.4%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #20 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,517
- Percentile
- 94.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: expert. Source rank: #20. Votes: 900. Organization: openai. License: Proprietary.
94.2% percentile inside its fair comparison set1,517Raw benchmark valueCI 1,497 - 1,537
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #33 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,473
- Percentile
- 91.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_business_and_management_and_financial_operations. Source rank: #35. Votes: 1640. Organization: openai. License: Proprietary.
91.3% percentile inside its fair comparison set1,473Raw benchmark valueCI 1,458 - 1,488
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #22 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,458
- Percentile
- 94.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_entertainment_and_sports_and_media. Source rank: #22. Votes: 2320. Organization: openai. License: Proprietary.
94.4% percentile inside its fair comparison set1,458Raw benchmark valueCI 1,445 - 1,471
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #29 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,484
- Percentile
- 92%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_legal_and_government. Source rank: #30. Votes: 785. Organization: openai. License: Proprietary.
92% percentile inside its fair comparison set1,484Raw benchmark valueCI 1,463 - 1,506
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #31 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,493
- Percentile
- 92%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_life_and_physical_and_social_science. Source rank: #32. Votes: 1538. Organization: openai. License: Proprietary.
92% percentile inside its fair comparison set1,493Raw benchmark valueCI 1,478 - 1,508
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #21 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,498
- Percentile
- 94.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_mathematical. Source rank: #22. Votes: 511. Organization: openai. License: Proprietary.
94.4% percentile inside its fair comparison set1,498Raw benchmark valueCI 1,472 - 1,525
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #75 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,467
- Percentile
- 78.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_medicine_and_healthcare. Source rank: #82. Votes: 672. Organization: openai. License: Proprietary.
78.6% percentile inside its fair comparison set1,467Raw benchmark valueCI 1,444 - 1,490
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #17 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,523
- Percentile
- 95.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_software_and_it_services. Source rank: #17. Votes: 3336. Organization: openai. License: Proprietary.
95.7% percentile inside its fair comparison set1,523Raw benchmark valueCI 1,513 - 1,533
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #29 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,465
- Percentile
- 92.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_writing_and_literature_and_language. Source rank: #31. Votes: 2559. Organization: openai. License: Proprietary.
92.5% percentile inside its fair comparison set1,465Raw benchmark valueCI 1,453 - 1,478
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #32 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,492
- Percentile
- 90.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: expert. Source rank: #33. Votes: 900. Organization: openai. License: Proprietary.
90.5% percentile inside its fair comparison set1,492Raw benchmark valueCI 1,473 - 1,512
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #56 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,437
- Percentile
- 85.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_business_and_management_and_financial_operations. Source rank: #59. Votes: 1640. Organization: openai. License: Proprietary.
85.1% percentile inside its fair comparison set1,437Raw benchmark valueCI 1,422 - 1,452
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #48 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,428
- Percentile
- 87.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_entertainment_and_sports_and_media. Source rank: #50. Votes: 2320. Organization: openai. License: Proprietary.
87.4% percentile inside its fair comparison set1,428Raw benchmark valueCI 1,415 - 1,441
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #56 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,454
- Percentile
- 84.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_legal_and_government. Source rank: #59. Votes: 785. Organization: openai. License: Proprietary.
84.2% percentile inside its fair comparison set1,454Raw benchmark valueCI 1,433 - 1,476
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #75 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,451
- Percentile
- 80.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_life_and_physical_and_social_science. Source rank: #80. Votes: 1538. Organization: openai. License: Proprietary.
80.2% percentile inside its fair comparison set1,451Raw benchmark valueCI 1,435 - 1,466
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #39 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,477
- Percentile
- 89.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_mathematical. Source rank: #41. Votes: 511. Organization: openai. License: Proprietary.
89.3% percentile inside its fair comparison set1,477Raw benchmark valueCI 1,451 - 1,503
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #119 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,418
- Percentile
- 65.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_medicine_and_healthcare. Source rank: #130. Votes: 672. Organization: openai. License: Proprietary.
65.9% percentile inside its fair comparison set1,418Raw benchmark valueCI 1,396 - 1,441
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #49 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,475
- Percentile
- 87.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_software_and_it_services. Source rank: #51. Votes: 3336. Organization: openai. License: Proprietary.
87.2% percentile inside its fair comparison set1,475Raw benchmark valueCI 1,465 - 1,485
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #50 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,439
- Percentile
- 86.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `gpt-6-astra-max`. Category: industry_writing_and_literature_and_language. Source rank: #53. Votes: 2559. Organization: openai. License: Proprietary.
86.9% percentile inside its fair comparison set1,439Raw benchmark valueCI 1,427 - 1,451
Vals Index
VALS-AI · Professional reasoning · Combined
Weighted model performance across economically relevant Vals tasks.
Rank #6 · Source label: openai/gpt-6-astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 63
- Percentile
- 87.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: vals_index; provider: OpenAI.
87.5% percentile inside its fair comparison set63Raw benchmark valueCI 61 - 65
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #25 · Source label: openai/gpt-6-astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 39.4%
- Percentile
- 64.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: OpenAI.
64.7% percentile inside its fair comparison set39.4%Raw benchmark valueCI 32.8% - 46.1%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #28 · Source label: openai/gpt-6-astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 5.4%
- Percentile
- 60.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: OpenAI.
60.9% percentile inside its fair comparison set5.4%Raw benchmark valueCI 3.1% - 7.7%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #25 · Source label: openai/gpt-6-astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 53.5%
- Percentile
- 65.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: OpenAI.
65.2% percentile inside its fair comparison set53.5%Raw benchmark valueCI 49.5% - 57.6%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #28 · Source label: openai/gpt-6-astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 48.5%
- Percentile
- 71.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: OpenAI.
71.6% percentile inside its fair comparison set48.5%Raw benchmark valueCI 44.3% - 52.7%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #14 · Source label: openai/gpt-6-astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 87.9%
- Percentile
- 86.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: OpenAI.
86.5% percentile inside its fair comparison set87.9%Raw benchmark valueCI 84.1% - 91.7%
SAGE
VALS-AI · Professional reasoning · Objective
Student Assessment with Generative Evaluation.
Rank #36 · Source label: openai/gpt-6-astra
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 46.4%
- Percentile
- 56.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: sage; provider: OpenAI.
56.8% percentile inside its fair comparison set46.4%Raw benchmark valueCI 39.6% - 53.1%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #1 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 83%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
100% percentile inside its fair comparison set83%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #4 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 82.2%
- Percentile
- 95.4%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
95.4% percentile inside its fair comparison set82.2%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #6 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 90.6%
- Percentile
- 93.8%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
93.8% percentile inside its fair comparison set90.6%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #1 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 58.4%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
100% percentile inside its fair comparison set58.4%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #19 · Source label: gpt-6-astra-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 100%
- Percentile
- 100%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
100% percentile inside its fair comparison set100%Raw benchmark value