Visible tradeoffsThis is a rubric-judged signal, so it is more structured than arena taste but still depends on the scoring rubric.
source
Scale Labs
metric
Score (%)
judge
Rubric
direction
higher better
group id
scale_prbench_legal_current
domain
Professional reasoning
What it measures vs what it misses
✓ Measures
Applied legal reasoning on professional-domain tasks.
✗ Misses
Broad chat preference. Search freshness or real-time retrieval quality.
Why this countsApplied legal reasoning on professional-domain tasks.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesBroad chat preference.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: Muse Spark 1.3. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: Muse Spark 1.1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-fable 5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: Muse Spark. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-opus-4-6 (Non-Thinking). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: Fable 5.1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5.6-sol (max). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5-pro. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5-pro. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5-pro. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-nano-to-gpt-5.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: o3-pro. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5.1-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: Gemini 3.8 Flash. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: o3. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: GPT 6 Astra. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: Grok 4.7 (xHigh). Collapse policy: highest reported score per canonical model.
47.6%
#17 · Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-opus-5-5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5.2-pro-2025-12-11. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-5.4 (High). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-opus-4-5-20251101-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gemini-3.1-pro. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: kimi-k2.5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-6-luna. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gemini-2.5-pro. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gemini-2.5-flash. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: kimi-k2-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-sonnet-4-5-20250929. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gemini-3-pro-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-oss-120b. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: mistral-medium-latest. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-6-sol. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: qwen.qwen3-235b-a22b-2507-v1:0. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: o4-mini. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: deepseek-v3p1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: deepseek-r1-0528. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-4.1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: kimi-k2-instruct. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model. Backfilled from Claude Opus 4.1 via approved benchmark identity mapping map-claude-opus-4-to-4-1.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model. Backfilled from Claude Opus 4.1 via approved benchmark identity mapping map-claude-opus-4-7-to-4-1.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: gpt-4.1-mini. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-prbench-legal. Reported model configuration: llama4-maverick-instruct-basic. Collapse policy: highest reported score per canonical model.