Visible tradeoffsThis is a rubric-judged signal, so it is more structured than arena taste but still depends on the scoring rubric.
source
Scale Labs
metric
Score (%)
judge
Rubric
direction
higher better
group id
scale_tutorbench_current
domain
Reasoning / math / science
What it measures vs what it misses
✓ Measures
How well a model tutors through multi-step academic problems. Instruction quality, pedagogy, and reasoning support on teaching-style prompts.
✗ Misses
Live classroom preference. Latency and cost.
Why this countsIt is one of the cleaner reads on deliberate reasoning strength rather than style or popularity.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesIt still misses product usability, latency, and whether the model stays correct in messy real workflows.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: Muse Spark. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-5.4-pro-2026-03-05. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gemini-2.5-pro-preview-06-05. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-5-2025-08-07. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-5-2025-08-07. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-5-2025-08-07. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-nano-to-gpt-5.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: o3-pro-2025-06-10. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: kimi-k2.5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-5.1-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-opus-4-6-thinking-max. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gemini-3-pro-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-5.2-2025-12-11. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gemini-3.1-pro-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: o3-2025-04-16-medium. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: o3-2025-04-16-high. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gemini-3.1-flash-lite-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-opus-4-5-20251101-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-opus-4-1-20250805-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-opus-4-5-20251101. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-4-opus-20250514-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-5.1-instant. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-sonnet-4-5-20250929-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-opus-4-1-20250805_anthropic. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-37-sonnet-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-sonnet-4-5-20250929. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: claude-opus-4-20250514. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: llama4-maverick. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-tutorbench. Reported model configuration: gpt-4o. Collapse policy: highest reported score per canonical model.