Visible tradeoffsThis is a rubric-judged signal, so it is more structured than arena taste but still depends on the scoring rubric.
source
Scale Labs
metric
Success rate (%)
judge
Rubric
direction
higher better
group id
scale_hil_current
domain
Coding
What it measures vs what it misses
✓ Measures
Recovery and execution quality on coding-heavy tasks with intervention paths.
✗ Misses
Pure chat fluency. Standalone latency metrics.
Why this countsIt tells you whether the model can generate, repair, and reason over code under evaluator pressure rather than marketing examples.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesIt does not fully capture repo-scale iteration, IDE ergonomics, or long debugging loops.
Leaderboard · this benchmark version
#1 · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Claude Fable 5.1. Collapse policy: highest reported score per canonical model.
61.5%
#2 · Claude Opus 5 (Adaptive Reasoning, Max Effort)
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Claude Opus 5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Claude Fable 5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: GLM 5.2. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Claude Opus 4.7. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Gemini 3.8 Flash. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: GPT-5.5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Claude Opus 4.6. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Claude Opus 4.8. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Gemini 3.1 Pro. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: GPT 5.6 Sol. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Gemini 3.5 Flash. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Grok-4.20. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Kimi-k2.6. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: GPT-5.4. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: Minimax-M2.5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-hil-bench. Reported model configuration: GPT-5.3-codex. Collapse policy: highest reported score per canonical model.