Visible tradeoffsThis is a rubric-judged signal, so it is more structured than arena taste but still depends on the scoring rubric.
source
Scale Labs
metric
APR (%)
judge
Rubric
direction
higher better
group id
scale_vtb_current
domain
Vision understanding
What it measures vs what it misses
✓ Measures
Visual interpretation and reasoning over benchmark images and prompts.
✗ Misses
Image generation quality. Tool-use orchestration beyond the judged task.
Why this countsIt is useful when the model must read charts, UI, screenshots, or visual scenes rather than text alone.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesIt does not tell you whether the model can generate or edit images well.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: Muse Spark 1.1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gpt-5.4-2026-03-05 (reasoning effort = high). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gemini-3.1-pro-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-opus-4-6-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gemini-3-pro-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gpt-5-2025-08-07-thinking. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gpt-5-2025-08-07-thinking. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gpt-5-2025-08-07-thinking. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-nano-to-gpt-5.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: o3-2025-04-16. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gemini-2.5-pro-preview-06-05. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: o4-mini-2025-04-16. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-sonnet-4-5-20250929-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-sonnet-4-5-20250929. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gpt-4.1-2025-04-14. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-opus-4-1-20250805-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model. Backfilled from Claude Opus 4.1 via approved benchmark identity mapping map-claude-opus-4-to-4-1.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model. Backfilled from Claude Opus 4.1 via approved benchmark identity mapping map-claude-opus-4-7-to-4-1.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: gemini-2.5-flash. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-sonnet-4. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-sonnet-4. Collapse policy: highest reported score per canonical model. Backfilled from Claude Sonnet 4 via approved benchmark identity mapping map-claude-sonnet-4-6-to-4.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: claude-sonnet-4-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: nova-premier. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: llama4-scout. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vtb. Reported model configuration: llama4-maverick. Collapse policy: highest reported score per canonical model.