Visible tradeoffsThis is a rubric-judged signal, so it is more structured than arena taste but still depends on the scoring rubric.
source
Scale Labs
metric
Score (%)
judge
Rubric
direction
higher better
group id
scale_vista_current
domain
Vision understanding
What it measures vs what it misses
✓ Measures
Visual reasoning and understanding.
✗ Misses
Image generation or editing quality.
Why this countsIt is useful when the model must read charts, UI, screenshots, or visual scenes rather than text alone.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesIt does not tell you whether the model can generate or edit images well.
Leaderboard · this benchmark version
#1 · Gemini 2.5 Pro Experimental (March 2025)
SL · Oct 3, 2026
Source label: Gemini 2.5 Pro Experimental (March 2025)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.5 Pro Experimental (March 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gemini-2.5-pro-preview-06-05. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5.4-pro-2026-03-05. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5-pro-2025-10-06. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5-pro-2025-10-06. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5-pro-2025-10-06. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-nano-to-gpt-5.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: o4-mini (high) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: o4-mini (medium) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: o3 Pro (high) (June 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gemini-3-pro-preview. Collapse policy: highest reported score per canonical model.
51.5%
#11 · Gemini 2.5 Pro Preview (May 06 2025)
SL · Oct 3, 2026
Source label: Gemini 2.5 Pro Preview (May 06 2025)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.5 Pro Preview (May 06 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5-mini-2025-08-07. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: o3 (high) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: o3 (medium) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.5 Flash Preview (May 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-sonnet-4-5-20250929-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-opus-4-1-20250805-thinking. Collapse policy: highest reported score per canonical model.
48.4%
#18 · Claude 3.7 Sonnet (Reasoning)
SL · Oct 3, 2026
Source label: Claude 3.7 Sonnet Thinking (Feb 2025)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude 3.7 Sonnet Thinking (Feb 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: o1 Pro (March 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.5 Flash (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude Opus 4 (Thinking). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gemini-3.1-flash-lite-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5.2-2025-12-11. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-opus-4-5-20251101-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-opus-4-6-thinking-max. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.0 Flash Thinking Experimental. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude Sonnet 4 (Thinking). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: GPT-4.1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-opus-4-5-20251101. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: o1 (December 2024). Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model. Backfilled from Claude Opus 4.1 via approved benchmark identity mapping map-claude-opus-4-7-to-4-1.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: claude-sonnet-4-5-20250929. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5.1-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude Opus 4. Collapse policy: highest reported score per canonical model.
43.5%
#36 · Gemini 2.0 Pro Experimental (Feb 2025)
SL · Oct 3, 2026
Source label: Gemini 2.0 Pro Experimental (Feb 2025)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.0 Pro Experimental (Feb 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude Sonnet 4. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude Sonnet 4. Collapse policy: highest reported score per canonical model. Backfilled from Claude Sonnet 4 via approved benchmark identity mapping map-claude-sonnet-4-6-to-4.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude 3.7 Sonnet (February 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: GPT-4.5 Preview (February 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: kimi-k2.5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: GPT-4.1 mini. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.0 Flash (February 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude 3.5 Sonnet (October 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude 3.5 Sonnet (June 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Llama 4 Maverick. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: ChatGPT-4o-latest (November 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 1.5 Pro. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: GPT-4o (August 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: gpt-5.1-instant. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Mistral Medium 3. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 1.5 Flash 002. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Pixtral Large (November 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Gemini 2.0 Flash Lite Preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Qwen2-VL-72B-Instruct. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Claude 3 Opus. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: GPT-4.1 nano. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Nova Pro. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Pixtral 12B (September 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Nova Lite. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Llama 3.2 90B Vision Instruct. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Llama 3.2 11B Vision-Instruct. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-vista. Reported model configuration: Phi 3.5 Vision-Instruct. Collapse policy: highest reported score per canonical model.