Visible tradeoffsThis is a rubric-judged signal, so it is more structured than arena taste but still depends on the scoring rubric.
source
Scale Labs
metric
Honesty score (%)
judge
Rubric
direction
higher better
group id
scale_mask_current
domain
Safety
What it measures vs what it misses
✓ Measures
Whether a model stays honest instead of covertly optimizing against the user.
✗ Misses
General capability breadth. Tool-use or retrieval quality.
Why this countsWhether a model stays honest instead of covertly optimizing against the user.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesGeneral capability breadth.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-opus-4-6 (Non-Thinking). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-sonnet-4-5-20250929-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude Sonnet 4 (Thinking). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-opus-4-1-20250805-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-opus-4-5-20251101-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-oss-120b. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5.4-pro-2026-03-05. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude Sonnet 4. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude Sonnet 4. Collapse policy: highest reported score per canonical model. Backfilled from Claude Sonnet 4 via approved benchmark identity mapping map-claude-sonnet-4-6-to-4.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude Opus 4 (Thinking). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-opus-4-1-20250805. Collapse policy: highest reported score per canonical model. Backfilled from Claude Opus 4.1 via approved benchmark identity mapping map-claude-opus-4-7-to-4-1.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-opus-4-5-20251101. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5.2-2025-12-11. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-oss-20b. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: claude-sonnet-4-5-20250929. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5.1-thinking. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5-pro-2025-10-06. Collapse policy: highest reported score per canonical model.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5-pro-2025-10-06. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-mini-to-gpt-5.
Fallback benchmark identity is visible for context but excluded from default ranking.
Identity
benchmark proxy (0.58)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5-pro-2025-10-06. Collapse policy: highest reported score per canonical model. Backfilled from GPT-5 via approved benchmark identity mapping map-gpt-5-4-nano-to-gpt-5.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o3 (high) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o3 (medium) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5-mini-2025-08-07. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o3 Pro (high) (June 2025). Collapse policy: highest reported score per canonical model.
82.5%
#25 · Claude 3.7 Sonnet (Thinking) (February 2025)
SL · Oct 3, 2026
Source label: Claude 3.7 Sonnet (Thinking) (February 2025)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude 3.7 Sonnet (Thinking) (February 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude Opus 4. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude 3 Opus. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o4-mini (high) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o4-mini (medium) (April 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude 3.5 Sonnet (October 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Claude 3.7 Sonnet (February 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: kimi-k2.5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt-5.1-instant. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o1-Pro. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: glm-4p5. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: GPT-4.1 nano. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Llama 3.1 405B Instruct. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: glm-4p5-air. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gpt 4o (November 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o1 (December 2024). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Deepseek R1 (Jan 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: GPT 4.5 Preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Mistral Magistral. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Qwen3-235B-A22B. Collapse policy: highest reported score per canonical model.
56.4%
#45 · Gemini 2.5 Pro Experimental (March 2025)
SL · Oct 3, 2026
Source label: Gemini 2.5 Pro Experimental (March 2025)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Gemini 2.5 Pro Experimental (March 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gemini-2.5-pro-preview-06-05. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Llama 3.2 90B Vision Instruct. Collapse policy: highest reported score per canonical model.
54.1%
#48 · Gemini 2.5 Pro Preview (May 06 2025)
SL · Oct 3, 2026
Source label: Gemini 2.5 Pro Preview (May 06 2025)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Gemini 2.5 Pro Preview (May 06 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: DeepSeek-R1-0528. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Llama 3.3 70B Instruct. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: GPT-4.1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: GPT-4.1 mini. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Llama 4 Maverick. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o3 mini (Low). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Gemini 2.0 Flash Thinking (January 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Gemini 2.5 Flash Preview (May 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Gemini 2.0 Flash. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o3 mini (Medium). Collapse policy: highest reported score per canonical model.
48.9%
#59 · Gemini 2.0 Pro Experimental (February 2025)
SL · Oct 3, 2026
Source label: Gemini 2.0 Pro Experimental (February 2025)
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Gemini 2.0 Pro Experimental (February 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gemini-3.1-flash-lite-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Mistral Large 2411. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: o3 mini (High). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: kimi-k2-instruct. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: DeepSeek-V3.1. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Deepseek V3 (March 2025). Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gemini-3-pro-preview. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: Mistral Medium 3. Collapse policy: highest reported score per canonical model.
Parsed from the public Scale Labs page for scale-mask. Reported model configuration: gemini-3.1-pro-preview. Collapse policy: highest reported score per canonical model.