Verified but agingThis is a rubric-judged signal, so it is more structured than arena taste but still depends on the scoring rubric.
source
BridgeBench
metric
Score (%)
judge
Rubric
direction
higher better
group id
bridgebench_ui_2026_05
domain
Coding
What it measures vs what it misses
✓ Measures
Completeness, visual quality, and interactivity for generated UI tasks.
✗ Misses
Full product design review. Accessibility audits beyond the benchmark rubric.
Why this countsIt tells you whether the model can generate, repair, and reason over code under evaluator pressure rather than marketing examples.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesIt does not fully capture repo-scale iteration, IDE ergonomics, or long debugging loops.