Verified but agingThis is a combined signal, so it bundles multiple inputs and should not be treated as one clean test.
source
BridgeBench
metric
Coverage (%)
judge
Combined
direction
higher better
group id
bridgebench_overall_coverage_2026_05
domain
Coding
What it measures vs what it misses
✓ Measures
How completely a model is represented across BridgeBench's overall benchmark set.
✗ Misses
The quality of the covered scores. Unpublished or private task coverage.
Why this countsIt tells you whether the model can generate, repair, and reason over code under evaluator pressure rather than marketing examples.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesIt does not fully capture repo-scale iteration, IDE ergonomics, or long debugging loops.