Vals combined benchmark across finance, coding, and education tasks.
Source · Vals AI Version · vals-ai snapshot 2026-06-24 Scores · 20
Test details
Verified but agingThis is a combined signal, so it bundles multiple inputs and should not be treated as one clean test.
source
Vals AI
metric
Accuracy (%)
judge
Combined
direction
higher better
group id
vals_multimodal_index_current
domain
Document understanding
What it measures vs what it misses
✓ Measures
Weighted model performance across multimodal professional Vals tasks.
✗ Misses
Pure text chat preference, media generation, latency, and cost.
Why this countsIt matters when the job is reading PDFs, tables, forms, or mixed-layout documents rather than plain chat.Same-test ruleThis percentile only compares models inside the exact benchmark/version group shown here. It is not a universal score.What it missesIt does not stand in for end-to-end document workflow quality or search relevance.