UAB
Home/Changelog
Changelog
Live · updated continuously
Operational history

Changelog

Parser changes, mapping fixes, methodology changes, and product releases stay visible because data plumbing changes what the site appears to know.
Entries · 7
Categories · parser / mapping / product / methodology

What changed this week

alert
229 review items still need manual judgment

The product keeps parser and mapping ambiguity visible instead of silently guessing.

models
Arena moved via real benchmark movement

80 benchmark rows were added, 4 removed, and 16276 existing rows changed value or evaluation date. Window: 2026-06-20T23:37:10Z -> 2026-06-24T03:37:55Z.

models
Artificial Analysis moved via real benchmark movement

0 benchmark rows were added, 1 removed, and 1 existing rows changed value or evaluation date. Window: 2026-08-19T04:01:26Z -> 2026-08-19T04:03:57Z.

models
BridgeBench moved via new benchmark coverage

10 benchmark rows were added, 10 removed, and 0 existing rows changed value or evaluation date. Window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z.

models
LiveBench moved via new benchmark coverage

651 benchmark rows were added, 0 removed, and 0 existing rows changed value or evaluation date. Window: 2026-08-19T04:01:41Z -> 2026-08-19T04:04:19Z.

product
Initial comparison-table release

Added comparison-table homepage, same-test normalization, per-cell source links, source pages, and custom-ranking preview.

models
Methodology contract published

Documented comparability rules, raw-vs-normalized behavior, and why unlike metrics are never averaged by default.

models
Artificial Analysis ID rule adopted

Stable model and creator IDs are now the preferred external identity keys when available.

2026-04-16
product
Initial comparison-table release Added comparison-table homepage, same-test normalization, per-cell source links, source pages, and custom-ranking preview.
2026-04-16
methodology
Methodology contract published Documented comparability rules, raw-vs-normalized behavior, and why unlike metrics are never averaged by default.
2026-04-15
mapping
Artificial Analysis ID rule adopted Stable model and creator IDs are now the preferred external identity keys when available.
2026-04-15
parser
BridgeBench parser fallback added Added alternate selectors for category headers after leaderboard markup drift.
2026-04-16
parser
LiveBench worker now feeds app bundle LiveBench records are now generated from the official public leaderboard dataset and merged into the catalog as a checked-in fragment with snapshot and parser metadata.
2026-04-16
mapping
Provider model registry added Current-model coverage now merges a generated registry sourced from official provider docs, with per-model verification links and a review queue for newly discovered names.
2026-04-16
methodology
Current models separated from historical benchmark identities New provider-verified variants such as GPT-5.4 and Claude Sonnet 4.5 now remain distinct from older benchmarked IDs so legacy scores are not silently relabeled as newer models.

What changed this week

alert
229 review items still need manual judgment

The product keeps parser and mapping ambiguity visible instead of silently guessing.

models
Arena moved via real benchmark movement

80 benchmark rows were added, 4 removed, and 16276 existing rows changed value or evaluation date. Window: 2026-06-20T23:37:10Z -> 2026-06-24T03:37:55Z.

Source-data window: 2026-06-20T23:37:10Z -> 2026-06-24T03:37:55Z

models
Artificial Analysis moved via real benchmark movement

0 benchmark rows were added, 1 removed, and 1 existing rows changed value or evaluation date. Window: 2026-08-19T04:01:26Z -> 2026-08-19T04:03:57Z.

Source-data window: 2026-08-19T04:01:26Z -> 2026-08-19T04:03:57Z

models
BridgeBench moved via new benchmark coverage

10 benchmark rows were added, 10 removed, and 0 existing rows changed value or evaluation date. Window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z.

Source-data window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z

models
LiveBench moved via new benchmark coverage

651 benchmark rows were added, 0 removed, and 0 existing rows changed value or evaluation date. Window: 2026-08-19T04:01:41Z -> 2026-08-19T04:04:19Z.

Source-data window: 2026-08-19T04:01:41Z -> 2026-08-19T04:04:19Z

product
Initial comparison-table release

Added comparison-table homepage, same-test normalization, per-cell source links, source pages, and custom-ranking preview.

Source-data window: 2026-04-16

models
Methodology contract published

Documented comparability rules, raw-vs-normalized behavior, and why unlike metrics are never averaged by default.

Source-data window: 2026-04-16

models
Artificial Analysis ID rule adopted

Stable model and creator IDs are now the preferred external identity keys when available.

Source-data window: 2026-04-15