UAB
Home/Find
Find
Live · updated continuously

Describe the job.

Set filters, get ranked options with verified scores and source dates attached.
Recommendation · Research assistant

Best evidence-backed choices for Research assistant with all public sources: Claude Fable 5, Claude Opus 4.8, Claude Opus 4.7, and GPT-5.5.

Claude Fable 5Anthropic · frontier
Verified but aging

The current evidence supports a shortlist, not a single winner.

Direct sources
16
Coverage
80%
Data version
Aug 20, 2026
Evidence behind the call
Data version Aug 20, 202616 direct source linksNo excluded sourcesAll public sources
What the data can support5 ready, 1 partial, 1 missing for the current task result.
Minimum source requirementenough direct data · enough direct data · enough direct data · enough direct data
Cost / latencyLatency is unavailable in the current verified source data.
Current top picksClaude Fable 5, Claude Opus 4.8, Claude Opus 4.7, GPT-5.5
Answer typeTop picks, not one winner
Coverage80% visible · 80% checked
PresetResearch assistant
SourcesAll public sources
Latest strong source dataAug 20, 2026
Where sources differ4 source rows behind this answer
1
Claude Fable 5Artificial Analysis · Humanity's Last Exam · 55.468 · #1 · 100pctl
used for answer
2
Claude Fable 5Vals AI · CorpFin v2 · 71.834 · #1 · 100pctl
used for answer
3
Claude Fable 5Vals AI · ProofBench · 77 · #1 · 100pctl
used for answer
4
Claude Fable 5Vals AI · MMLU Pro · 91.502 · #1 · 100pctl
used for answer
Ranked options

Direct matches first, then weaker or missing-data cases

Direct matches stay strict; strong models with indirect data still surface below. Open a row for its scores, source links, and caveats.

Primary group · Direct-match leaders
#1Claude Fable 5Anthropic · frontier100.080% visible
enough direct data · recent · strong · clearfrontier-priced price band from registry metadataLatency is unavailable in the current verified source data.
Verified but aging3% score spread · 68.6% recent data · exact alias
Fit score100.0
Strongest source datareasoning math science · search tool use

Claude Fable 5 is strongest on Reasoning / math / science and Search / tool use for this preset.

  • Last updated: Latest visible source row is 0 days old.
  • Where sources differ: 3-point cross-domain spread; warning threshold is 30.
  • Data checks: No open review item was matched to this recommendation.

The verified evidence is decent, but too much of it is aging to treat this as a clean current winner.

Data parser or model matching changes recently moved Artificial Analysis, Vals AI, Arena.

Verified rows
17
Hand-checked rows
0
Copied rows
0
Backfilled rows
0
Headline lane
Aug 20, 2026
Background data
No extra context
Formula recommendation-fit-v2.0.0.
Measured by
public third-party sources
Source basis
exact alias
Compared fairly because
same test setup, version groups, source-balanced before averaging
Open model
Open compare
#2Claude Opus 4.8Anthropic · frontier93.280% visible
enough direct data · recent · strong · clearfrontier-priced price band from registry metadataLatency is unavailable in the current verified source data.
Verified but aging14.9% score spread · 68.6% recent data · exact alias
Fit score93.2
Strongest source datareasoning math science · long context

Claude Opus 4.8 is strongest on Reasoning / math / science and Long context for this preset.

  • Last updated: Latest visible source row is 0 days old.
  • Where sources differ: 15-point cross-domain spread; warning threshold is 30.
  • Data checks: No open review item was matched to this recommendation.

The verified evidence is decent, but too much of it is aging to treat this as a clean current winner.

Data parser or model matching changes recently moved Artificial Analysis, Vals AI, Arena.

Verified rows
17
Hand-checked rows
0
Copied rows
0
Backfilled rows
0
Headline lane
Aug 20, 2026
Background data
No extra context
Formula recommendation-fit-v2.0.0.
Measured by
public third-party sources
Source basis
exact alias
Compared fairly because
same test setup, version groups, source-balanced before averaging
Open model
Open compare
#3Claude Opus 4.7Anthropic · frontier85.7official80% visible
Official company result included
enough direct data · recent · strong · clearfrontier-priced price band from registry metadataLatency is unavailable in the current verified source data.
Verified but aging27.9% score spread · 63.2% recent data · exact alias
Fit score85.7
Strongest source datadocument understanding · reasoning math science

Claude Opus 4.7 is strongest on Document understanding and Reasoning / math / science for this preset.

  • Last updated: Latest visible source row is 65 days old.
  • Where sources differ: 28-point cross-domain spread; warning threshold is 30.
  • Data checks: No open review item was matched to this recommendation.

Provider-official evidence is self-reported company data. It can support recency, but it is not third-party test data.

The verified evidence is decent, but too much of it is aging to treat this as a clean current winner.

Some visible coverage is coming from provider-official source links while independent coverage catches up.

Verified rows
20
Hand-checked rows
2
Copied rows
0
Backfilled rows
0
Headline lane
Aug 20, 2026
Background data
No extra context
Formula recommendation-fit-v2.0.0.
Measured by
mixed independent and provider-official sources
Source basis
exact alias
Compared fairly because
same test setup, version groups, source-balanced before averaging
Open model
Open compare
#4GPT-5.5OpenAI · frontier82.3official80% visible
Official company result included
enough direct data · recent · strong · open reviewfrontier-priced price band from registry metadataLatency is unavailable in the current verified source data.
Verified but aging5.7% score spread · 61.5% recent data · exact alias
Fit score82.3
Strongest source datadocument understanding · long context

GPT-5.5 is strongest on Document understanding and Long context for this preset.

  • Last updated: Latest visible source row is 56 days old.
  • Where sources differ: 6-point cross-domain spread; warning threshold is 30.
  • Data checks: Open review items affect this model, source, or benchmark context.

Provider-official evidence is self-reported company data. It can support recency, but it is not third-party test data.

The verified evidence is decent, but too much of it is aging to treat this as a clean current winner.

Some visible coverage is coming from provider-official source links while independent coverage catches up.

Verified rows
30
Hand-checked rows
2
Copied rows
0
Backfilled rows
0
Headline lane
Aug 20, 2026
Background data
No extra context
Formula recommendation-fit-v2.0.0.
Measured by
mixed independent and provider-official sources
Source basis
exact alias
Compared fairly because
same test setup, version groups, source-balanced before averaging
Open model
Open compare
#5Qwen3.5 27BQwen · budget79.360% visible
Source status unavailablebudget price bandLatency unavailable
Visible tradeoffs9.4% score spread · 94% recent data · exact alias
Fit score79.3
Strongest source datasearch tool use · reasoning math science

Qwen3.5 27B is strongest on Search / tool use and Reasoning / math / science for this preset.

The visible evidence mix still leans on weaker or split signals, especially around Long context, source verification state, and any backfilled or relay evidence still in play.

Data parser or model matching changes recently moved Artificial Analysis, Arena.

Verified rows
7
Hand-checked rows
0
Copied rows
0
Backfilled rows
0
Headline lane
Aug 20, 2026
Background data
No extra context
Formula recommendation-fit-v2.0.0.

No source link clears the minimum source requirement.

Measured by
public third-party sources
Source basis
exact alias
Compared fairly because
same test setup, version groups, source-balanced before averaging
Open model
Open compare
Evidence & limits

Read the argument before you commit

Why these options made the listtop reasons behind the current answer
  • Current shortlist: Claude Fable 5, Claude Opus 4.8, Claude Opus 4.7, and GPT-5.5.
  • Claude Fable 5 is the strongest exact-match option still visible.
  • Claude Fable 5 currently leads the fit score at 100.0, but the evidence is still too mixed for a single headline winner.
What to pressure testwhere the current answer is still fragile
  • No single winner: The current public evidence is only strong enough to support a shortlist, not one winner.
  • Strongest alternative · Claude Opus 4.8: Claude Opus 4.8 is strongest on Reasoning / math / science and Long context for this preset.
  • Evidence risk: The verified evidence is decent, but too much of it is aging to treat this as a clean current winner.
What would flip the answerthe assumptions the result rests on
  • If you tighten benchmark spread: Claude Fable 5 still holds if you care more about aligned evidence than upside.
  • If you tighten recency: The winner becomes less stable quickly because the last-updated score is only 69 in the visible source data.
  • If you require open-weight: No open-weight model currently clears the same evidence floor.
  • If cost and speed matter more: No clearly cheaper alternative currently clears the same evidence floor.
Why this is not a clean winlimitations to keep in mind
  • The current evidence supports a shortlist, not a single winner.
  • Claude Opus 4.8 remains close enough that a different scoring recipe can still flip the public answer.
Source links6 rows behind this answer
Data limitswhere the public data is thin
  • cost: partial · Only registry price bands are ready; exact price calculations are not shown.
  • latency: missing · Latency is not available in the current verified shortlist data.
Not chosen6 well-known models left off

Gemini 2.5 Pro

Closest option

  • Gemini 2.5 Pro has direct evidence on part of this preset, but not enough to clear the exact-match floor.
Open model

GPT-5

Closest option

  • GPT-5 has direct evidence on part of this preset, but not enough to clear the exact-match floor.
Open model

Grok 4

Closest option

  • Grok 4 has direct evidence on part of this preset, but not enough to clear the exact-match floor.
Open model

Llama 4 Maverick

Closest option

  • Llama 4 Maverick has direct evidence on part of this preset, but not enough to clear the exact-match floor.
Open model

GPT-5.4

Closest option

  • Missing benchmark coverage in Embeddings / retrieval.
  • GPT-5.4 has direct evidence on part of this preset, but not enough to clear the exact-match floor.
Open model

Gemini 2.0 Pro Experimental

Known current model

  • Current generated catalog does not have enough matching source links for this task preset.
Open model
Needs more source data12 tracked models with thin public data

Gemini 2.5 Pro

Google · 100% visible · 100% direct · 0% indirect

Gemini 2.5 Pro has direct evidence on part of this preset, but not enough to clear the exact-match floor.

document understandingreasoning math science

GPT-5

OpenAI · 100% visible · 100% direct · 0% indirect

GPT-5 has direct evidence on part of this preset, but not enough to clear the exact-match floor.

embeddings retrievaldocument understanding

Grok 4

xAI · 100% visible · 100% direct · 0% indirect

Grok 4 has direct evidence on part of this preset, but not enough to clear the exact-match floor.

reasoning math sciencelong context

Llama 4 Maverick

Meta · 100% visible · 100% direct · 0% indirect

Llama 4 Maverick has direct evidence on part of this preset, but not enough to clear the exact-match floor.

long contextreasoning math science

GPT-5.4

OpenAI · 80% visible · 80% direct · 0% indirect

GPT-5.4 has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
reasoning math sciencedocument understanding

Grok 4.3

xAI · 80% visible · 80% direct · 0% indirect

Grok 4.3 has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
long contextreasoning math science

GPT-5.4 mini

OpenAI · 80% visible · 80% direct · 0% indirect

GPT-5.4 mini has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
reasoning math sciencedocument understanding

GPT-5.4 nano

OpenAI · 80% visible · 80% direct · 0% indirect

GPT-5.4 nano has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
reasoning math sciencelong context

Gemini 3.5 Flash

Google · 80% visible · 80% direct · 0% indirect

Gemini 3.5 Flash has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
reasoning math sciencedocument understanding

Gemini 3 Pro Preview

Google · 80% visible · 80% direct · 0% indirect

Gemini 3 Pro Preview has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
reasoning math sciencedocument understanding

Claude Sonnet 4.6

Anthropic · 80% visible · 80% direct · 0% indirect

Claude Sonnet 4.6 has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
reasoning math sciencedocument understanding

Grok 4.20

xAI · 80% visible · 80% direct · 0% indirect

Grok 4.20 has direct evidence on part of this preset, but not enough to clear the exact-match floor.

  • Missing benchmark coverage in Embeddings / retrieval.
reasoning math sciencesearch tool use

Shareable claims with evidence

The product should generate public claims worth checking, not just filter state.

Open change report
alert
8 review items still need manual judgment

The product keeps parser and mapping ambiguity visible instead of silently guessing.

Open
models
Arena moved via real benchmark movement

80 benchmark rows were added, 4 removed, and 16276 existing rows changed value or evaluation date. Window: 2026-06-20T23:37:10Z -> 2026-06-24T03:37:55Z.

Open
models
Artificial Analysis moved via real benchmark movement

28 benchmark rows were added, 0 removed, and 5949 existing rows changed value or evaluation date. Window: 2026-06-20T23:37:17Z -> 2026-06-24T03:38:09Z.

Open
models
BridgeBench moved via new benchmark coverage

10 benchmark rows were added, 10 removed, and 0 existing rows changed value or evaluation date. Window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z.

Open
sources
LLMBase moved via source updated leaderboard

The saved raw source snapshot changed relative to the previous run. Window: 2026-06-20T23:37:24Z -> 2026-06-24T03:38:25Z.

Open

What changed this week

alert
229 review items still need manual judgment

The product keeps parser and mapping ambiguity visible instead of silently guessing.

models
Arena moved via real benchmark movement

80 benchmark rows were added, 4 removed, and 16276 existing rows changed value or evaluation date. Window: 2026-06-20T23:37:10Z -> 2026-06-24T03:37:55Z.

Source-data window: 2026-06-20T23:37:10Z -> 2026-06-24T03:37:55Z

models
Artificial Analysis moved via real benchmark movement

0 benchmark rows were added, 1 removed, and 1 existing rows changed value or evaluation date. Window: 2026-08-20T04:01:14Z -> 2026-08-20T04:04:20Z.

Source-data window: 2026-08-20T04:01:14Z -> 2026-08-20T04:04:20Z

models
BridgeBench moved via new benchmark coverage

10 benchmark rows were added, 10 removed, and 0 existing rows changed value or evaluation date. Window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z.

Source-data window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z

models
LiveBench moved via new benchmark coverage

651 benchmark rows were added, 0 removed, and 0 existing rows changed value or evaluation date. Window: 2026-08-20T04:01:28Z -> 2026-08-20T04:04:39Z.

Source-data window: 2026-08-20T04:01:28Z -> 2026-08-20T04:04:39Z

product
Initial comparison-table release

Added comparison-table homepage, same-test normalization, per-cell source links, source pages, and custom-ranking preview.

Source-data window: 2026-04-16

models
Methodology contract published

Documented comparability rules, raw-vs-normalized behavior, and why unlike metrics are never averaged by default.

Source-data window: 2026-04-16

models
Artificial Analysis ID rule adopted

Stable model and creator IDs are now the preferred external identity keys when available.

Source-data window: 2026-04-15

Data version

Current snapshot.

Published Aug 20, 2026Model list checked2 stale sources216 claim warnings9 providers · 1154 tracked models

Freshness is calculated at request time. Failed sources retain their last verified rows with a visible age penalty. View live data status.

Quick routes

Jump straight to a page.

Resolve a recommendation into a public reportbest open model for long-context researchResearch assistantOpen page
Send a shortlist into compare modecompare gpt-5, claude opus, gemini proEveryday chatbotOpen page
Open a head-to-head debate pagegpt-5 vs claude opusEveryday chatbotOpen page
Open a source-difference reportbenchmark controversy for livebench codingCoding copilotOpen page
Open the latest public movementwhat changed this weekEveryday chatbotOpen page
Jump straight to an entity pageopen model gpt-5Open-weight shortlistOpen page