Gemini 3.5 Flash
Closest option
- Gemini 3.5 Flash has direct evidence on part of this preset, but not enough to clear the exact-match floor.
The main warning is thinner evidence on Chat / text.
Direct matches stay strict; strong models with indirect data still surface below. Open a row for its scores, source links, and caveats.
claude-fable-5-high is strongest on Reasoning / math / science and Chat / text for this preset.
The main warning is thinner evidence on Chat / text.
claude-opus-4-6-high is strongest on Chat / text and Reasoning / math / science for this preset.
The main warning is thinner evidence on Reasoning / math / science.
Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) is strongest on Reasoning / math / science and Long context for this preset.
The main warning is thinner evidence on Chat / text.
claude-opus-5.5-high is strongest on Reasoning / math / science and Long context for this preset.
The main warning is thinner evidence on Chat / text.
Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) is strongest on Long context and Reasoning / math / science for this preset.
The main warning is thinner evidence on Chat / text.
No source link clears the minimum source requirement.
Closest option
Closest option
Closest option
Closest option
Closest option
Known current model
Google · 100% visible · 100% direct · 0% indirect
Gemini 3.5 Flash has direct evidence on part of this preset, but not enough to clear the exact-match floor.
OpenAI · 100% visible · 100% direct · 0% indirect
GPT-5.4 has direct evidence on part of this preset, but not enough to clear the exact-match floor.
OpenAI · 100% visible · 100% direct · 0% indirect
GPT-5.6 Sol has direct evidence on part of this preset, but not enough to clear the exact-match floor.
OpenAI · 100% visible · 100% direct · 0% indirect
GPT-5.6 Terra has direct evidence on part of this preset, but not enough to clear the exact-match floor.
DeepSeek · 100% visible · 100% direct · 0% indirect
deepseek-v4-pro has direct evidence on part of this preset, but not enough to clear the exact-match floor.
Zhipu · 100% visible · 100% direct · 0% indirect
glm-5.2-max has direct evidence on part of this preset, but not enough to clear the exact-match floor.
OpenAI · 100% visible · 100% direct · 0% indirect
GPT-5.4 mini has direct evidence on part of this preset, but not enough to clear the exact-match floor.
OpenAI · 100% visible · 100% direct · 0% indirect
GPT-5.4 nano has direct evidence on part of this preset, but not enough to clear the exact-match floor.
OpenAI · 100% visible · 100% direct · 0% indirect
GPT-5.5 has direct evidence on part of this preset, but not enough to clear the exact-match floor.
OpenAI · 100% visible · 100% direct · 0% indirect
GPT-5.6 Luna has direct evidence on part of this preset, but not enough to clear the exact-match floor.
xAI · 100% visible · 100% direct · 0% indirect
Grok 4.3 has direct evidence on part of this preset, but not enough to clear the exact-match floor.
xAI · 100% visible · 100% direct · 0% indirect
grok-4.5 has direct evidence on part of this preset, but not enough to clear the exact-match floor.
The product keeps parser and mapping ambiguity visible instead of silently guessing.
The saved raw source snapshot changed relative to the previous run. Window: 2026-10-03T11:41:25Z -> 2026-10-03T11:42:08Z.
10 benchmark rows were added, 10 removed, and 0 existing rows changed value or evaluation date. Window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z.
The saved raw source snapshot changed relative to the previous run. Window: 2026-10-03T11:41:42Z -> 2026-10-03T11:42:34Z.
Added comparison-table homepage, same-test normalization, per-cell source links, source pages, and custom-ranking preview.
The product keeps parser and mapping ambiguity visible instead of silently guessing.
The saved raw source snapshot changed relative to the previous run. Window: 2026-10-03T11:41:25Z -> 2026-10-03T11:42:08Z.
10 benchmark rows were added, 10 removed, and 0 existing rows changed value or evaluation date. Window: 2026-05-15T16:26:59Z -> 2026-05-15T16:34:46Z.
The saved raw source snapshot changed relative to the previous run. Window: 2026-10-03T11:41:42Z -> 2026-10-03T11:42:34Z.
Added comparison-table homepage, same-test normalization, per-cell source links, source pages, and custom-ranking preview.
Documented comparability rules, raw-vs-normalized behavior, and why unlike metrics are never averaged by default.
Stable model and creator IDs are now the preferred external identity keys when available.
Added alternate selectors for category headers after leaderboard markup drift.
Freshness is calculated at request time. Failed sources retain their last verified rows with a visible age penalty. View live data status.
best open model for long-context researchResearch assistantOpen pagecompare gpt-5, claude opus, gemini proEveryday chatbotOpen pagegpt-5 vs claude opusEveryday chatbotOpen pagebenchmark controversy for livebench codingCoding copilotOpen pagewhat changed this weekEveryday chatbotOpen pageopen model gpt-5Open-weight shortlistOpen page