Benchmarks
The benchmarks.
What each one measures, where it comes from, and how recently it moved.
Loading benchmark profiles
| Intelligence Index Combined intelligence | AA | Text | Chat / text | 556 | Oct 3, 2026 |
| Humanity's Last Exam Hard knowledge | AA | Text | Reasoning / math / science | 530 | Oct 3, 2026 |
| GPQA Scientific reasoning | AA | Text | Reasoning / math / science | 494 | Oct 3, 2026 |
| CritPt Physics reasoning | AA | Text | Reasoning / math / science | 458 | Oct 3, 2026 |
| Long Context Reasoning Long-document reasoning | AA | Document | Long context | 456 | Oct 3, 2026 |
| AA-Omniscience accuracy Knowledge | AA | Text | Chat / text | 454 | Oct 3, 2026 |
| AA-Omniscience non-hallucination Knowledge honesty | AA | Text | Chat / text | 454 | Oct 3, 2026 |
| Text Arena Blind chat preference | AR | Text | Chat / text | 377 | Oct 2, 2026 |
| Input price API cost | AA | Text | Chat / text | 355 | Oct 3, 2026 |
| Output price API cost | AA | Text | Chat / text | 355 | Oct 3, 2026 |
| IFBench Instruction following | AA | Text | Chat / text | 349 | Oct 3, 2026 |
| Tau2-Bench Telecom Tool use | AA | Text | Search / tool use | 340 | Oct 3, 2026 |
| Terminal-Bench Hard Terminal agents | AA | Code | Coding | 333 | Oct 3, 2026 |
| Time to first token Latency | AA | Text | Chat / text | 265 | Oct 3, 2026 |
| Output Speed Generation speed | AA | Text | Chat / text | 263 | Oct 3, 2026 |
| Time to first answer token Latency | AA | Text | Chat / text | 263 | Oct 3, 2026 |
| MMMU-Pro Visual reasoning | AA | Vision | Vision understanding | 209 | Oct 3, 2026 |
| SciCode Scientific coding | AA | Code | Coding | 199 | Oct 3, 2026 |
| Text to Image Blind image generation preference | AA | Image | Image generation | 157 | Oct 3, 2026 |
| Vision Arena Blind multimodal preference | AR | Vision | Vision understanding | 140 | Oct 2, 2026 |