| Intelligence Index AA · index Text · Chat / text | 4494%exact aliasverified runtime Row details- Raw value
- 44
- Percentile
- 94%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 35n/aexact aliasverified runtimeContext only Row details- Raw value
- 35
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| Coding Index AA · index Code · Coding | 7494.7%exact aliasverified runtime Row details- Raw value
- 74
- Percentile
- 94.7%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
Source row | 39n/aexact aliasverified runtimeContext only Row details- Raw value
- 39
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| Agentic Index AA · index Code · Coding | 4688.7%exact aliasverified runtime Row details- Raw value
- 46
- Percentile
- 88.7%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
Source row | 26n/aexact aliasverified runtimeContext only Row details- Raw value
- 26
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| GPQA AA · % Text · Reasoning / math / science | 88.5%92.7%exact aliasverified runtime Row details- Raw value
- 88.5%
- Percentile
- 92.7%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 85.7%n/aexact aliasverified runtimeContext only Row details- Raw value
- 85.7%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| Humanity's Last Exam AA · % Text · Reasoning / math / science | 33.3%92.1%exact aliasverified runtime Row details- Raw value
- 33.3%
- Percentile
- 92.1%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 29.6%n/aexact aliasverified runtimeContext only Row details- Raw value
- 29.6%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| CritPt AA · % Text · Reasoning / math / science | 5.1%88.9%exact aliasverified runtime Row details- Raw value
- 5.1%
- Percentile
- 88.9%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 8.9%n/aexact aliasverified runtimeContext only Row details- Raw value
- 8.9%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| SciCode AA · % Code · Coding | 50.1%94.8%exact aliasverified runtime Row details- Raw value
- 50.1%
- Percentile
- 94.8%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 38.5%n/aexact aliasverified runtimeContext only Row details- Raw value
- 38.5%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| AA-Omniscience accuracy AA · % Text · Chat / text | 44.7%93.2%exact aliasverified runtime Row details- Raw value
- 44.7%
- Percentile
- 93.2%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 18.6%n/aexact aliasverified runtimeContext only Row details- Raw value
- 18.6%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| AA-Omniscience non-hallucination AA · % Text · Chat / text | 45.9%83.1%exact aliasverified runtime Row details- Raw value
- 45.9%
- Percentile
- 83.1%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 67%n/aexact aliasverified runtimeContext only Row details- Raw value
- 67%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| GDPval-AA AA · rating Text · Professional reasoning | 1,49089.4%exact aliasverified runtime Row details- Raw value
- 1,490
- Percentile
- 89.4%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
Source row | 1,109n/aexact aliasverified runtimeContext only Row details- Raw value
- 1,109
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| Long Context Reasoning AA · % Document · Long context | 72.3%90.4%exact aliasverified runtime Row details- Raw value
- 72.3%
- Percentile
- 90.4%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Opus 4.7 (Non-reasoning, High Effort)
Source row | 66%n/aexact aliasverified runtimeContext only Row details- Raw value
- 66%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |