| Intelligence Index AA · index Text · Chat / text | 5298.4%exact aliasverified runtime Row details- Raw value
- 52
- Percentile
- 98.4%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Source row | 23n/aexact aliasverified runtimeContext only Row details- Raw value
- 23
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| Humanity's Last Exam AA · % Text · Reasoning / math / science | 50%95.8%exact aliasverified runtime Row details- Raw value
- 50%
- Percentile
- 95.8%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Source row | 29.6%n/aexact aliasverified runtimeContext only Row details- Raw value
- 29.6%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| CritPt AA · % Text · Reasoning / math / science | 31.1%98.5%exact aliasverified runtime Row details- Raw value
- 31.1%
- Percentile
- 98.5%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Source row | 8.9%n/aexact aliasverified runtimeContext only Row details- Raw value
- 8.9%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| SciCode AA · % Code · Coding | 57.3%88.4%exact aliasverified runtime Row details- Raw value
- 57.3%
- Percentile
- 88.4%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Source row | 41%n/aexact aliasverified runtimeContext only Row details- Raw value
- 41%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| AA-Omniscience accuracy AA · % Text · Chat / text | 53%92.1%exact aliasverified runtime Row details- Raw value
- 53%
- Percentile
- 92.1%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Source row | 18.6%n/aexact aliasverified runtimeContext only Row details- Raw value
- 18.6%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| AA-Omniscience non-hallucination AA · % Text · Chat / text | 37.1%70.4%exact aliasverified runtime Row details- Raw value
- 37.1%
- Percentile
- 70.4%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Source row | 67%n/aexact aliasverified runtimeContext only Row details- Raw value
- 67%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |
| Long Context Reasoning AA · % Document · Long context | 79.7%87%exact aliasverified runtime Row details- Raw value
- 79.7%
- Percentile
- 87%
- Last updated
- recent
- Eligibility
- headline eligible
- Identity
- provider alias (0.94)
- Source label
- Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Source row | 70%n/aexact aliasverified runtimeContext only Row details- Raw value
- 70%
- Percentile
- n/a
- Last updated
- recent
- Eligibility
- benchmark_derived_model
- Identity
- provider alias (0.94)
- Source label
- A.X-K2
Source row | |