APEX-Agents-AA
AA · Professional reasoning · Objective
Long-horizon agentic task completion.
Rank #8 · Source label: GLM-5.2 (Max)
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Artificial Analysis
- Raw value
- 33.7%
- Percentile
- 78.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Artificial Analysis public leaderboard field `apexAgents`.
78.1% percentile inside its fair comparison set33.7%Raw benchmark value
Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #40 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,502
- Percentile
- 88.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: expert. Source rank: #41. Votes: 4851. Organization: zai. License: MIT.
88.1% percentile inside its fair comparison set1,502Raw benchmark valueCI 1,493 - 1,511
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #45 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,465
- Percentile
- 88.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_business_and_management_and_financial_operations. Source rank: #48. Votes: 8755. Organization: zai. License: MIT.
88.1% percentile inside its fair comparison set1,465Raw benchmark valueCI 1,458 - 1,473
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #37 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,446
- Percentile
- 90.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_entertainment_and_sports_and_media. Source rank: #39. Votes: 11229. Organization: zai. License: MIT.
90.4% percentile inside its fair comparison set1,446Raw benchmark valueCI 1,439 - 1,453
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #25 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,491
- Percentile
- 93.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_legal_and_government. Source rank: #25. Votes: 3713. Organization: zai. License: MIT.
93.1% percentile inside its fair comparison set1,491Raw benchmark valueCI 1,480 - 1,501
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #24 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,499
- Percentile
- 93.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_life_and_physical_and_social_science. Source rank: #24. Votes: 7518. Organization: zai. License: MIT.
93.9% percentile inside its fair comparison set1,499Raw benchmark valueCI 1,492 - 1,507
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #26 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,494
- Percentile
- 93%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_mathematical. Source rank: #27. Votes: 2480. Organization: zai. License: MIT.
93% percentile inside its fair comparison set1,494Raw benchmark valueCI 1,482 - 1,506
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #30 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,492
- Percentile
- 91.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_medicine_and_healthcare. Source rank: #30. Votes: 3358. Organization: zai. License: MIT.
91.6% percentile inside its fair comparison set1,492Raw benchmark valueCI 1,481 - 1,503
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #42 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,506
- Percentile
- 89.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_software_and_it_services. Source rank: #42. Votes: 17871. Organization: zai. License: MIT.
89.1% percentile inside its fair comparison set1,506Raw benchmark valueCI 1,500 - 1,512
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #36 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,459
- Percentile
- 90.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_writing_and_literature_and_language. Source rank: #38. Votes: 12098. Organization: zai. License: MIT.
90.7% percentile inside its fair comparison set1,459Raw benchmark valueCI 1,453 - 1,466
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #41 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,486
- Percentile
- 87.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: expert. Source rank: #42. Votes: 4851. Organization: zai. License: MIT.
87.8% percentile inside its fair comparison set1,486Raw benchmark valueCI 1,477 - 1,495
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #41 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,450
- Percentile
- 89.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_business_and_management_and_financial_operations. Source rank: #44. Votes: 8755. Organization: zai. License: MIT.
89.2% percentile inside its fair comparison set1,450Raw benchmark valueCI 1,443 - 1,457
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #26 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,449
- Percentile
- 93.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_entertainment_and_sports_and_media. Source rank: #27. Votes: 11229. Organization: zai. License: MIT.
93.3% percentile inside its fair comparison set1,449Raw benchmark valueCI 1,443 - 1,456
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #23 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,485
- Percentile
- 93.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_legal_and_government. Source rank: #24. Votes: 3713. Organization: zai. License: MIT.
93.7% percentile inside its fair comparison set1,485Raw benchmark valueCI 1,475 - 1,495
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #22 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,490
- Percentile
- 94.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_life_and_physical_and_social_science. Source rank: #23. Votes: 7518. Organization: zai. License: MIT.
94.4% percentile inside its fair comparison set1,490Raw benchmark valueCI 1,482 - 1,497
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #29 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,488
- Percentile
- 92.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_mathematical. Source rank: #31. Votes: 2480. Organization: zai. License: MIT.
92.1% percentile inside its fair comparison set1,488Raw benchmark valueCI 1,475 - 1,500
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #28 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,475
- Percentile
- 92.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_medicine_and_healthcare. Source rank: #29. Votes: 3358. Organization: zai. License: MIT.
92.2% percentile inside its fair comparison set1,475Raw benchmark valueCI 1,464 - 1,486
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #37 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,486
- Percentile
- 90.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_software_and_it_services. Source rank: #39. Votes: 17871. Organization: zai. License: MIT.
90.4% percentile inside its fair comparison set1,486Raw benchmark valueCI 1,480 - 1,491
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #28 · Source label: glm-5.2-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,462
- Percentile
- 92.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.2-max`. Category: industry_writing_and_literature_and_language. Source rank: #30. Votes: 12098. Organization: zai. License: MIT.
92.8% percentile inside its fair comparison set1,462Raw benchmark valueCI 1,456 - 1,469
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #36 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 31.3%
- Percentile
- 48.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: Zhipu AI.
48.5% percentile inside its fair comparison set31.3%Raw benchmark valueCI 24.9% - 37.6%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #22 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 7.1%
- Percentile
- 69.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: Zhipu AI.
69.6% percentile inside its fair comparison set7.1%Raw benchmark valueCI 3.2% - 11%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #40 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 84.1%
- Percentile
- 71.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: Zhipu AI.
71.9% percentile inside its fair comparison set84.1%Raw benchmark valueCI 83.2% - 85%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #37 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 49.7%
- Percentile
- 47.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: Zhipu AI.
47.8% percentile inside its fair comparison set49.7%Raw benchmark valueCI 48% - 51.4%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #59 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 40.8%
- Percentile
- 38.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: Zhipu AI.
38.9% percentile inside its fair comparison set40.8%Raw benchmark valueCI 36.5% - 45%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #44 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 83.5%
- Percentile
- 55.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: Zhipu AI.
55.2% percentile inside its fair comparison set83.5%Raw benchmark valueCI 79.6% - 87.5%
SkillsBench
VALS-AI · Professional reasoning · Objective
Applied professional skills tasks.
Rank #26 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 45.1%
- Percentile
- 24.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: skillsbench; provider: Zhipu AI.
24.2% percentile inside its fair comparison set45.1%Raw benchmark valueCI 36.5% - 53.7%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #49 · Source label: zai/glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 73.3%
- Percentile
- 62.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: Zhipu AI.
62.8% percentile inside its fair comparison set73.3%Raw benchmark valueCI 71.6% - 75%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #41 · Source label: glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 73.7%
- Percentile
- 38.5%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
38.5% percentile inside its fair comparison set73.7%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #46 · Source label: glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 73.2%
- Percentile
- 30.8%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
30.8% percentile inside its fair comparison set73.2%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #37 · Source label: glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 79.3%
- Percentile
- 44.6%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
44.6% percentile inside its fair comparison set79.3%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #48 · Source label: glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 45.9%
- Percentile
- 27.7%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
27.7% percentile inside its fair comparison set45.9%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #55 · Source label: glm-5.2
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 96.1%
- Percentile
- 20%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
20% percentile inside its fair comparison set96.1%Raw benchmark value