Text Arena · Expert
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #19 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,518
- Percentile
- 94.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: expert. Source rank: #19. Votes: 1797. Organization: zai. License: MIT.
94.5% percentile inside its fair comparison set1,518Raw benchmark valueCI 1,504 - 1,532
Text Arena · Industry Business And Management And Financial Operations
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #30 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,475
- Percentile
- 92.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_business_and_management_and_financial_operations. Source rank: #31. Votes: 3346. Organization: zai. License: MIT.
92.1% percentile inside its fair comparison set1,475Raw benchmark valueCI 1,464 - 1,486
Text Arena · Industry Entertainment And Sports And Media
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #36 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,447
- Percentile
- 90.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_entertainment_and_sports_and_media. Source rank: #38. Votes: 4756. Organization: zai. License: MIT.
90.6% percentile inside its fair comparison set1,447Raw benchmark valueCI 1,437 - 1,456
Text Arena · Industry Legal And Government
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #50 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,470
- Percentile
- 86%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_legal_and_government. Source rank: #55. Votes: 1578. Organization: zai. License: MIT.
86% percentile inside its fair comparison set1,470Raw benchmark valueCI 1,455 - 1,486
Text Arena · Industry Life And Physical And Social Science
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #19 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,505
- Percentile
- 95.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_life_and_physical_and_social_science. Source rank: #19. Votes: 2984. Organization: zai. License: MIT.
95.2% percentile inside its fair comparison set1,505Raw benchmark valueCI 1,494 - 1,516
Text Arena · Industry Mathematical
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #18 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,499
- Percentile
- 95.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_mathematical. Source rank: #18. Votes: 939. Organization: zai. License: MIT.
95.2% percentile inside its fair comparison set1,499Raw benchmark valueCI 1,480 - 1,519
Text Arena · Industry Medicine And Healthcare
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #25 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,497
- Percentile
- 93.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_medicine_and_healthcare. Source rank: #25. Votes: 1282. Organization: zai. License: MIT.
93.1% percentile inside its fair comparison set1,497Raw benchmark valueCI 1,479 - 1,514
Text Arena · Industry Software And It Services
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #29 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,512
- Percentile
- 92.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_software_and_it_services. Source rank: #29. Votes: 6662. Organization: zai. License: MIT.
92.6% percentile inside its fair comparison set1,512Raw benchmark valueCI 1,504 - 1,520
Text Arena · Industry Writing And Literature And Language
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #27 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,466
- Percentile
- 93.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_writing_and_literature_and_language. Source rank: #29. Votes: 5117. Organization: zai. License: MIT.
93.1% percentile inside its fair comparison set1,466Raw benchmark valueCI 1,457 - 1,475
Text Arena · Expert · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena expert leaderboard.
Rank #17 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,514
- Percentile
- 95.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: expert. Source rank: #17. Votes: 1797. Organization: zai. License: MIT.
95.1% percentile inside its fair comparison set1,514Raw benchmark valueCI 1,500 - 1,528
Text Arena · Industry Business And Management And Financial Operations · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_business_and_management_and_financial_operations leaderboard.
Rank #30 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,461
- Percentile
- 92.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_business_and_management_and_financial_operations. Source rank: #31. Votes: 3346. Organization: zai. License: MIT.
92.1% percentile inside its fair comparison set1,461Raw benchmark valueCI 1,450 - 1,471
Text Arena · Industry Entertainment And Sports And Media · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_entertainment_and_sports_and_media leaderboard.
Rank #30 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,446
- Percentile
- 92.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_entertainment_and_sports_and_media. Source rank: #32. Votes: 4756. Organization: zai. License: MIT.
92.2% percentile inside its fair comparison set1,446Raw benchmark valueCI 1,436 - 1,455
Text Arena · Industry Legal And Government · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_legal_and_government leaderboard.
Rank #39 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,469
- Percentile
- 89.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_legal_and_government. Source rank: #41. Votes: 1578. Organization: zai. License: MIT.
89.1% percentile inside its fair comparison set1,469Raw benchmark valueCI 1,454 - 1,484
Text Arena · Industry Life And Physical And Social Science · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_life_and_physical_and_social_science leaderboard.
Rank #20 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,496
- Percentile
- 94.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_life_and_physical_and_social_science. Source rank: #20. Votes: 2984. Organization: zai. License: MIT.
94.9% percentile inside its fair comparison set1,496Raw benchmark valueCI 1,485 - 1,507
Text Arena · Industry Mathematical · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_mathematical leaderboard.
Rank #13 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,499
- Percentile
- 96.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_mathematical. Source rank: #13. Votes: 939. Organization: zai. License: MIT.
96.6% percentile inside its fair comparison set1,499Raw benchmark valueCI 1,480 - 1,518
Text Arena · Industry Medicine And Healthcare · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_medicine_and_healthcare leaderboard.
Rank #25 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,477
- Percentile
- 93.1%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_medicine_and_healthcare. Source rank: #26. Votes: 1282. Organization: zai. License: MIT.
93.1% percentile inside its fair comparison set1,477Raw benchmark valueCI 1,460 - 1,494
Text Arena · Industry Software And It Services · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_software_and_it_services leaderboard.
Rank #29 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,492
- Percentile
- 92.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_software_and_it_services. Source rank: #29. Votes: 6662. Organization: zai. License: MIT.
92.6% percentile inside its fair comparison set1,492Raw benchmark valueCI 1,484 - 1,500
Text Arena · Industry Writing And Literature And Language · No Style Control
AR · Professional reasoning · Human
Observed user preference in Arena's Text Arena industry_writing_and_literature_and_language leaderboard.
Rank #21 · Source label: glm-5.3-max
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Arena
- Raw value
- 1,467
- Percentile
- 94.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Arena leaderboard dataset row `glm-5.3-max`. Category: industry_writing_and_literature_and_language. Source rank: #22. Votes: 5117. Organization: zai. License: MIT.
94.7% percentile inside its fair comparison set1,467Raw benchmark valueCI 1,458 - 1,476
Vals Index
VALS-AI · Professional reasoning · Combined
Weighted model performance across economically relevant Vals tasks.
Rank #16 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 54
- Percentile
- 62.5%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: vals_index; provider: Fireworks AI.
62.5% percentile inside its fair comparison set54Raw benchmark valueCI 51 - 56
Legal Research Bench
VALS-AI · Professional reasoning · Objective
Applied legal research tasks.
Rank #7 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 49%
- Percentile
- 91.2%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_research; provider: Zhipu AI.
91.2% percentile inside its fair comparison set49%Raw benchmark valueCI 42.2% - 55.8%
Harvey's Legal Agent Benchmark
VALS-AI · Professional reasoning · Objective
Completing legal work with documents, spreadsheets, presentations, and file-system tools.
Rank #20 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 8.3%
- Percentile
- 73.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: hlab; provider: Zhipu AI.
73.9% percentile inside its fair comparison set8.3%Raw benchmark valueCI 4.4% - 12.2%
LegalBench
VALS-AI · Professional reasoning · Objective
Academic legal reasoning tasks.
Rank #28 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 84.8%
- Percentile
- 80%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: legal_bench; provider: Zhipu AI.
80% percentile inside its fair comparison set84.8%Raw benchmark valueCI 84.1% - 85.6%
Finance Agent v2
VALS-AI · Professional reasoning · Objective
Core financial analyst tasks for agentic models.
Rank #17 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 55.8%
- Percentile
- 76.8%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: fabv2; provider: Fireworks AI.
76.8% percentile inside its fair comparison set55.8%Raw benchmark valueCI 51.8% - 59.9%
MedCode
VALS-AI · Professional reasoning · Objective
Medical billing support and coding tasks.
Rank #47 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 42.9%
- Percentile
- 51.6%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medcode; provider: Zhipu AI.
51.6% percentile inside its fair comparison set42.9%Raw benchmark valueCI 38.7% - 47%
MedScribe
VALS-AI · Professional reasoning · Objective
Administrative documentation support for doctors.
Rank #9 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 88.8%
- Percentile
- 91.7%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: medscribe; provider: Zhipu AI.
91.7% percentile inside its fair comparison set88.8%Raw benchmark valueCI 84.9% - 92.7%
Public Benefits Bench
VALS-AI · Professional reasoning · Objective
Answering SNAP benefits questions across the public-benefits lifecycle.
Rank #8 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 68.5%
- Percentile
- 84.4%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: public-benefits-bench; provider: Zhipu AI.
84.4% percentile inside its fair comparison set68.5%Raw benchmark valueCI 66.2% - 70.9%
SkillsBench
VALS-AI · Professional reasoning · Objective
Applied professional skills tasks.
Rank #24 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 47.5%
- Percentile
- 30.3%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: skillsbench; provider: Zhipu AI.
30.3% percentile inside its fair comparison set47.5%Raw benchmark valueCI 38.6% - 56.4%
TaxEval v2
VALS-AI · Professional reasoning · Objective
Answer quality on tax questions and responses.
Rank #63 · Source label: zai/glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- Vals AI
- Raw value
- 72.4%
- Percentile
- 51.9%
- Last updated
- recent
- Eligibility
- headline eligible
Parsed from Vals AI BenchmarkView overall scores. Vals slug: tax_eval_v2; provider: Zhipu AI.
51.9% percentile inside its fair comparison set72.4%Raw benchmark valueCI 70.6% - 74.1%
Data analysis
LB · Professional reasoning · Objective
Structured data manipulation and table reasoning accuracy.
Rank #51 · Source label: glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 70.2%
- Percentile
- 23.1%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category: Data Analysis. Tasks scored: 3.
23.1% percentile inside its fair comparison set70.2%Raw benchmark value
Overall
LB · Professional reasoning · Objective
Average objective performance across LiveBench's current public category mix.
Rank #32 · Source label: glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 76.1%
- Percentile
- 52.3%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Category averages included: 7.
52.3% percentile inside its fair comparison set76.1%Raw benchmark value
Consecutive events
LB · Professional reasoning · Objective
Objective consecutive events score in LiveBench.
Rank #44 · Source label: glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 72.3%
- Percentile
- 33.8%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: consecutive_events. Category: Data Analysis.
33.8% percentile inside its fair comparison set72.3%Raw benchmark value
Table join
LB · Professional reasoning · Objective
Objective table join score in LiveBench.
Rank #63 · Source label: glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 42.3%
- Percentile
- 4.6%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablejoin. Category: Data Analysis.
4.6% percentile inside its fair comparison set42.3%Raw benchmark value
Table reformat
LB · Professional reasoning · Objective
Objective table reformat score in LiveBench.
Rank #58 · Source label: glm-5.3
verified runtimeexact alias
Raw row drilldownsource row, percentile, last updated, eligibility
- Source
- LiveBench
- Raw value
- 96.1%
- Percentile
- 20%
- Last updated
- stale
- Eligibility
- headline eligible
Derived from the official LiveBench website leaderboard table. Task: tablereformat. Category: Data Analysis.
20% percentile inside its fair comparison set96.1%Raw benchmark value