Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
GLM-5.2 · Evaluations · Kaino
GLM-5.2 logo
model evaluation

GLM-5.2

Z.AI

Z.AI flagship text foundation model for long-horizon coding and agentic engineering, with 1M context, 128K output, thinking modes, and MIT-licensed open weights.

modelZ.AIGLM
78.2KAINO SCORERecommended
Evaluated Jul 31, 202611 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness87
  • Technical capability84
  • Coding & agentic84
  • Speed & availability83
  • Developer experience82
  • Reasoning & knowledge78
  • Pricing clarity77
  • Risk & evidence76
  • Adoption signal76
  • Multimodal & I/O55

Kainotomic evaluation

GLM-5.2 has credible high-end coding evidence: 78.7% ±1.9% on Epoch’s SWE-bench Verified leaderboard, plus a 51 Artificial Analysis Intelligence Index. Its documented 1M-token context, 128K output, open weights, and MIT license are meaningful technical and deployment advantages. It exceeds the published GLM-4.6 anchor in capability, coding, speed, developer experience, and public signal on materially stronger current evidence. Coding evidence is mixed rather than uniformly leading. The independent DeepSWE v1.1 result—44% ±2%, 14th of 18 displayed models—prevents a score near GPT-5.5, GPT-5.6 Sol, or Claude Opus 4.8, all of which have substantially higher catalog coding anchors. It is broadly comparable to Kimi K2.5 on coding-oriented use, but trails its 86 coding score because the supplied independent agent benchmark is weak. No GLM-5.2 entry was found on LiveCodeBench V5. Value and serving evidence are strong: Artificial Analysis lists $1.40/M input and $4.40/M output, 115.8 output tokens/s and 1.43s median TTFT, while provider tracking reports 15 providers and lower-cost alternatives. Pricing clarity is lower than the price score because supplied official pricing excerpts omit a durable rate table. Multimodal capability remains conservatively low: the supplied materials establish a text model, not comparable image/audio I/O. Official Terminal Bench and SWE-bench Pro claims are useful but vendor-reported and not treated as independent performance proof.

Strengths

  • Independent 78.7% ±1.9% SWE-bench Verified result.
  • 1M-token context, 128K output, MIT-licensed open weights, and public repository/model card.
  • Competitive independently tracked API price, latency, throughput, and multi-provider availability.

Caveats

  • DeepSWE v1.1 is only 44% ±2%, ranked 14th of 18 displayed models.
  • No matching GLM-5.2 result on the checked LiveCodeBench V5 leaderboard.
  • Supplied evidence supports text workflows; it does not substantiate broad multimodal I/O.