Skip to content
news22 min read

GLM-5.3 Ships at GLM-5.2's Exact Price — and One Old Parameter Now Fails

GLM-5.3 landed on August 18, 2026 at GLM-5.2's exact rates, $1.40 per million input tokens and $4.40 per million output tokens, and scores 60 against 53 on version 4.1.1 of the Artificial Analysis Intelligence Index. Reasoning can no longer be disabled: that request now fails.

Author
Anthony M.
22 min readVerified August 27, 2026Tested hands-on
GLM-5.3 scores 60 on the Artificial Analysis Intelligence Index at the same list price as GLM-5.2, which scores 53
Measured by Artificial Analysis on Intelligence Index version 4.1.1, read August 22, 2026: GLM-5.3 at max effort scores 60 and GLM-5.2 at max effort scores 53, on identical list rates and near-identical throughput.

Z.ai released GLM-5.3 on August 18, 2026. On version 4.1.1 of the Artificial Analysis Intelligence Index, an evaluation run by a third party rather than by Z.ai, GLM-5.3 at max reasoning effort scores 60, against 53 for GLM-5.2 at its own max effort, read August 22, 2026. The list price did not move by a cent: Z.ai's pricing page bills GLM-5.3 at $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens — the exact rates already charged for GLM-5.2 and GLM-5.1. Z.ai's own documentation states that GLM-5.3 uses the same base model as GLM-5.2 and that every improvement comes from post-training. One change is not backward compatible: thinking.type: "disabled" is no longer supported, and a request that still sends it will fail.

Key takeaways

  • Seven index points for zero extra dollars per token. Artificial Analysis moves GLM-5.3 to 60 from GLM-5.2's 53 on the same version of the same index, while Z.ai's pricing page shows identical rates for GLM-5.3, GLM-5.2 and GLM-5.1.
  • It ties, it does not lead. Among models from Chinese labs on that leaderboard, GLM-5.3 shares the top score of 60 with Kimi K3 at max effort. It is the faster and cheaper of the two on Artificial Analysis's own measurements — 93 output tokens per second against 38, and $0.68 per index task against $0.84.
  • Identical per-token price, higher measured cost per task. Artificial Analysis puts GLM-5.3 at $0.68 per Intelligence Index task against $0.44 for GLM-5.2, because the new model emitted about 21 percent more output tokens across the same nine evaluations.
  • The migration order matters. Z.ai's documentation says to switch thinking.type to enabled with reasoning_effort set to low before pointing the model ID at glm-5.3. Do it the other way around and the request fails.
  • We could not confirm the open-source claim. Z.ai's release notes claim state of the art "among open-source models", but we did not find published GLM-5.3 weights on Hugging Face on August 22, 2026, and Artificial Analysis now labels GLM-5.3 a proprietary model where it labels GLM-5.2 open weights.

What did Z.ai actually ship on August 18, 2026?

Z.ai's release notes carry a single entry dated 2026-08-18 for GLM-5.3, and it makes two claims. The first is coding: a 50 percent gain over GLM-5.2 on Z.ai Code Bench, plus state of the art "among open-source models" on public benchmarks including Terminal Bench 3.0. The second is cybersecurity: that GLM-5.3 "matches Mythos 5 in white-box code review and vulnerability discovery", and that in collaboration with multiple cybersecurity teams it was tested on real-world targets and identified 2,436 vulnerabilities, of which 1,097 were medium or high severity.

Z.ai Code Bench is Z.ai's own benchmark. A 50 percent gain measured on it is a vendor result, not an independent one, and the release notes do not publish the underlying task set or the harness. The model page adds one more vendor benchmark to the coding claim, Agents' Last Exam in its CLI form, alongside Terminal Bench 3.0.

The dedicated model page carries the specification detail. GLM-5.3 accepts text only — there is no vision input — with a 1 million token context window and a maximum output length of 128,000 tokens. It reaches the API through three protocols: an OpenAI Chat Completion endpoint at https://api.z.ai/api/coding/paas/v4, an OpenAI Response endpoint at https://api.z.ai/api/v1, and an Anthropic Messages endpoint at https://api.z.ai/api/anthropic. There is a restriction buried in that section worth reading twice: anyone who has previously subscribed to a GLM Coding Plan, including an expired subscription, can currently reach the model API only through the OpenAI Chat Completion-compatible protocol. Z.ai says it will improve this in later iterations.

The same page describes the Coding Plan quota as points-based, and states that calls made during off-peak hours — which Z.ai defines as including all day on weekends — consume only 50 percent of the standard points. That is a pricing lever on the subscription side with no equivalent on the per-token side, where the rates are flat.

The price did not move, and that is the unusual part

GLM-5.1, GLM-5.2 and GLM-5.3 shown side by side as three identical display cases on one shared price card
Z.ai list prices per million tokens, read from the vendor pricing page on August 22, 2026. GLM-5.3, GLM-5.2 and GLM-5.1 sit on identical rates.

Z.ai's pricing page, read on August 22, 2026, lists GLM-5.3 at $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens, with cached input storage marked free for a limited time. The rows for GLM-5.2 and GLM-5.1 carry the same three numbers. A generation of capability was added to the same line on the invoice.

ModelInput per 1M tokensCached input per 1M tokensOutput per 1M tokens
GLM-5.3$1.40$0.26$4.40
GLM-5.2$1.40$0.26$4.40
GLM-5.1$1.40$0.26$4.40
GLM-5-Turbo$1.20$0.24$4.00
GLM-5$1.00$0.20$3.20
GLM-4.7$0.60$0.11$2.20

Source: Z.ai pricing documentation, read August 22, 2026. Cached input storage is listed as free for a limited time on every row above.

The cache discount is the same on both models and it is steep: $0.26 against $1.40 is an 81 percent reduction on cached input, which is the figure Artificial Analysis also displays on both model pages. For a workload that re-sends a large stable prefix — a repository, a specification, a long system prompt — that discount, not the headline input rate, is the number that decides the bill.

Holding a price flat across a capability jump is a decision, not a default. Anthropic made the same one in July when Claude Opus 5 launched on Opus 4.8's price card, and DeepSeek made it elsewhere in its catalog when its cheapest model overtook its own flagship at unchanged rates. What each of those has in common with GLM-5.3 is that the vendor absorbed the improvement rather than repricing it.

What did the independent evaluation actually measure?

Artificial Analysis scores GLM-5.3 at max reasoning effort at 60 on version 4.1.1 of its Intelligence Index, read August 22, 2026. That version aggregates nine evaluations, which Artificial Analysis names on the page: GDPval-AA v2, Tau-cubed Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. The same index version scores GLM-5.2 at max effort at 53, which makes the seven-point gap a like-for-like comparison rather than two readings taken against different rulers.

Model (effort)Intelligence Index v4.1.1Cost per index taskMedian output tokens per second
Claude Opus 5 (max)63$2.3458
Claude Fable 5 (with fallback)62$3.1470
GPT-5.6 Sol (max)61$1.2370
Grok 4.6 (high)61$0.8469
Kimi K3 (max)60$0.8438
GLM-5.3 (max)60$0.6893
Qwen3.8 Max58$1.1345
GLM-5.2 (max)53$0.4492
DeepSeek V4 Pro 0813 (max)53$0.2581

Source: Artificial Analysis model leaderboard, Intelligence Index version 4.1.1, read August 22, 2026. Effort configurations are listed separately by the evaluator and are not interchangeable.

Two things in that table are easy to misread. The first is that effort levels are separate entries, not variants of one number: Claude Opus 5 appears at 63 at max, 63 at xhigh, 61 at high, 59 at medium and 52 at low. Comparing a max-effort score against a medium-effort score describes a configuration choice, not a model. The second is that Artificial Analysis ranks each model inside a comparison class rather than against the whole field. Its GLM-5.3 page displays an intelligence rank of 9 out of 186, where those 186 are the models it groups with GLM-5.3 — proprietary and open-weight models sharing the same blended price band. GLM-5.2's page shows rank 4 out of 107 against a different set. The two rank numbers describe different contests and cannot be subtracted from each other.

The speed reading barely moved: 93.2 output tokens per second for GLM-5.3 against 92.5 for GLM-5.2. Whatever post-training did, it did not cost throughput.

Is GLM-5.3 the strongest model from a Chinese lab?

It is tied for the strongest, not alone at the top. On the Artificial Analysis leaderboard read August 22, 2026, Kimi K3 at max effort also scores 60, and the evaluator lists it above GLM-5.3. Below them, Qwen3.8 Max and Qwen3.8 2.4T A95B both score 58, and DeepSeek V4 Pro 0813 at max effort scores 53, level with GLM-5.2. Anyone writing that GLM-5.3 is "the best Chinese model" is reporting a tie as a win.

Where GLM-5.3 does separate itself from Kimi K3 is on the two axes that are not intelligence. Artificial Analysis measures GLM-5.3 at 93 output tokens per second against 38 for Kimi K3 at max — roughly two and a half times the throughput — and at $0.68 per index task against $0.84. Same score, faster, cheaper to push through the same evaluation. For an agent loop that spends its day waiting on tokens, that difference is the whole argument. Our existing head-to-head on GLM-5.2 against Kimi K3 was written against the older GLM scores and now understates the GLM side.

Counting effort configurations separately, nine leaderboard entries score 60 or above. They come from six distinct models across five vendors. GLM-5.3 is inside that group; it is not at the front of it.

Same price per token, higher cost per task

Identical per-token rate for GLM-5.2 and GLM-5.3 next to the different output token volumes Artificial Analysis recorded across the index
Per-token rates are identical, but the measured cost per Intelligence Index task is not: Artificial Analysis recorded 170 million output tokens for GLM-5.3 against 140 million for GLM-5.2, read August 22, 2026.

A flat price per token does not mean a flat price per job. Artificial Analysis measures cost per Intelligence Index task at $0.68 for GLM-5.3 and $0.44 for GLM-5.2 — about 55 percent more for the newer model, on rates that are identical to the cent. The evaluator also publishes what it spent in total: $1,238.50 to run GLM-5.3 through the index, against $843.44 for GLM-5.2, roughly 47 percent more.

The reason sits in the token counts. Artificial Analysis records 170 million output tokens generated by GLM-5.3 across the index against 140 million for GLM-5.2, about 21 percent more, and flags GLM-5.3 as very verbose against a median of 72 million for its comparison class. Reasoning is always on and the default effort is max, so the model thinks at length unless told otherwise. The gap between plus 21 percent output tokens and plus 47 percent total spend implies input volume grew as well, which is what happens when longer reasoning traces are carried forward through multi-turn evaluations. Artificial Analysis publishes the totals but not that decomposition, so we are describing the direction, not the arithmetic.

The practical version: if your budget is modeled on tokens, nothing changed. If it is modeled on completed tasks, assume the same job costs meaningfully more on GLM-5.3 at default settings than it did on GLM-5.2, and that reasoning_effort: low is the lever that brings it back down.

Same base model, all gains from post-training

The most interesting sentence in the GLM-5.3 documentation is one most vendors would not write. Z.ai states that GLM-5.3 "uses the same base model as GLM-5.2, with all improvements driven by post-training". No new pre-training run, no new parameter count, no new architecture — a seven-point move on an independent index attributed entirely to what happened after the base model was frozen.

Vendors usually leave that ambiguous, because a version bump implies a new model and a new model implies new investment. Saying it out loud is a claim about where the remaining headroom is: in reinforcement learning, in tool-use training, in the shaping of long-horizon behavior, rather than in the next order of magnitude of pre-training compute. The Emergent Cyber Capability section of the same page makes the point explicitly, describing cybersecurity ability as improving "as the scale of post-training continues to expand".

It also has a mundane consequence worth knowing before a migration: the specification sheet does not change. Same 1 million token context, same text-only input, and an output ceiling documented at 128,000 tokens. If a system was engineered around GLM-5.2's shape, GLM-5.3 fits the same slot — with one exception, which is the next section.

The breaking change: no reasoning-disabled mode

GLM-5.3 always operates with reasoning enabled. Z.ai's documentation states that disabling reasoning is no longer supported, and that a request still using thinking.type: "disabled" will fail. This is a silent trap for any codebase that was calling GLM-5.2 in non-reasoning mode for cheap, fast, low-stakes calls: the model ID swap alone turns a working request into an error.

ParameterAccepted valuesDefaultBehavior
thinking.typeenabled onlyenabledReasoning cannot be disabled
reasoning_effortlow, high, maxmaxLightweight, enhanced and deep reasoning respectively

Source: Z.ai GLM-5.3 model documentation, read August 22, 2026.

The migration path, in the order Z.ai specifies

Z.ai's migration notice is explicit about sequencing: change thinking.type to enabled and set reasoning_effort to low before updating the model ID to glm-5.3. Otherwise, the documentation says, the request will fail. Doing it in that order means the failure mode never reaches production: step one is a no-op against GLM-5.2, step two is the actual switch.

The call that breaks:

{
  "model": "glm-5.2",
  "thinking": { "type": "disabled" }
}

Step one — still on the old model, already compatible:

{
  "model": "glm-5.2",
  "thinking": { "type": "enabled" },
  "reasoning_effort": "low"
}

Step two — flip the model ID only once step one is deployed everywhere:

{
  "model": "glm-5.3",
  "thinking": { "type": "enabled" },
  "reasoning_effort": "max"
}

For complex tasks such as coding, Z.ai recommends max, and max is also what the API applies when reasoning_effort is absent. That default is the one to be deliberate about: an omitted parameter now buys the most expensive reasoning tier, which is exactly the setting Artificial Analysis measured at $0.68 per task.

One more thing to check before scheduling the migration. If your organization has ever held a GLM Coding Plan subscription, expired ones included, Z.ai's documentation says the model API is currently reachable only through the OpenAI Chat Completion-compatible protocol. A service written against the Anthropic Messages endpoint at https://api.z.ai/api/anthropic may need a transport change as well as a parameter change.

The cybersecurity numbers, and what each one is attached to

Which GLM-5.3 claims come from Z.ai and which come from an independent evaluator
The index score comes from a third party. The coding gain, the cybersecurity totals and the Mythos 5 comparison are all published by Z.ai.

Z.ai publishes three separate cybersecurity claims for GLM-5.3, and they are attached to three different things. The 2,436 vulnerabilities figure, of which 1,097 are described as medium or high severity, comes from the release notes and is attributed to testing on real-world targets alongside multiple cybersecurity teams — not to CyberGym. The CyberGym claim is separate and sits on the model page: best performance to date on that vulnerability discovery benchmark. The third is a comparison: that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery. All three are Z.ai's own statements.

The Mythos 5 comparison deserves the most caution, and not because of Z.ai. Anthropic's Claude Mythos 5 sits behind invitation-only trusted access, which is precisely why its behavior in cyber evaluations is documented mainly through Anthropic's own disclosures. A claim to match a model that almost nobody outside a small access list can benchmark is a claim that almost nobody outside that list can check.

The CyberGym claim runs into a structural issue we have covered before. When Microsoft announced 96 percent on CyberGym in July, the score turned out to belong to a harness running two models rather than to a single model. CyberGym results are reported for a system: model plus scaffolding plus tool access plus attempt budget. Z.ai's page states an outcome without stating the harness, which means the number cannot be lined up against another lab's CyberGym number even when both are honestly reported.

Z.ai also reports that GLM-5.3's scores on vulnerability exploitation benchmarks exceed twice those of GLM-5.2, and that the gains grow the further the model progresses along the exploitation chain. That is a relative claim against the vendor's own previous model, which is the kind of comparison a vendor is best placed to make and an outsider is least able to audit.

The open-source claim we could not confirm

Z.ai's release notes claim state of the art "among open-source models" for GLM-5.3. On August 22, 2026 we looked for published GLM-5.3 weights and did not find them. Querying the Hugging Face API for the zai-org account sorted by creation date, the most recent public repository is GLM-5.2, dated June 16, 2026 — the same account that published GLM-5, GLM-5.1 and GLM-5.2. A global Hugging Face search for "GLM-5.3" returns three repositories, all from community accounts, and the only GGUF among them contains just a .gitattributes file and a README.

On ModelScope, the two repository paths that serve GLM-5.1 and GLM-5.2 both answer with HTTP 200, while the equivalent GLM-5.3 paths under both ZhipuAI and zai-org return not-found. We did not find a working search endpoint on that platform, so we checked direct paths only. Z.ai's own launch post at z.ai/blog/glm-5.3, linked from the documentation, returned no server-rendered text to our fetch, so we did not use it as a source.

What we are stating is the outcome of a search, not the state of the world. We did not find published weights in the places where the previous three GLM releases published theirs. That is different from saying the weights are not published, and the distinction is not pedantic — a release can land on a mirror, behind a gate, or on a platform we did not query.

One independent data point does lean the same way. Artificial Analysis labels GLM-5.3 a proprietary model on its page and places it in a proprietary comparison class of 186 models, while it labels GLM-5.2 an open weights model and ranks it inside a class of 107. Whatever GLM-5.2's status — its Hugging Face repository carries an MIT license and has been downloaded more than 2.7 million times — the evaluator did not carry that classification forward to GLM-5.3.

Two readings of the release-notes sentence stay open, and we are not choosing between them. It may mean GLM-5.3 is itself open and released somewhere we did not look. It may equally mean GLM-5.3 leads the field of open-source models on those benchmarks — a claim about the competition, not about GLM-5.3's own license. The sentence supports both, and the evidence available on August 22, 2026 does not settle it.

Who should move, and who should wait

The case for moving is straightforward for anyone already paying GLM-5.2 rates for agentic coding: the same invoice line buys a model that measures seven points higher on an independent index at effectively unchanged throughput. The case for waiting is narrower but real, and it comes down to three specific situations.

Wait if you rely on a non-reasoning mode. Teams using GLM-5.2 with reasoning disabled for classification, routing, extraction or any high-volume low-stakes call have no equivalent on GLM-5.3. The nearest option is reasoning_effort: low, which is not the same thing — it is lightweight reasoning, not no reasoning, and it emits tokens accordingly.

Wait if you need self-hosting. If your deployment depends on running the weights yourself, GLM-5.2 remains the newest GLM release we could find published, under an MIT license. Until GLM-5.3 weights turn up somewhere verifiable, an on-premises plan cannot be built on it.

Wait if you need vision. GLM-5.3 is text-only. Z.ai's multimodal line runs separately through GLM-5V-Turbo and GLM-4.6V, and this release does not change that split.

For everyone else, the migration is a two-line parameter change made in the documented order, followed by a decision about reasoning_effort that will drive your bill more than the model swap does. Our published comparisons — including Claude Opus 5 against GLM-5.2 and GLM-5.2 against DeepSeek V4 — were written against GLM-5.2's numbers and should be read as describing the previous generation until we refresh them.

What would change this reading

Three things. If GLM-5.3 weights appear under a permissive license on Hugging Face or ModelScope, the open-source sentence resolves in the first direction and the self-hosting caveat disappears; Artificial Analysis would presumably reclassify the model, which would also move its displayed rank into a different class. If an independent party publishes a CyberGym run for GLM-5.3 with the harness described, the cybersecurity claims become checkable rather than merely reported. And index scores are re-run: Artificial Analysis states version 4.1.1 on both GLM pages today, but a later version, or a re-measurement of any model in the table above, can move these numbers. Every figure attributed to Artificial Analysis in this article was read on August 22, 2026.

The bottom line

GLM-5.3 is a post-training release that bought seven index points without touching the price card, which is a genuinely good outcome for anyone already on GLM-5.2 rates. It is not the runaway result the phrase "best Chinese model" would suggest — it ties Kimi K3 at 60 and trails Claude Opus 5, Claude Fable 5, GPT-5.6 Sol and Grok 4.6 — but it is meaningfully faster and cheaper per task than the model it ties with. The claims that would make it more than that, on CyberGym and against Mythos 5, are all Z.ai's own, and the open-source framing does not match what we could find published on August 22, 2026. The one item that belongs on a sprint board today is the parameter change: the reasoning-disabled mode is gone, and the model ID must be the last thing you change, not the first.

GLM-5.3 FAQ

What is GLM-5.3?

GLM-5.3 is Z.ai's flagship language model, released on August 18, 2026. Z.ai's documentation states that it uses the same base model as GLM-5.2, with all improvements driven by post-training. It accepts text input only, carries a 1 million token context window and a maximum output length of 128,000 tokens, and always runs with reasoning enabled at one of three effort levels: low, high or max.

How much does GLM-5.3 cost?

Z.ai's pricing page, read on August 22, 2026, lists GLM-5.3 at $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens, with cached input storage free for a limited time. Those are the same three rates the page shows for GLM-5.2 and GLM-5.1, so the newer model did not come with a price increase.

Is GLM-5.3 more expensive to run than GLM-5.2?

Per token, no — the rates are identical. Per completed task, yes. Artificial Analysis measured a cost of $0.68 per Intelligence Index task for GLM-5.3 against $0.44 for GLM-5.2, about 55 percent more, because GLM-5.3 emitted roughly 21 percent more output tokens across the same nine evaluations. Reasoning is always on and the default effort is max, so a call that omits reasoning_effort buys the most expensive tier.

What breaks when I switch to GLM-5.3?

Any request that sends thinking.type set to disabled. Z.ai's documentation states that disabling reasoning is no longer supported on GLM-5.3 and that such a request will fail. The documented migration path is to change thinking.type to enabled and set reasoning_effort to low before updating the model ID to glm-5.3, so that the parameter change is already deployed when the model swap happens.

How does GLM-5.3 score on an independent benchmark?

Artificial Analysis scores GLM-5.3 at max reasoning effort at 60 on version 4.1.1 of its Intelligence Index, read on August 22, 2026. That version aggregates nine evaluations: GDPval-AA v2, Tau-cubed Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. The same index version scores GLM-5.2 at max effort at 53.

Is GLM-5.3 the best Chinese AI model?

It is tied for the best score on that index, not alone at the top. On the Artificial Analysis leaderboard read August 22, 2026, Kimi K3 at max effort also scores 60 and is listed above GLM-5.3. Qwen3.8 Max scores 58 and DeepSeek V4 Pro 0813 at max effort scores 53. GLM-5.3 does beat Kimi K3 on the other two measurements: 93 output tokens per second against 38, and $0.68 per index task against $0.84.

Are GLM-5.3 weights open source?

We could not confirm it. On August 22, 2026 the most recent public repository on the Hugging Face account that published GLM-5, GLM-5.1 and GLM-5.2 was still GLM-5.2, dated June 16, 2026, and a global search for GLM-5.3 returned only community repositories, the only GGUF among them containing just a .gitattributes file and a README. On ModelScope, the paths that serve GLM-5.1 and GLM-5.2 respond while the equivalent GLM-5.3 paths return not-found. That is the result of a search, not proof that no weights exist.

Why do Z.ai's release notes say state of the art among open-source models?

The sentence supports two readings and we are not choosing between them. It may mean GLM-5.3 is itself an open-source model released somewhere we did not look. It may equally mean GLM-5.3 leads the field of open-source models on those benchmarks, which is a statement about the competition rather than about GLM-5.3's own license. Artificial Analysis labels GLM-5.3 a proprietary model where it labels GLM-5.2 open weights, which is one independent data point but not a resolution.

What are Z.ai's cybersecurity claims for GLM-5.3?

There are three, attached to different things. The release notes report 2,436 vulnerabilities identified on real-world targets in collaboration with multiple cybersecurity teams, of which 1,097 are described as medium or high severity — that figure is not a CyberGym result. The model page separately claims the best performance to date on the CyberGym vulnerability discovery benchmark, and reports vulnerability exploitation scores more than twice GLM-5.2's. The release notes also state that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery. All three are Z.ai's own statements.

Does GLM-5.3 support images or video?

No. Z.ai's documentation states that GLM-5.3 currently supports text-only inputs. Vision runs through a separate line in the catalog, including GLM-5V-Turbo and GLM-4.6V, and this release does not change that split.

Which API endpoint should I call for GLM-5.3?

Z.ai documents three: an OpenAI Chat Completion endpoint at https://api.z.ai/api/coding/paas/v4, an OpenAI Response endpoint at https://api.z.ai/api/v1, and an Anthropic Messages endpoint at https://api.z.ai/api/anthropic. There is a restriction attached: anyone who has previously subscribed to a GLM Coding Plan, including an expired subscription, can currently access the model API only through the OpenAI Chat Completion-compatible protocol.

Should I upgrade from GLM-5.2 to GLM-5.3?

Move if you are already paying GLM-5.2 rates for agentic coding, since the same per-token price buys seven more index points at effectively unchanged throughput — 93.2 output tokens per second against 92.5. Wait in three cases: if you rely on a non-reasoning mode, which GLM-5.3 does not offer; if you self-host, since GLM-5.2 is the newest GLM release with weights we could find published; and if you need vision, since GLM-5.3 is text-only.

Sources

Every measurement attributed to Artificial Analysis was read on August 22, 2026 on version 4.1.1 of the Intelligence Index. Independent index scores are re-run over time and can move; the version of the index is stated wherever a score appears. Prices are vendor list prices read on the same date and exclude any subscription plan discount.

Related Articles

Was this review helpful?
Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.