Open Source / News

GLM-5.3: we tested Z.ai's open coding model, and who really runs it

GLM-5.3 from Z.ai is free to download and cheap to use. On our private coding test it was nearly perfect, but it thought for minutes per job. And it was not Z.ai that answered.

Table of our test: GLM-5.3 passed 182 of 186 secret checks and GLM-5.3-Flash 184, next to DeepSeek V4 Pro, Mistral Large 4, Claude Sonnet 5.5 and GPT-6.1 Sol
Our own test through OpenRouter, 6 and 7 October 2026. Chart by Not an AI App.

GLM-5.3 is the newest big AI model from Z.ai, whose service is run by a company in Singapore. You can download it, and it is popular with programmers. We tested GLM-5.3 and its small brother GLM-5.3-Flash on our new coding test, next to DeepSeek, Mistral, Claude and GPT. GLM-5.3 got 182 of 186 secret checks right, and Flash 184. But both thought for minutes per job, where Claude and GPT needed under half a minute.

What is GLM-5.3?

Screenshot of Z.ai's documentation: GLM-5.3 is its latest flagship model and uses the same base model as GLM-5.2, with all improvements from post-training
From Z.ai's model page. Screenshot from 7 October 2026, cropped.

GLM-5.3 came out in August 2026. On its model page, Z.ai says it "uses the same base model as GLM-5.2". All the gains come from extra training afterwards, mostly on long coding and agent tasks. Z.ai says it is 50% better than GLM-5.2 on its own coding test.

You can download both versions on Hugging Face. GLM-5.3-Flash uses the MIT licence: you can use it for almost anything. GLM-5.3 has its own licence. It is also free, but a company that sells AI as a service and earns more than $10 billion a year needs a security check by Z.ai first. That only applies to a handful of giant companies.

It always thinks, on the highest setting

Screenshot of Z.ai's documentation: reasoning cannot be turned off, and the reasoning effort setting is low, high or max, with max as the default
From Z.ai's model page. Screenshot from 7 October 2026, cropped.

GLM-5.3 always "thinks" before it answers: it writes hidden notes to itself first. You cannot turn that off. You can only choose how hard it thinks: low, high or max. The standard setting is max, Z.ai's documentation shows.

That makes it careful, but also slow and wordy. In our test, GLM-5.3 wrote about 32,000 tokens per job on average, most of it thinking. Claude Sonnet 5.5 wrote about 2,000 for the same jobs.

What does it cost?

Screenshot of Z.ai's price list: GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens; GLM-5.3-Flash $0.15 and $0.50
Z.ai's price list. Screenshot from 7 October 2026, cropped.

On Z.ai's price list, GLM-5.3 costs $1.40 per million tokens read and $4.40 per million written. GLM-5.3-Flash costs $0.15 and $0.50. That is cheaper per token than Claude Sonnet 5.5 or GPT-6.1 Sol, which cost $2 and $10.

There is also a "GLM Coding Plan", a subscription for coding tools, where use outside busy hours counts for half.

We tested through OpenRouter, a service that gives access to many models. There the price can differ, because different companies run the model, each with its own price. That is why two tries with almost the same number of tokens could cost 4 cents once and 9 cents another time.

Our test: three small coding jobs

Our hardest test job: reading CSV text with commas, line breaks and doubled quotes inside quoted fields, and naming the right line in error messages
Diagram by Not an AI App. The full secret tests stay private.

We used the same private test as for Claude Fable 5.1. The tests in our public test kit are known by now, so we keep these secret. There are three jobs:

  • D. Fix a money reader: code that turns "€ 1.234,56" into cents, with several bugs. 28 secret checks.
  • E. Split a restaurant bill: shared dishes and a tip, with cents that must add up exactly. 17 checks.
  • F. Read a CSV file: a spreadsheet saved as text, where commas, quotes and line breaks can hide inside a field. 17 checks.

Here is one check from job E. Two friends order dishes for 70 and 230 cents and add a 3% tip. The tip of 9 cents must then be split 2 and 7, not 3 and 6. Each model did each job three times, in one request through OpenRouter, with the default settings.

The results: careful, but slow

Chart of cost and time per job: GLM-5.3 about 11.0 cents and 4 minutes, GLM-5.3-Flash about 0.7 cents, Claude Sonnet 5.5 and GPT-6.1 Sol under half a minute
Our own test, prices as OpenRouter charged them. Chart by Not an AI App.

Here is how the models did, with all three jobs together:

  • GLM-5.3: 182 of 186 secret checks. In one try at job E, it split the tip wrong, so the amounts did not add up exactly. In one try at job F, it got empty fields wrong. About 4 minutes and 11.0 cents per job.
  • GLM-5.3-Flash: 184 of 186. It slipped on small details of job F twice. About 3 minutes and only 0.7 cents per job.
  • DeepSeek V4 Pro: 180 of 186, about 4 minutes per job.
  • Mistral Large 4: 181 of 186. It read "1234,5" wrong on job D every time.
  • Claude Sonnet 5.5 and GPT-6.1 Sol: every check right, in about 17 seconds and 24 seconds per job.

So both GLM models are close to the best. The small Flash even scored a little higher than the big GLM-5.3, for a fraction of the price.

Note the cost of the big GLM-5.3: per token it is cheaper than Claude Sonnet at Z.ai, but in our test it cost more per job. It wrote a lot, and OpenRouter charged more per written token than Z.ai's own price list. The price you pay is waiting time: minutes instead of seconds.

One more thing we noticed: Mistral Large 4 thought much longer than on its launch day. On 6 October it answered in seconds; on 7 October some answers took minutes, and one request got no answer within 15 minutes. We ran that one again.

Who actually ran GLM-5.3?

Chart of which companies ran GLM-5.3 and GLM-5.3-Flash for our requests through OpenRouter; Z.ai itself was not one of them
From the provider field in OpenRouter's answers. Chart by Not an AI App.

This surprised us. Because GLM-5.3 can be downloaded, many companies run it. When we asked OpenRouter for GLM-5.3, it sent our requests to Decart, Mistral, PrimeIntellect, SiliconFlow, Wafer. Yes, Mistral: the French AI company also runs other companies' open models. For Flash it used Near AI, OpenInference, Parasail, Relace, StreamLake. Z.ai itself answered none of them.

That matters for your data. If you use GLM through a service like OpenRouter, the company that runs it gets your text, and its rules apply. If you want Z.ai's own rules, use Z.ai's own service.

Where does your data go at Z.ai?

Screenshot of Z.ai's terms for its API: the company does not store the content customers or users send or generate; it is processed in real time and not saved
From Z.ai's data processing terms for its API. Screenshot from 7 October 2026, cropped.

Z.ai's privacy page has separate terms for its API. In them, Z.ai says it does "not store any of the content" that customers or their users send or get back. It is handled in real time and not saved on its servers.

The same terms say Z.ai generally runs its services from Singapore, so your data is "generally processed in Singapore". The company behind Z.ai is registered there too. For the free chat app, the general privacy policy applies, which works differently.

So with Z.ai's own API, your code is not kept, according to Z.ai. Through OpenRouter, it depends on which company answers.

How we checked

Diagram of our method in three steps: read Z.ai's pages, run three coding jobs three times on six models, check secret tests, time, cost and which company answered
Diagram by Not an AI App.

On 7 October 2026 we read Z.ai's model page, its price list, its privacy page and API terms, and the model pages on Hugging Face. The screenshots are our own captures of those pages, cropped only.

We ran the test on 6 and 7 October through OpenRouter, with default settings and three tries per job per model. A request with no answer within 15 minutes counted as an error; that happened once, with Mistral, and we ran it again. The costs are what OpenRouter charged. The owner of this site paid for the requests. The charts were made by us. None of the images was made by AI.

This was a small test with three short jobs. It says nothing about long projects, where Z.ai says GLM-5.3 is strongest. For other open models we tested, see Qwen 3.8.