GLM-5.3: we tested Z.ai's open coding model, and who really runs it
GLM-5.3 from Z.ai is free to download and cheap to use. On our private coding test it was nearly perfect, but it thought for minutes per job. And it was not Z.ai that answered.
GLM-5.3 is the newest big AI model from Z.ai, whose service is run by a company in Singapore. You can download it, and it is popular with programmers. We tested GLM-5.3 and its small brother GLM-5.3-Flash on our new coding test, next to DeepSeek, Mistral, Claude and GPT. GLM-5.3 got 182 of 186 secret checks right, and Flash 184. But both thought for minutes per job, where Claude and GPT needed under half a minute.
What is GLM-5.3?
GLM-5.3 came out in August 2026. On its model page, Z.ai says it "uses the same base model as GLM-5.2". All the gains come from extra training afterwards, mostly on long coding and agent tasks. Z.ai says it is 50% better than GLM-5.2 on its own coding test.
You can download both versions on Hugging Face. GLM-5.3-Flash uses the MIT licence: you can use it for almost anything. GLM-5.3 has its own licence. It is also free, but a company that sells AI as a service and earns more than $10 billion a year needs a security check by Z.ai first. That only applies to a handful of giant companies.
It always thinks, on the highest setting
GLM-5.3 always "thinks" before it answers: it writes hidden notes to itself first. You cannot turn that off. You can only choose how hard it thinks: low, high or max. The standard setting is max, Z.ai's documentation shows.
That makes it careful, but also slow and wordy. In our test, GLM-5.3 wrote about 32,000 tokens per job on average, most of it thinking. Claude Sonnet 5.5 wrote about 2,000 for the same jobs.
What does it cost?
On Z.ai's price list, GLM-5.3 costs $1.40 per million tokens read and $4.40 per million written. GLM-5.3-Flash costs $0.15 and $0.50. That is cheaper per token than Claude Sonnet 5.5 or GPT-6.1 Sol, which cost $2 and $10.
There is also a "GLM Coding Plan", a subscription for coding tools, where use outside busy hours counts for half.
We tested through OpenRouter, a service that gives access to many models. There the price can differ, because different companies run the model, each with its own price. That is why two tries with almost the same number of tokens could cost 4 cents once and 9 cents another time.
Our test: three small coding jobs
We used the same private test as for Claude Fable 5.1. The tests in our public test kit are known by now, so we keep these secret. There are three jobs:
- D. Fix a money reader: code that turns "€ 1.234,56" into cents, with several bugs. 28 secret checks.
- E. Split a restaurant bill: shared dishes and a tip, with cents that must add up exactly. 17 checks.
- F. Read a CSV file: a spreadsheet saved as text, where commas, quotes and line breaks can hide inside a field. 17 checks.
Here is one check from job E. Two friends order dishes for 70 and 230 cents and add a 3% tip. The tip of 9 cents must then be split 2 and 7, not 3 and 6. Each model did each job three times, in one request through OpenRouter, with the default settings.
The results: careful, but slow
Here is how the models did, with all three jobs together:
- GLM-5.3: 182 of 186 secret checks. In one try at job E, it split the tip wrong, so the amounts did not add up exactly. In one try at job F, it got empty fields wrong. About 4 minutes and 11.0 cents per job.
- GLM-5.3-Flash: 184 of 186. It slipped on small details of job F twice. About 3 minutes and only 0.7 cents per job.
- DeepSeek V4 Pro: 180 of 186, about 4 minutes per job.
- Mistral Large 4: 181 of 186. It read "1234,5" wrong on job D every time.
- Claude Sonnet 5.5 and GPT-6.1 Sol: every check right, in about 17 seconds and 24 seconds per job.
So both GLM models are close to the best. The small Flash even scored a little higher than the big GLM-5.3, for a fraction of the price.
Note the cost of the big GLM-5.3: per token it is cheaper than Claude Sonnet at Z.ai, but in our test it cost more per job. It wrote a lot, and OpenRouter charged more per written token than Z.ai's own price list. The price you pay is waiting time: minutes instead of seconds.
One more thing we noticed: Mistral Large 4 thought much longer than on its launch day. On 6 October it answered in seconds; on 7 October some answers took minutes, and one request got no answer within 15 minutes. We ran that one again.
Who actually ran GLM-5.3?
This surprised us. Because GLM-5.3 can be downloaded, many companies run it. When we asked OpenRouter for GLM-5.3, it sent our requests to Decart, Mistral, PrimeIntellect, SiliconFlow, Wafer. Yes, Mistral: the French AI company also runs other companies' open models. For Flash it used Near AI, OpenInference, Parasail, Relace, StreamLake. Z.ai itself answered none of them.
That matters for your data. If you use GLM through a service like OpenRouter, the company that runs it gets your text, and its rules apply. If you want Z.ai's own rules, use Z.ai's own service.
Where does your data go at Z.ai?
Z.ai's privacy page has separate terms for its API. In them, Z.ai says it does "not store any of the content" that customers or their users send or get back. It is handled in real time and not saved on its servers.
The same terms say Z.ai generally runs its services from Singapore, so your data is "generally processed in Singapore". The company behind Z.ai is registered there too. For the free chat app, the general privacy policy applies, which works differently.
So with Z.ai's own API, your code is not kept, according to Z.ai. Through OpenRouter, it depends on which company answers.
How we checked
On 7 October 2026 we read Z.ai's model page, its price list, its privacy page and API terms, and the model pages on Hugging Face. The screenshots are our own captures of those pages, cropped only.
We ran the test on 6 and 7 October through OpenRouter, with default settings and three tries per job per model. A request with no answer within 15 minutes counted as an error; that happened once, with Mistral, and we ran it again. The costs are what OpenRouter charged. The owner of this site paid for the requests. The charts were made by us. None of the images was made by AI.
This was a small test with three short jobs. It says nothing about long projects, where Z.ai says GLM-5.3 is strongest. For other open models we tested, see Qwen 3.8.

