Qwen 3.8: we tested Alibaba's open AI models against Claude and GPT
Qwen 3.8 is Alibaba's newest AI family, and most of it is free to download. On our new coding test it was almost always right, but it took 8 to 12 minutes per answer.
Qwen is the family of AI models from Alibaba, the big Chinese tech company. Qwen 3.8 is its newest one, and you can download most of it for free. People search for it a lot. So we tested four versions of Qwen 3.8 on a new coding test, next to Claude, GPT, DeepSeek and Mistral. Qwen 3.8 got almost every answer right. But it was slow: on average it took 8 to 12 minutes per answer, where Claude and GPT needed less than half a minute.
What is Qwen 3.8?
Alibaba launched Qwen 3.8-Max on 3 August 2026 and called it "the most capable model in the Qwen family to date", in its announcement. It has 2.4 trillion parameters. Parameters are the numbers a model learns during training. For each word, about 95 billion of them do the work.
It was the first time Alibaba gave away the files of its biggest model. They appeared on Hugging Face five days later, as Qwen3.8 2.4T. A much smaller version, Qwen3.8 27B, came out around the same time. It has been downloaded more than six million times.
Which Qwen 3.8 is which?
The names can be confusing, so here they are side by side:
- Qwen3.8 Max: the big model on Alibaba's own service.
- Qwen3.8 2.4T: the same big model as a download. Other companies run it too.
- Qwen3.8 27B: a much smaller model you can download. It uses the Apache 2.0 licence, which lets almost anyone use it for almost anything.
- Qwen3.8 Flash: a small, very cheap model on Alibaba's service.
- Max Prime: sold on OpenRouter as a faster Max, for twice the price. Alibaba's page on its fast mode promises 1.5 to 2 times the speed.
What does it cost?
On Alibaba's price list, Max costs $2 per million tokens read and $6 per million written, in its international region. Flash costs $0.15 and $0.47. New users get one million free tokens for Max, valid for 90 days, but only in the Singapore region.
That makes Max cheaper than GPT-6.1 Sol and Claude Sonnet 5.5, which cost $2 and $10. But in our test, price per token was not the whole story, as you will see. Qwen thought for so long that it wrote many more tokens per answer.
Our new test: two small jobs
Our old test is now public, so new models may have seen its answers. For this article we wrote a new one, and we keep its secret tests private. It has two jobs:
- D. Fix a money reader. A short piece of code should turn Dutch amounts like "€ 1.234,56" into cents, and back. It has several bugs. The model has to fix it so it works "for all valid inputs".
- E. Split a restaurant bill. From a one-page description, the model writes code that splits a bill between friends. Shared dishes are split evenly, a tip is added, and the cents must add up exactly.
Then we ran our secret tests: 28 checks for job D and 17 for job E. A few examples:
- "€ 0,5" must give 50 cents.
- "1.23,45" must be refused, because the dots are in the wrong place.
- A €10 bottle of wine for three people must come to 334, 333 and 333 cents.
Before the test, we checked that our own answers pass every one.
Each model did each job three times, in one request through OpenRouter, a service that gives access to many models. We used the default settings. If a model gave no answer within 15 minutes, we counted it as an error.
The results: right, but slow
All eight models got most of it right. The big differences were in waiting time and reliability:
- Claude Sonnet 5.5 and GPT-6.1 Sol passed every secret check, every time. They took 15 seconds and 24 seconds per answer on average. All six tries together cost 14 cents and 7 cents.
- Qwen3.8 Flash also passed everything, for 12 cents in total. But it thought for about 11 minutes per answer.
- Qwen3.8 2.4T passed everything too, but took about 8 minutes per answer and cost $1.32, the most of all.
- Qwen3.8 Max passed job E twice. Once, on job E, it gave no usable answer, even when we tried again. On job D it once got "12" and "€1.000" wrong. To be fair: a plain amount without cents, like "12", is not one of the examples in the task.
- Qwen3.8 27B gave no usable answer in three of its six tries. Once it ran out of room while still thinking, once it sent back an empty answer, and once it failed twice in a row: first a broken answer, then no answer within 15 minutes.
For comparison: DeepSeek V4 Pro slipped on job D in two tries, and Mistral Large 4 read "1234,5" wrong every time, even though that exact example is in the task.
Our advice: Qwen 3.8 is good at getting things right, but it thinks a very long time. For work where you wait for the answer, that is a problem. For jobs that can run in the background, it can be a cheap choice, especially Flash.
Where does your data go?
Alibaba has several regions for its AI service, including Singapore, Frankfurt and Virginia. Its documentation explains two things you choose:
- The region decides where your data is stored.
- The "deployment scope" decides where the model does its work. "Global" means anywhere in the world. Some regions also offer a local scope, such as "EU" in Frankfurt.
Alibaba says it sends the data encrypted, and your stored data stays in the region you picked. If you use Qwen through another company, such as the ones that answered our requests through OpenRouter, that company's rules apply.
Using Qwen 3.8 from Europe
For Europeans this is the important part. In Alibaba's price list for Frankfurt, Qwen3.8-Max is only listed with the "Global" scope, at $1.65 and $4.95 per million tokens. Only the older Qwen3-Max is listed with an "EU" scope.
So your data can be stored in Germany, while Qwen 3.8 Max does its work somewhere else in the world. If that matters for your work, you can run the open 2.4T or 27B model yourself, or use a company that runs it in Europe.
Can you use the free download for anything?
The small 27B model uses Apache 2.0, a well-known open licence: you can use it for almost anything.
The big 2.4T model has its own licence. It is free too, with two exceptions. A product with more than 100 million users a month, or more than $20 million revenue a month, must show the model's name. And a company that sells AI as a service, with more than $50 million revenue a year, needs a separate licence from Alibaba. For nearly everyone else, it is free to use.
How we checked
On 6 October 2026 we read Alibaba's announcement, its price list, its pages on regions and fast mode, and the model pages and licence on Hugging Face. The screenshots are our own captures of those pages, cropped only. For the price table we zoomed the page out, so the whole table fits.
We ran the test on 6 and 7 October through OpenRouter, with default settings, three tries per job per model. Four requests came back broken, all from Qwen models; we ran those again. A try that still failed counts as zero. One extra Qwen Max try ran by mistake; we left it out. The test cost about $3.25 that OpenRouter reported to us; the failed requests may have cost a little more. The owner of this site paid for it. The charts were made by us. None of the images was made by AI.
This was a small test with two clear jobs. Speeds can change from day to day. Earlier tests with the same kind of jobs: DeepSeek V4 and Mistral Large 4.

