How we test AI models on coding, and how you can run our test yourself
Our three coding tasks and the secret tests we use to score AI models, explained step by step. Download the test kit and run it on any AI yourself.
In many of our articles we write "we tested it". But what does that mean? Here is our test, in full. We give every AI model the same three small coding tasks. Then we check the work with secret tests: checks we wrote ourselves, which the AI never sees. You can download the test kit and try it on any AI yourself.
Task A: fix a bug
The first task is a small file of 28 lines. It turns a time like "1h30m" into seconds, and back. But it has bugs. For "1h30m" it gives 1,800 seconds instead of 5,400, because one line replaces the total instead of adding to it. It also reads "1.5h" as 5 hours, and it turns 0 seconds into nothing instead of "0s".
The model gets this message, word for word: "The tests in test/ fail. Fix src/duration.js so that it does what its comments promise, for all valid inputs, not only the cases in the existing tests."
The important words are "for all valid inputs". The two tests in the folder only check a few cases. A good fix also works for times the tests do not try.
Task B: build an invoice calculation
The second task starts with an empty function and a description of one page. The model has to write a calculation for a Dutch invoice, with all amounts in cents. The description has seven rules. A few of them:
- A quantity can have up to three decimals, like 1.5 hours.
- Amounts are rounded to whole cents. Half a cent goes up.
- VAT, the Dutch sales tax, is worked out on the total per tax rate, not on every line.
- Negative prices are allowed, for credit notes.
This looks easy, but it is full of traps. Rounding is where most mistakes in real invoice software happen.
Task C: add a feature to an app
The third task is a small to-do app that you use by typing commands. Its code is split over three files. The model has to add due dates: store them, sort the list by them, mark late tasks as OVERDUE, and refuse dates that do not exist, like 30 February. Everything that already worked must keep working.
This task needs a model that can run the app and check its own work. So we only gave it to the "agents": Codex with GPT-6.1 Sol and GPT-6 Astra, and Claude Code with three Claude models.
The secret tests
The secret tests are the heart of our test. They are small checks we wrote ourselves. The AI never sees them. We add them only after it has finished, and count how many pass.
For task A, for example, there is a list of inputs that must be refused. "1h banana" must be refused, because there is junk after the time. "-5m" must be refused, because a time cannot be negative. Claude Haiku 4.5 accepted both.
One test checks "1h1h", the same unit twice. The task does not say this is wrong, and every model so far has accepted it. We keep that test on purpose: it shows which models check more than they are asked. None did yet.
For task B, one test uses a quantity of 1.005. Mistral Large 4 refused that valid amount in two of its three tries, because of a tiny rounding error inside the computer.
Two ways we run a model
Not every model is used the same way, so we test in two ways:
- One request. We send the task and the files in one message through OpenRouter, a service that gives access to many models. The model answers with new code, and we save it. It cannot run anything itself. We used this for Grok, DeepSeek and Mistral, with tasks A and B, three tries each.
- An agent. The model works in a real folder on our computer, through Codex or Claude Code. It reads the files, changes them and runs the tests, like a programmer would. We used this for GPT-6.1 Sol, GPT-6 Astra and the Claude models, with all three tasks, two tries each.
Either way, the same secret tests decide the score. We use the default settings, with medium "thinking effort" for the agents, and a fresh copy of the folder for every try.
What we write down
For every try we save the answer, the score on the secret tests, the cost and the number of tokens. Tokens are the small pieces of text an AI reads and writes. We also note which company ran the model, and any errors.
We only report times when they are fair. In our Claude Code pricing test our computer was busy during part of the test, so we left the times out. And on Mistral's launch day, some requests failed the first time. We wrote that down too.
Run our test yourself
You can run the same test on any AI. Download the test kit. It has the three tasks, the secret tests, a short guide, and a small script that gives the score. You only need Node.js, a free program for running JavaScript.
1. Copy a task folder.
2. Give the task to your AI. The guide shows exactly what we sent.
3. Run the script: node check.mjs a-duration your-folder.
It shows how many secret tests passed, and which ones failed. Now that these secret tests are public, future models may have seen them. So for new articles we will make a new set.
What our test cannot tell you
Our test is small on purpose. It shows whether a model gets short, clear coding jobs right, where it slips on tricky details, and what one job costs.
It does not show how a model handles a big project with hundreds of files, a vague question, or a long conversation. It only uses JavaScript. And it does not say which model is "the best": most strong models pass almost everything here.
You can read the results in our articles on GPT-6.1 Sol and Astra, Grok 4.7, Codex vs Claude Code, DeepSeek V4, Claude Code pricing and Mistral Large 4.

