Anthropic / News

Codex vs Claude Code: we gave both the same three coding tasks

We gave Codex and Claude Code the same three programming jobs. Both did equally well. Claude Code was four times faster; Codex cost half as much for developers.

Four facts from our test: the same score on our secret tests, 16 minutes against 4, $0.76 against $1.67 at developer prices, and the same price per word
Chart by Not an AI App from our own test on 4 and 5 October 2026: Codex with GPT-6.1 Sol against Claude Code with Sonnet 5.5.

Many people ask which AI coding helper is better: OpenAI's Codex or Anthropic's Claude Code. So we gave both the same three programming jobs. Both did the work equally well. Claude Code was about four times faster. But for developers who pay per use, it would have cost about twice as much.

What we compared

Screenshot of Anthropic's price list: Claude Sonnet 5.5 costs $2 per million tokens in and $10 per million tokens out
Anthropic's price list for Claude Sonnet 5.5. Screenshot from 5 October 2026, cropped.

Codex and Claude Code are AI helpers for programmers. You give them a job. They read the code, try things out and change files until the job is done.

Each tool can use different AI models. We picked two models that cost exactly the same: GPT-6.1 Sol in Codex and Claude Sonnet 5.5 in Claude Code. For developers, both cost $2 per million tokens of text in and $10 per million tokens out. Tokens are the small pieces of text an AI reads and writes. You can check the prices on OpenAI's page and Anthropic's page.

Because the price per token is the same, any difference comes from the tools themselves. The Codex results are the same ones we used in our test of GPT-6.1 Sol against GPT-6 Astra.

How we tested

Diagram of our test: the same three small projects, each tool twice, and the same secret tests run after each job
Diagram by Not an AI App: how we tested.

We made three small programming projects. Each came with a short note saying what to do:

  • A. Fix a bug. A piece of code that turns "1h30m" into seconds gave wrong answers.
  • B. Build something from a description. A one-page description of a Dutch invoice, with tax, discounts and rounding.
  • C. Add a new feature. Due dates for a small to-do list program.

Each tool did each job twice, starting from a clean copy each time. Both used the same middle setting for how hard the AI thinks.

Then came our secret tests. These are small checks we wrote ourselves, such as "does 1h30m give 5,400 seconds?". The AI never saw them. We only ran them after it had finished. That way the AI could not aim for our checks. We also tried every secret test on our own answers first, to make sure the tests were right.

Who did the better job? Neither

Table of twelve jobs: Codex and Claude Code passed exactly the same secret tests on all three tasks, both times
All twelve jobs. We ran our secret tests only after each job was done.

Both tools passed exactly the same secret tests, in every try:

JobCodexClaude Code
A. Fix a bug23 of 24, twice23 of 24, twice
B. Build from a description15 of 15, twice15 of 15, twice
C. Add a feature5 of 5, twice5 of 5, twice

One test in job A failed for both tools every time. That was our mistake. It checked a rule we never wrote in the job note. So we do not count it against either tool.

Both tools also wrote their own tests, as the job notes asked. In two tries, Claude Code wanted to use a program we had not allowed. It was stopped, and it finished the job another way. That did not cost it any points.

Which one was faster? Claude Code, by a lot

Bar chart of time per task: fixing a bug 76 against 23 seconds, building from a description 129 against 40, adding a feature 272 against 42, Codex against Claude Code
Chart by Not an AI App: average of two jobs per task.

Speed was the big difference:

JobCodex, averageClaude Code, average
A. Fix a bug76 seconds23 seconds
B. Build from a description129 seconds40 seconds
C. Add a feature272 seconds42 seconds
All six triesabout 16 minutesabout 4 minutes

Claude Code was faster on every job, every time. The gap was biggest on job C. One Codex try took more than six minutes there. The other took three.

For one small job, that hardly matters. If you do dozens of jobs a day, it adds up.

Which one costs less? Codex, for developers

Bar chart of what one job would cost at developer prices: Claude Code costs more on every task
Chart by Not an AI App: our estimate from how much text each tool read and wrote.

Most people use these tools with a monthly subscription. Then you do not pay per job. But developers can also pay per use, through the API, the paid connection for programmers. So we worked out what our jobs would cost that way:

JobCodex, per tryClaude Code, per try
A. Fix a bug$0.06$0.24
B. Build from a description$0.11$0.29
C. Add a feature$0.21$0.31
All six tries$0.76$1.67

How can that be, if the price per token is the same? It is about how the tools work. Claude Code starts every job by saving a long set of instructions, so it does not have to send them again and again. Saving costs extra: twice the normal price. Codex also reuses a lot of text, but more cheaply.

We did this sum ourselves. Claude Code's own estimate said "unknown", because this version did not know Sonnet 5.5's price yet.

What if you use a subscription?

Screenshot of Anthropic's prices for saved instructions: saving them for an hour costs twice the normal price, reading them back one tenth
Anthropic's prices for saving and reusing instructions. Screenshot from 5 October 2026, cropped.

With a ChatGPT or Claude subscription you pay a fixed amount each month. What matters then is how fast you use up your allowance.

For Codex, OpenAI publishes how much each model uses. By that measure, our six Codex tries used about 19 of the "credits" that make up a plan's allowance. We did not find a similar list for Claude Code, so we cannot compare this part.

Our advice: for clear, everyday programming jobs, both tools did the same work. If you are short on time, Claude Code was faster. If you pay per use, Codex was cheaper.

How we checked

List of all twelve jobs with tool, task, seconds and secret test score
Our record of every job, set in a frame by us.

We ran the Codex tries on 4 October 2026 and the Claude Code tries on 5 October 2026, on the same Windows laptop, with the owner's own subscriptions. We used Codex version 0.160.0 with GPT-6.1 Sol, and Claude Code version 2.1.280 with Claude Sonnet 5.5. We chose Sonnet 5.5 by name, because this version of Claude Code still picked the older Sonnet 5 by default. We timed every try, wrote down how much text each tool used, and ran the same secret tests afterwards.

The costs are our own estimates. We used the token counts each tool reported and the prices on Anthropic's and OpenAI's pages from 5 October 2026. Codex did not report one small extra charge, so its cost may be a little higher. The screenshots are our own. The charts and the record of every job were made by us from our results. None of the images was made by AI.

This was a small test: three jobs, two tries per tool. Bigger projects may give different results. Want to know more about Claude Code? Read our test of Claude Code mods.