OpenAI / News

GPT-6.1 Sol vs GPT-6 Astra in Codex: we gave both the same three coding tasks

We gave GPT-6.1 Sol and GPT-6 Astra the same three coding tasks in Codex, twice each. Both passed the same hidden tests. Astra used about six times the credits.

Bar chart of average credits per run: fixing a bug 1.5 for GPT-6.1 Sol and 9.8 for GPT-6 Astra, building from a spec 2.8 and 19.2, adding a feature 5.2 and 29.9
Chart by Not an AI App from our own test of 4 October 2026. Credits are estimated from the token counts Codex reported and OpenAI's published credit rates.

OpenAI says its new GPT-6.1 Sol delivers "near-Astra performance" at a lower cost. In Codex, OpenAI's coding agent, that difference is counted in credits: per token, GPT-6 Astra uses five times as many. We gave both models the same three coding tasks in Codex, twice each, and graded the results with tests the models never saw. Both passed exactly the same tests. Astra used about six times as many credits.

What Sol and Astra are

Screenshot of OpenAI's model page for GPT-6.1 Sol: near-Astra performance for complex work at a lower cost, $2 input and $10 output per million tokens
OpenAI's API page for GPT-6.1 Sol. Screenshot taken on 4 October 2026, cropped.

OpenAI introduced GPT-6.1 Sol at its developer conference on 29 September 2026. Its model page describes it as "near-Astra performance for complex work at a lower cost". GPT-6 Astra is, in OpenAI's words, "our most capable model for the most demanding work".

On paper the two are close. Both have a context window of 1,050,000 tokens, a maximum output of 128,000 tokens and the same knowledge cutoff, 30 April 2026. The difference is the price. Through the API, Sol costs $2 per million input tokens and $10 per million output tokens; Astra costs $10 and $50.

In Codex with a ChatGPT subscription you do not pay per token. Instead, each model uses up your plan's allowance at its own rate, which OpenAI expresses in credits. That is what we wanted to measure: what do you get, and what does it use, when you pick one or the other for ordinary coding work?

How we tested

Diagram of our test in three steps: three small projects, Codex runs each twice per model with the same settings, and we grade afterwards with hidden tests
Diagram by Not an AI App: how we tested.

We wrote three small JavaScript projects, each with a short task file, the way you might hand a job to a colleague:

  • A. Fix a bug. A function that turns text like "1h30m" into seconds added the units up wrongly. The task: make it do what its comments promise, for all valid inputs.
  • B. Build from a spec. A one-page description of a Dutch invoice calculation with VAT, discounts, credit lines, EU reverse charge and rounding rules. The task: implement it and add tests.
  • C. Add a feature across files. Due dates for a small to-do command-line tool, with validation, sorting, an overdue marker, a new command and exit codes.

Each model got each task twice, in a fresh copy of the project, through the Codex command-line tool (version 0.160.0) on Windows 11. Both ran at medium reasoning effort, the default for Sol, and the models took turns so that neither got a quieter moment on OpenAI's servers.

After each run we copied in a set of hidden tests: 24 for task A, 15 for B and 5 for C, the last ones running the whole tool. The models never saw these tests. Before the first run we checked every hidden test against our own reference solutions; all of them passed.

The results: the same score twelve times out of twelve

Table of twelve runs: both models passed 23 of 24 hidden tests on task A, 15 of 15 on task B and 5 of 5 on task C, with times between 72 and 368 seconds
All twelve runs from our own test file. Hidden tests were copied in only after each run had finished.

Both models scored the same on every task, in both runs:

TaskGPT-6.1 SolGPT-6 Astra
A. Fix a bug23 of 24, twice23 of 24, twice
B. Build from a spec15 of 15, twice15 of 15, twice
C. Add a feature5 of 5, twice5 of 5, twice

The one test that failed in task A failed for both models in every run, and that is on us. It checked that "1h1h", the same unit twice, is rejected. Our task file never said so, and both models reasonably accepted it as two hours. We count that test as unfair, which leaves both models with full marks.

The solutions themselves were remarkably alike. For the invoice task, both models avoided ordinary floating-point arithmetic and calculated with exact integers, the safe way to handle money. Their code had nearly the same structure, down to the helper functions.

We did find two small differences. In task A, Astra added a test file of its own in both runs, which the task did not ask for; Sol did not. And the comment "into whole seconds" was read in two ways, not by model but by run: each model once rounded "1.5s" up to 2 seconds and once dropped the half second. Our hidden tests did not cover that case, and both readings fit the comment. It is a reminder that a vague instruction gets a different answer from run to run, whichever model you pick.

What it used: about six times the credits

Screenshot of OpenAI's Codex pricing page: GPT-6 Astra costs 250 credits per million input tokens and 1,250 per million output tokens, GPT-6.1 Sol 50 and 250
The credit rates on OpenAI's Codex pricing page. Screenshot taken on 4 October 2026, cropped.

Codex reports how many tokens each run used. We converted those counts into credits with the rates on OpenAI's Codex pricing page: per million tokens, Astra costs 250 credits for input, 25 for cached input and 1,250 for output; Sol costs 50, 2.5 and 250.

TaskSol, average per runAstra, average per run
A. Fix a bug1.5 credits9.8 credits
B. Build from a spec2.8 credits19.2 credits
C. Add a feature5.2 credits29.9 credits
All six runs19 credits118 credits

Most of the gap comes from the rate per token. Astra did use somewhat more tokens over its six runs, about 17% more input and 6% more output, which turns the factor of five into roughly six. Time was not the difference either: Sol needed 954 seconds for its six runs, Astra 1,072. The slowest single run was Sol's: 368 seconds and 22 commands on task C, where its second attempt took 176 seconds.

Most input tokens in every run came from the instructions and tools our Codex installation loads at the start. Both models got exactly the same, so the comparison is fair, but your own numbers will depend on your setup. These are estimates; Codex's usage dashboard is the official count.

What that means on Plus and Pro

Screenshot of OpenAI's usage estimates: on Plus, 5 to 45 local messages per five hours with GPT-6 Astra and 15 to 160 with GPT-6.1 Sol; Pro has no five-hour limit
OpenAI's own estimates of how many local messages fit in five hours on Plus. Screenshot taken on 4 October 2026, cropped.

On a subscription, credits translate into how far your allowance stretches. OpenAI's own estimates for the Plus plan make the gap visible: in a five-hour period, roughly 5 to 45 local messages with Astra, against 15 to 160 with GPT-6.1 Sol. OpenAI stresses these are estimates, not fixed limits.

We ran our test on a ChatGPT Pro plan. According to the same page, "Pro plans currently have no five-hour limit", although weekly limits may apply. Our twelve runs fitted without any limit message. On Plus, with OpenAI's estimate of 5 to 45 Astra messages per five hours, six runs like ours could already use a large part of that window.

When the allowance runs out, Plus and Pro users can buy extra credits. Then the factor of five per token becomes a factor in what you pay.

When Astra could still be worth it

Screenshot of OpenAI's model page for GPT-6 Astra: our most capable model for the most demanding work, $10 input and $50 output per million tokens
OpenAI's API page for GPT-6 Astra. Screenshot taken on 4 October 2026, cropped.

Our tasks were small and clearly described. That is a large share of everyday coding work, but it is not where OpenAI says Astra is strongest. OpenAI's model guide in Codex recommends Astra "for the hardest end-to-end work" and adds that it "is better at asking focused questions and incorporating your guidance". For repeated, long-running work "when cost matters", it points to Sol.

Our test does not contradict that, and it cannot confirm it either. Three tasks with two runs each is a small sample. Larger projects, vague requirements or work across code and documents could show differences that ours did not.

What our result does suggest: for well-defined coding tasks of this size, starting with Sol and switching to Astra only when Sol gets stuck is a reasonable default.

Who can use GPT-6.1 Sol in Codex

Screenshot of OpenAI's Codex models page: the GPT-6.1 Sol rollout includes Plus, Pro, Business, Enterprise and Edu in Codex, not Free and Go
Who gets GPT-6.1 Sol in Codex, according to OpenAI. Screenshot taken on 4 October 2026, cropped.

According to OpenAI's Codex model page, the GPT-6.1 Sol rollout "includes Plus, Pro, Business, Enterprise, and Edu in Codex in the desktop app and CLI". In Enterprise and Edu workspaces an administrator has to switch it on first. Free and Go plans are not included at launch.

In the Codex app and command-line tool you choose the model per task. On the command line that is the -m option, for example -m gpt-6.1-sol. This is also how we ran the test.

This follows our test of Claude Code mods, where we also kept to what we could run and measure ourselves.

How we checked

We ran the test on 4 October 2026 with the owner's ChatGPT Pro account, through the Codex command-line tool 0.160.0 on Windows 11, at medium reasoning effort, in a sandbox that allowed writing only inside the test folder. Each of the twelve runs started from a fresh copy of the project. We measured the time of each run, took the token counts from Codex's own output and graded the result with the hidden tests afterwards. The full results, the task files and the hidden tests are kept in our test folder.

That same day we read OpenAI's model pages for GPT-6.1 Sol and GPT-6 Astra, the API price list, and the Codex pricing and models pages. The screenshots are our own captures of those pages, cropped only. The charts and diagrams were made by us from our results file. None of the images is AI-generated.

This was a small test of two models on three tasks. It says something about well-defined everyday coding work, not about every kind of project. Credits are our estimate from published rates; OpenAI's usage dashboard is the official count.