GPT-6.1 Sol vs GPT-6 Astra in Codex: we gave both the same three coding tasks
We gave GPT-6.1 Sol and GPT-6 Astra the same three coding tasks in Codex, twice each. Both passed the same hidden tests. Astra used about six times the credits.
OpenAI says its new GPT-6.1 Sol delivers "near-Astra performance" at a lower cost. In Codex, OpenAI's coding agent, that difference is counted in credits: per token, GPT-6 Astra uses five times as many. We gave both models the same three coding tasks in Codex, twice each, and graded the results with tests the models never saw. Both passed exactly the same tests. Astra used about six times as many credits.
What Sol and Astra are
OpenAI introduced GPT-6.1 Sol at its developer conference on 29 September 2026. Its model page describes it as "near-Astra performance for complex work at a lower cost". GPT-6 Astra is, in OpenAI's words, "our most capable model for the most demanding work".
On paper the two are close. Both have a context window of 1,050,000 tokens, a maximum output of 128,000 tokens and the same knowledge cutoff, 30 April 2026. The difference is the price. Through the API, Sol costs $2 per million input tokens and $10 per million output tokens; Astra costs $10 and $50.
In Codex with a ChatGPT subscription you do not pay per token. Instead, each model uses up your plan's allowance at its own rate, which OpenAI expresses in credits. That is what we wanted to measure: what do you get, and what does it use, when you pick one or the other for ordinary coding work?
How we tested
We wrote three small JavaScript projects, each with a short task file, the way you might hand a job to a colleague:
- A. Fix a bug. A function that turns text like "1h30m" into seconds added the units up wrongly. The task: make it do what its comments promise, for all valid inputs.
- B. Build from a spec. A one-page description of a Dutch invoice calculation with VAT, discounts, credit lines, EU reverse charge and rounding rules. The task: implement it and add tests.
- C. Add a feature across files. Due dates for a small to-do command-line tool, with validation, sorting, an overdue marker, a new command and exit codes.
Each model got each task twice, in a fresh copy of the project, through the Codex command-line tool (version 0.160.0) on Windows 11. Both ran at medium reasoning effort, the default for Sol, and the models took turns so that neither got a quieter moment on OpenAI's servers.
After each run we copied in a set of hidden tests: 24 for task A, 15 for B and 5 for C, the last ones running the whole tool. The models never saw these tests. Before the first run we checked every hidden test against our own reference solutions; all of them passed.
The results: the same score twelve times out of twelve
Both models scored the same on every task, in both runs:
| Task | GPT-6.1 Sol | GPT-6 Astra |
|---|---|---|
| A. Fix a bug | 23 of 24, twice | 23 of 24, twice |
| B. Build from a spec | 15 of 15, twice | 15 of 15, twice |
| C. Add a feature | 5 of 5, twice | 5 of 5, twice |
The one test that failed in task A failed for both models in every run, and that is on us. It checked that "1h1h", the same unit twice, is rejected. Our task file never said so, and both models reasonably accepted it as two hours. We count that test as unfair, which leaves both models with full marks.
The solutions themselves were remarkably alike. For the invoice task, both models avoided ordinary floating-point arithmetic and calculated with exact integers, the safe way to handle money. Their code had nearly the same structure, down to the helper functions.
We did find two small differences. In task A, Astra added a test file of its own in both runs, which the task did not ask for; Sol did not. And the comment "into whole seconds" was read in two ways, not by model but by run: each model once rounded "1.5s" up to 2 seconds and once dropped the half second. Our hidden tests did not cover that case, and both readings fit the comment. It is a reminder that a vague instruction gets a different answer from run to run, whichever model you pick.
What it used: about six times the credits
Codex reports how many tokens each run used. We converted those counts into credits with the rates on OpenAI's Codex pricing page: per million tokens, Astra costs 250 credits for input, 25 for cached input and 1,250 for output; Sol costs 50, 2.5 and 250.
| Task | Sol, average per run | Astra, average per run |
|---|---|---|
| A. Fix a bug | 1.5 credits | 9.8 credits |
| B. Build from a spec | 2.8 credits | 19.2 credits |
| C. Add a feature | 5.2 credits | 29.9 credits |
| All six runs | 19 credits | 118 credits |
Most of the gap comes from the rate per token. Astra did use somewhat more tokens over its six runs, about 17% more input and 6% more output, which turns the factor of five into roughly six. Time was not the difference either: Sol needed 954 seconds for its six runs, Astra 1,072. The slowest single run was Sol's: 368 seconds and 22 commands on task C, where its second attempt took 176 seconds.
Most input tokens in every run came from the instructions and tools our Codex installation loads at the start. Both models got exactly the same, so the comparison is fair, but your own numbers will depend on your setup. These are estimates; Codex's usage dashboard is the official count.
What that means on Plus and Pro
On a subscription, credits translate into how far your allowance stretches. OpenAI's own estimates for the Plus plan make the gap visible: in a five-hour period, roughly 5 to 45 local messages with Astra, against 15 to 160 with GPT-6.1 Sol. OpenAI stresses these are estimates, not fixed limits.
We ran our test on a ChatGPT Pro plan. According to the same page, "Pro plans currently have no five-hour limit", although weekly limits may apply. Our twelve runs fitted without any limit message. On Plus, with OpenAI's estimate of 5 to 45 Astra messages per five hours, six runs like ours could already use a large part of that window.
When the allowance runs out, Plus and Pro users can buy extra credits. Then the factor of five per token becomes a factor in what you pay.
When Astra could still be worth it
Our tasks were small and clearly described. That is a large share of everyday coding work, but it is not where OpenAI says Astra is strongest. OpenAI's model guide in Codex recommends Astra "for the hardest end-to-end work" and adds that it "is better at asking focused questions and incorporating your guidance". For repeated, long-running work "when cost matters", it points to Sol.
Our test does not contradict that, and it cannot confirm it either. Three tasks with two runs each is a small sample. Larger projects, vague requirements or work across code and documents could show differences that ours did not.
What our result does suggest: for well-defined coding tasks of this size, starting with Sol and switching to Astra only when Sol gets stuck is a reasonable default.
Who can use GPT-6.1 Sol in Codex
According to OpenAI's Codex model page, the GPT-6.1 Sol rollout "includes Plus, Pro, Business, Enterprise, and Edu in Codex in the desktop app and CLI". In Enterprise and Edu workspaces an administrator has to switch it on first. Free and Go plans are not included at launch.
In the Codex app and command-line tool you choose the model per task. On the command line that is the -m option, for example -m gpt-6.1-sol. This is also how we ran the test.
This follows our test of Claude Code mods, where we also kept to what we could run and measure ourselves.
How we checked
We ran the test on 4 October 2026 with the owner's ChatGPT Pro account, through the Codex command-line tool 0.160.0 on Windows 11, at medium reasoning effort, in a sandbox that allowed writing only inside the test folder. Each of the twelve runs started from a fresh copy of the project. We measured the time of each run, took the token counts from Codex's own output and graded the result with the hidden tests afterwards. The full results, the task files and the hidden tests are kept in our test folder.
That same day we read OpenAI's model pages for GPT-6.1 Sol and GPT-6 Astra, the API price list, and the Codex pricing and models pages. The screenshots are our own captures of those pages, cropped only. The charts and diagrams were made by us from our results file. None of the images is AI-generated.
This was a small test of two models on three tasks. It says something about well-defined everyday coding work, not about every kind of project. Credits are our estimate from published rates; OpenAI's usage dashboard is the official count.

