September 30, 2026 • 9 min read

Claude Sonnet 5.5 vs GPT-6.1 Sol: Same $2 Price, Very Different Bill. Which Should Your Business Use?

Anthropic released Claude Sonnet 5.5 on 28 September. The next day, OpenAI released GPT-6.1 Sol. Both are the mid-priced "workhorse" model of their family, and both have exactly the same list price: $2 per million tokens in and $10 per million tokens out. Most comparisons will stop there and call it a tie. It is not a tie. When an independent tester ran both through the same set of tasks, one of them reached the same score for roughly a quarter of the cost per task. This article explains why the same price produces a very different bill, where each model is the better choice, and how to test them on your own work before you switch.

What Anthropic and OpenAI released this week

  • Claude Sonnet 5.5 (Anthropic, 28 September). Anthropic calls it "a clear upgrade over Claude Sonnet 5" that runs more than 30% faster and "costs up to 30% less for most work" than Sonnet 5, because it needs fewer tokens to finish a job. It sits below Opus 5.5 and is aimed at well-scoped everyday tasks, bug fixes, and polished documents, slides and spreadsheets. It is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure.
  • GPT-6.1 Sol (OpenAI, 29 September). OpenAI says it "nearly matches GPT-6 Astra's intelligence" on coding, computer use and professional work "at one-fifth of Astra's standard input and output token prices." It went live in the API, and in ChatGPT Work and Codex for Plus and higher ChatGPT plans. OpenAI says it is not yet in the normal ChatGPT chat.
  • What is coming next. Anthropic says a cheaper Claude Haiku 5.5 will join the family "in the coming weeks." OpenAI says a faster Ultrafast version of GPT-6.1 Sol is coming "in the coming days."

The list prices are the same. Almost.

These are the standard prices each company published, per million tokens. A token is roughly three quarters of a word.

Per 1M tokensClaude Sonnet 5.5GPT-6.1 Sol
Input$2.00$2.00
Output$10.00$10.00
Cached input (cache read)$0.20$0.10
Cache write$2.50$2.50

The one line that differs is cached input, and it matters more than it looks. Any AI feature built into your software, whether a support bot, an invoice reader or a call summariser, sends the same instructions and background with every request. That repeated part is what gets cached. For an agent that re-reads the same long context on every step, it can be a large part of the bill, and GPT-6.1 Sol charges half as much for it.

But the bigger difference is not on the price list at all. It is how many tokens each model burns to finish the same job.

The same score, a very different bill

Artificial Analysis is an independent company that runs every major model through the same set of tests and publishes two numbers: an Intelligence Index score, and the average cost of completing its test tasks at the model's real prices. Both models can be run at different "effort" levels. More effort means more thinking, a higher score, and a higher bill.

Here is what it takes for each model to reach a score of 52 on that index, from the leaderboard on 30 September.

Cost per task to reach an Intelligence Index score of 52
Artificial Analysis leaderboard, 30 September 2026. Average USD per index task.
GPT-6.1 Sol (max effort)$0.72
GPT-6 Astra (xhigh effort)$2.31
Claude Sonnet 5.5 (xhigh effort)$2.74
Same score, same list price for Sol and Sonnet, and Sonnet costs about 3.8 times as much per task.

The full picture across effort levels is below, with each model named exactly as Artificial Analysis lists it. Three things stand out.

Model and effortIndex scoreCost per task
Claude Opus 5.5 (max, with fallback)58$5.98
Claude Sonnet 5.5 (max, with fallback)56$7.60
GPT-6 Astra (max)53$3.26
Claude Sonnet 5.5 (xhigh, with fallback)52$2.74
GPT-6.1 Sol (max)52$0.72
GPT-6.1 Sol (xhigh)51$0.39
GPT-6.1 Sol (high)50$0.32
GPT-6.1 Sol (medium)48$0.21
Claude Sonnet 5.5 (high, with fallback)47$1.08
Claude Sonnet 5.5 (medium, with fallback)41$0.59
  • Sonnet 5.5 reaches higher. At maximum effort it scores 56, second only to Opus 5.5 and above GPT-6 Astra's best score of 53. GPT-6.1 Sol tops out at 52. If you need the best answer a mid-priced model can give, Sonnet 5.5 is the one.
  • Sol is far cheaper for good-enough work. Sol at high effort scores 50 for $0.32 a task. Sonnet 5.5 at high effort scores 47 for $1.08. In our view, a score in the high 40s is plenty for most routine business jobs, and Sol gets there for less than a third of the cost.
  • Sonnet at full effort costs more than Opus. Sonnet 5.5 at max effort costs $7.60 a task, more than Opus 5.5 at max ($5.98), for a lower score. If you are going to turn Sonnet all the way up, you should use Opus instead.
A model's price per token tells you what one word costs. What you actually pay depends on how many words it uses to finish the job.

Why the same price gives a different bill

Modern models "think" before they answer, and that thinking is billed as output tokens at $10 per million. Two models with the same price can use very different amounts of thinking to reach the same answer. Artificial Analysis's cost per task captures this, which is why it is a better guide than the price list.

Anthropic's own launch material makes the same point in its favour when it compares Sonnet 5.5 with the older Sonnet 5. One customer, Balyasny Asset Management, told Anthropic that Sonnet 5.5 used about 121,000 tokens per answer on its finance tasks, where Sonnet 5 used 497,000. That is a real improvement over Sonnet 5. It does not tell you how Sonnet 5.5 compares with a rival model, and the independent numbers above suggest Sol is more economical still.

Speed is where Sonnet 5.5 wins. At medium effort, which Anthropic says is the default in the Claude apps, Artificial Analysis measured Sonnet 5.5 starting its answer in about 1.3 seconds, against about 5.7 seconds for GPT-6.1 Sol at the same setting. Sonnet also wrote faster once it started, at 89 tokens a second against 59. For a chatbot where a customer is waiting, that first delay is what people feel. The trade-off is quality and cost: at medium effort, Sol scored 48 for $0.21 a task, and Sonnet scored 41 for $0.59.

Where Sonnet 5.5 is still the better choice

Cheapest per task is not the same as best for your job. Anthropic's benchmarks are its own and should be read that way, but several of its customers describe specific gains that fit particular kinds of work.

  • Documents, slides and design. Anthropic says Sonnet 5.5 is strongest at "polished documents, slides, and spreadsheets." In one internal test, two experts judged a 10-slide operating review it drafted from earnings materials "ready to send as is." If your team produces client-facing documents, test Sonnet first.
  • Support tickets. Zendesk told Anthropic that across hundreds of real support cases, Sonnet 5.5 "made fewer wrong decisions" and processed tickets 20% faster than the Claude models it uses in production today.
  • You already use Claude. Moving an existing Claude setup to another vendor means rewriting prompts and retesting everything. Upgrading from Sonnet 5 to 5.5 costs the same per token and, by Anthropic's figures, less per task. One catch for developers: if you run Sonnet with thinking off, Anthropic says you must switch to a new "between_tools" setting before you move to 5.5.
  • Replies need to feel instant. For live chat and other work where a person is waiting, the medium-effort speed figures above favour Sonnet 5.5. Test whether its medium-effort answers are good enough for your job, because that is where its speed advantage is.
  • Your staff use the free apps. Claude's free plan includes Sonnet models. ChatGPT's free plan runs GPT-5.6 Luna, and OpenAI says GPT-6.1 Sol is not yet in regular ChatGPT chat. For a staff member on a free account, Claude currently gives access to a Sonnet model rather than a budget one.

One more difference to know about. Anthropic says higher-risk cybersecurity requests on Sonnet 5.5 "will visibly fall back to Sonnet 5." Routine bug fixing is not affected, but a security team may notice it.

How to choose for your own business

The benchmark gives you a starting point, not an answer. The only number that matters is the cost of your job at a quality you accept. Testing that takes an afternoon, not a project.

  • Collect 20 real examples. Take 20 real inputs from the job you want to automate: support emails, invoices, call transcripts, whatever it is. Real ones, not invented ones.
  • Run both models at medium and high effort. That gives you four results per example. Keep your prompt identical across all of them.
  • Mark each answer pass or fail yourself. Do not ask a model to grade the answers. You know what a correct invoice or a good reply looks like.
  • Divide cost by passes. Both APIs report the tokens each run used, so you can work out what each run cost. The model with the lowest cost per passing answer wins that job, even if its list price looks higher.

The answer is often different for different jobs, which is why the approach we described in our open-weight models article still holds. Send each job to the cheapest model that does it well, rather than picking one model for everything. Based on the numbers available today, our view is simple: GPT-6.1 Sol at medium or high effort is the value pick for high-volume background work, and Sonnet 5.5 earns its higher cost per task on customer-facing writing, documents, and replies where speed is worth paying for. We expect Haiku 5.5 to shake up the budget end when it arrives, and we will update this comparison when it does.

What this means for the AI inside your software

If you pay for software that uses AI, you rarely see which model runs underneath. It still decides your running cost. A call-analysis product like CallSentinel transcribes and scores every call, and a website chatbot like our AI Chat Assistant answers every visitor question. Both are exactly the kind of high-volume, repeated work where a quarter of the per-task cost adds up quickly. Ask any AI vendor you use whether they re-test models when prices change. The answer tells you whether the savings in this article will ever reach your invoice.

For the wider picture, our AI API pricing guide covers the price changes already scheduled for January and what the same job costs across providers, and our Claude Opus 5.5 guide covers when the top model is worth its price. If you want help working out which model fits a job in your business, talk to us.

Tags

Claude Sonnet 5.5 GPT-6.1 Sol Claude Sonnet 5.5 vs GPT-6.1 Sol Claude Sonnet 5.5 price GPT-6.1 Sol price AI API cost per task Best AI model for business Artificial Analysis

Share this article

JM

Jamil Malik

Founder & Lead Engineer, IO Snack

Started building software for businesses in 2015 — first as a solo developer, then as IO Snack. Builds and runs the CRM, ERP, POS, AI voice and call-analytics systems the articles here draw on, so the numbers come from production, not from a press release.

All articles by Jamil Malik

Want This Working in Your Business?

Tell us what you are trying to fix and we will tell you which module does it — or whether you need one at all.

Need help with your project?

Chat with us on WhatsApp