If you are adding AI to a support chat, an invoice reader or a call summariser, the first question is usually "which model is cheapest?" The price list gives a quick answer, and it is the wrong one. Two models with similar per-token prices can differ by more than twenty times in what a finished job costs, because cheap models spend very different amounts of "thinking" on each task. This article puts the lowest-priced models from OpenAI, Google, Anthropic and DeepSeek side by side, using the vendors' own price pages and an independent cost-per-task measurement, and says which one we would start with for which job.
What the cheapest models cost per million tokens
A token is roughly three quarters of an English word. Every vendor charges separately for what you send in (the question plus your instructions) and what comes back (the answer). These are the standard list prices from each vendor's pricing page, checked on 3 October 2026.
| Model | Input per 1M | Output per 1M | Vendor |
|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | OpenAI |
| DeepSeek V4.1 Flash | $0.15 to $0.30 | $0.60 to $1.20 | DeepSeek |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | |
| Claude Haiku 4.5 | $1.00 | $5.00 | Anthropic |
| Gemini 3.8 Flash | $0.75 until 31 Dec 2026, then $1.50 | $3.75 until 31 Dec 2026, then $7.50 | |
| GPT-6.1 Sol (for reference) | $2.00 | $10.00 | OpenAI |
DeepSeek publishes a range because it charges half price off-peak. Anthropic and Google both state a 50% Batch discount for work that can wait, OpenAI lists separate Batch prices, and all three charge less for repeated, cached input. Google also still lists the older Gemini 2.5 Flash-Lite at $0.10 and $0.40, but says it is limiting the 2.5 models to users who have already used them and points new projects to 3.5 Flash-Lite or 3.8 Flash, so we left it out. Google says that on its free tier your content may be used to improve its products, while paid-tier content is not, so a business app belongs on the paid tier.
What 10,000 real jobs cost at list price
Take a typical support ticket: about 3,000 tokens in (your instructions plus the customer's message and history) and 500 tokens out (the reply). That split is our assumption, not a vendor figure, so replace it with your own. Here is what 10,000 such tickets cost on list price alone.
| Model | Cost per ticket | Cost for 10,000 tickets |
|---|---|---|
| GPT-6 Luna | $0.00055 | $5.50 |
| DeepSeek V4.1 Flash | $0.00075 to $0.0015 | $7.50 to $15.00 |
| Gemini 3.1 Flash-Lite | $0.0015 | $15.00 |
| Gemini 3.5 Flash-Lite | $0.00215 | $21.50 |
| Gemini 3.8 Flash (to 31 Dec) | $0.0041 | $41.25 |
| Claude Haiku 4.5 | $0.0055 | $55.00 |
| GPT-6.1 Sol | $0.011 | $110.00 |
On this view, 10,000 tickets cost under $110 on any of them, and the spread between cheapest and dearest is about twenty times. But this table assumes every model uses exactly 500 tokens to answer. That is the assumption that breaks.
Cost per task: the number that changes the ranking
Modern models can "reason" before they answer. The reasoning tokens are billed as output, and different models spend very different amounts of them. Artificial Analysis is an independent tester that runs models through the same set of tasks and publishes an Intelligence Index score plus the average cost of completing its tasks. Their index combines 10 evaluations covering agent work, coding, scientific reasoning and general knowledge, and its published cost covers the model calls being tested. The figures below are from its leaderboard on 3 October 2026.
Luna at low effort and Gemini 3.5 Flash-Lite score the same 22. Luna's price per token is lower, but that is a 3x gap on input and 5x on output, not 27x. The rest of the gap is that Luna at low effort uses far fewer tokens to finish. Haiku 4.5, the cheapest Claude, scores lower (17) at a higher cost.
In plain terms, "effort" is a setting that tells the model how long to think before it answers. At low effort it answers almost straight away and spends few tokens. At high or max effort it writes out long chains of reasoning first, and you pay for every word of it even though your customer never sees it. That is why a model with a cheap price list can still produce an expensive bill, and why a model that looks pricier per token can come out cheaper per job when it needs less thinking to reach the same answer.
The picture is the same one level up, where the models are good enough for real drafting and triage work.
Two lessons. First, the cheapest model on the price list is not always the cheapest per job, and the "Flash" and "Flash-Lite" names say nothing about how much a model thinks. Second, among the cheap models GPT-6 Luna is the outlier on cost per task: Artificial Analysis names Luna at low effort, together with Granite 4.2 3B, as having the lowest cost per task of the models it tracks.
The catch: effort level, speed and what the score means
Luna is not free of trade-offs, and the same table shows them.
- Effort is a dial, and it trades money for time. Artificial Analysis measured Luna at several effort levels. Low scores 22 at $0.0045 and finishes in about 6 seconds. High scores 33 at $0.03 and takes about 16 seconds. Max scores 38 at $0.07 but takes about 133 seconds. A customer waiting in a chat window will not wait two minutes, so max effort is for background work such as overnight reports.
- The index is a hard test. It includes Humanity's Last Exam, terminal-based agent work and scientific reasoning. A score of 22 does not mean a model cannot answer "where is my order?". It means it will struggle with long, multi-step jobs. In our view, a low score is a warning for agents that must act without a human check, and nearly irrelevant for classifying a ticket or drafting a short reply.
- Sol is not that much dearer where it counts. GPT-6.1 Sol at medium effort scores 48 for $0.21 per task, and at low effort it scores 42 for $0.13. If a wrong answer costs you a customer, paying more per job for a score in the 40s can be the cheaper decision overall.
- Google's price moves on 1 January. Gemini 3.8 Flash input and output both double from 1 January 2027, from $0.75 and $3.75 to $1.50 and $7.50 per million tokens, according to Google's pricing page. We covered that in what doubles on 1 January 2027. Budget on the higher price.
The cheapest model is the one that finishes your job at the lowest total bill, and the only way to know that is to count what each model spends on your own work.
Which cheap model we would start with, by job
- Sorting and labelling (tickets, leads, call tags). Start with GPT-6 Luna at low or medium effort. A score around 22 to 30 is enough to classify, and the cost per task is the lowest listed. Answers are short and fast.
- Customer-facing chat. Start with Luna at high effort, or GPT-6.1 Sol at low effort if answers must be right more often. Keep total response time under about 15 seconds, which rules out max effort.
- Long overnight jobs (summaries, report drafts, bulk extraction). Use the Batch discount, which Anthropic and Google both state at 50% (OpenAI lists its own Batch prices), and run a higher effort level because nobody is waiting.
- Work where the data must stay under strict rules. Check the vendor's data terms before price. Use paid tiers only, and ask where your provider processes data. This is a legal question, not a cost one, and no leaderboard answers it.
Test your own jobs before you commit
- Collect 50 real examples. Take real tickets, invoices or calls, including the awkward ones, not the clean demo cases.
- Run the same 50 through three models. Luna at two effort levels, one Flash model and one Sol or Sonnet run is plenty.
- Count tokens and seconds, not just answers. Your vendor dashboard shows tokens used per request. Multiply by the list price from the first table.
- Score the answers blind. Have someone who knows the business mark each answer as usable or not without seeing which model wrote it.
- Pick on cost per usable answer. Total spend divided by the number of answers someone would actually send. That is the number that matters.
Leaderboard numbers move when models are re-tested, and prices change, so treat everything above as a snapshot from 3 October 2026 and re-check before you sign anything.
Where to go from here
If you would rather not run this test yourself, the AI Chat Assistant and CallSentinel run on hosted AI models, so the cost question is one we deal with in production. For the bigger models, read Claude Sonnet 5.5 vs GPT-6.1 Sol, and for self-hosting instead of paying per token, the open-weight cost guide. If you want help choosing for your own workload, talk to us.