If you have a budget line for AI, go and look at the rate you used to build it. There is a good chance it expires. Google's public pricing page states that Gemini 3.8 Flash costs $0.75 per million input tokens "through December 31, 2026" and $1.50 "starting January 1, 2027" — input, output, caching and cache storage all exactly double, on a date that is already printed. OpenAI's pricing page carries a similar note against GPT-5.6 Sol. Anthropic cancelled a scheduled increase of its own last month. None of this is rumour or leak; all of it is published by the vendors, and almost nobody has put it into a spreadsheet. This article does that, and covers the three other mechanics that can move your AI bill without any headline rate changing at all.
What Actually Changes on 1 January 2027
Google released Gemini 3.8 Flash on 2 September 2026, describing it in its own release notes as "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows". It shipped at the same introductory rate as 3.7 Flash and 3.6 Flash before it. Read the pricing page rather than the launch post and the introductory part is explicit:
- Input. $0.75 per million tokens through 31 December 2026, then $1.50.
- Output. $3.75 per million tokens through 31 December 2026, then $7.50.
- Context caching. $0.075 per million tokens, then $0.15.
- Cache storage. $0.50 per million tokens per hour, then $1.00.
- Batch tier. $0.375 in and $1.875 out, then $0.75 and $3.75. Batch is half price, and the half price doubles too.
Every one of those is a clean 2x. There is no tapering and no grandfathering mentioned. The same dated wording appears against Gemini 3.7 Flash and 3.6 Flash, so moving to an older Flash model does not escape it either.
Take a modest production workload — 100 million input tokens and 10 million output tokens a month, which is what a busy document-processing or support-triage job looks like. On Gemini 3.8 Flash that is $112.50 a month today and $225.00 from January. Over a year, the same work costs $1,350 more, for the same volume, the same model and the same code.
The same month of work, before and after
100M input + 10M output tokens on Gemini 3.8 Flash, at Google's published rates
Through 31 December 2026
$112.50
From 1 January 2027
$225.00
A difference of $1,350 a year on one workload. Source: Google's Gemini API pricing page.
Three Other Ways a Quoted Price Stops Being the Price You Pay
The dated increase is the easy one to plan for, because it announces itself. The other three do not.
1. Context length is a price tier, not just a limit
OpenAI publishes two columns for its flagship models: short context and long context. GPT-6 Astra is $10 per million input tokens and $50 per million output in the short column, and $20 and $75 in the long one. Feed the same job a longer prompt and the rate itself changes. On our reference workload that is $1,500 a month against $2,750 — roughly 1.8x for identical token counts, decided by how long each individual request happened to be.
Google does the same thing on Gemini 3.1 Pro Preview, where prompts up to 200k tokens are $2.00 per million input and prompts above it are $4.00. If your application stuffs a growing conversation history or a document set into each call, your average rate drifts upward on its own as usage grows. Nothing changed on the pricing page. Your bill still went up.
2. Data residency is a multiplier
This one matters specifically to readers in Pakistan and the Gulf, where "where does the data physically sit" is now a routine question in a procurement conversation. OpenAI states that regional processing endpoints "are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency". Anthropic applies a 1.1x multiplier "on all token pricing categories" when you pin inference to the US, cache reads and writes included. Residency is a legitimate requirement and often a non-negotiable one, but it is priced, and the price is a percentage of everything rather than a fixed fee.
3. The token count itself can change under a flat price
This is the one almost nobody budgets for, and it is stated plainly in Anthropic's own documentation. Claude 4.7 and later models use a newer tokenizer which, in Anthropic's words, "produces approximately 30% more tokens for the same text", with the exact increase depending on the content and workload shape.
Think about what that means for a forecast. The advertised price per million tokens can hold perfectly steady while the number of tokens the same paragraph turns into goes up by roughly a third. If you upgrade a model and your bill rises without your traffic rising, this is a candidate explanation, and it will not appear anywhere on a pricing page as a price change — because it is not one.
The unit price of AI is not a fact about the market. It is a term with a date on it, a tier, a multiplier and a unit of measurement, and any of the four can move while the other three sit still.
What the Same Job Costs Right Now, Across Providers
Every figure below is arithmetic on the vendors' own published rates, for the same 100 million input and 10 million output tokens a month, standard tier, short context where a provider distinguishes one. No benchmark, no opinion about quality — just what the meter reads.
| Model | Input / output per 1M | Cost for the month |
|---|---|---|
| GPT-5.6 Luna | $0.20 / $1.20 | $32 |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $55 |
| Gemini 3.8 Flash (to 31 Dec) | $0.75 / $3.75 | $112.50 |
| Claude Haiku 4.5 | $1.00 / $5.00 | $150 |
| Gemini 3.8 Flash (from 1 Jan) | $1.50 / $7.50 | $225 |
| Gemini 3.5 Flash | $1.50 / $9.00 | $240 |
| Claude Sonnet 5 | $2.00 / $10.00 | $300 |
| GPT-5.6 Terra | $2.00 / $12.00 | $320 |
| GPT-6 Astra | $10.00 / $50.00 | $1,500 |
| GPT-6 Astra (long context) | $20.00 / $75.00 | $2,750 |
The spread between the top and bottom of that table is about 86x. Most business tasks — reading an invoice, classifying a support message, summarising a call — do not need the row at the bottom. The single largest cost decision most teams make is not which vendor they sign with; it is whether the cheap model was ever tried on the boring 80% of the work.
Why the Cheaper Model Is Sometimes the Newer One
There is a habit, carried over from every other kind of software, of assuming the older version is the budget option. In AI pricing right now that is often backwards.
- Gemini 3.5 Flash costs $1.50 in and $9.00 out. Gemini 3.8 Flash, released months later, is currently $0.75 and $3.75. The newer model is less than half the price of the one it succeeded — until 1 January, when it becomes slightly cheaper rather than dramatically cheaper.
- Claude Sonnet 4.6 is $3 in and $15 out. Claude Sonnet 5 is $2 and $10. Again, newer and cheaper.
- Anthropic cancelled an increase. Its documentation states that the $2/$10 Sonnet 5 pricing, "announced at launch as introductory pricing through August 31, 2026, is now the standard price", and that the scheduled rise to $3/$15 on 1 September "will not occur". Introductory windows can close upward, and they can also simply be made permanent.
So "we are on the older model to keep costs down" deserves checking against the current page rather than the memory of it. Staying put can be the expensive choice, and the pin that holds a team on an old model is usually the cost of re-testing prompts, not the token rate.
How We Handle This in Our Own Metered AI
We bill AI usage inside our own products, so these mechanics are not theoretical for us. Three practices have earned their place, and each one exists because the alternative caused a problem.
- Meter what was actually spent, not what you assumed. Cost comes from the token counts the provider returns on each response, recorded per request. An estimate built from an average prompt length stops matching reality the moment the workload shifts, and you find out at the end of the month.
- Fix the rate at the moment you quote. If you tell a customer what a job will cost, that quoted rate has to travel with the job through to settlement. Otherwise a price change between the quote and the run silently lands on someone — either the customer, who was promised a number, or you.
- Never charge a customer for your own failover. If your first-choice provider is rate-limited and the request falls through to a more expensive one, that is your resilience decision, not the customer's purchase. The cost of the fallback belongs to whoever chose to have a fallback.
The same discipline applies whether you are building on these APIs or buying software built on them. If a vendor cannot tell you how their AI charges are derived, they are carrying this exposure somewhere, and eventually it arrives on your invoice or in their margin. One of those is worse for you than the other, but neither is good.
What to Do Between Now and the End of December
Concrete, in the order that pays back fastest.
- Re-read the pricing page for every model you call. Not a tracker site, not a blog — the vendor's own page, where the dated wording lives. Note every "through" and "starting" date you find, with the model name against it.
- Rebuild your 2027 forecast at the January rates. If the doubled number is uncomfortable, you have found out in September rather than in January.
- Check the long-context column. Look at how many of your requests cross the threshold. Trimming a prompt below a tier boundary is often the single cheapest optimisation available.
- Try the smaller model on your boring work. Route the routine 80% to a cheap model and keep the expensive one for what genuinely needs it. On our reference workload that is the difference between $32 and $1,500 a month.
- Turn on caching, and understand what it costs. Cache reads are a fraction of input price, but writes cost more than input, so caching only pays once content is genuinely reused.
- Ask your software vendors the question directly. "Is any AI cost in my subscription based on an introductory rate that expires?" is a fair question, and the answer tells you a lot.
This matters most where AI is running continuously rather than occasionally. Our AI Cam platform judges camera events around the clock, and IO Snack Media and CallSentinel both meter AI work per job, which is exactly why the quoting-and-settlement discipline above is built into them rather than bolted on. If your AI spending sits alongside the rest of your books, IO Snack Accounts is where it becomes a cost line you can actually see against revenue. For the wider picture of what these systems are being asked to do, our earlier pieces on agentic AI in business automation and on what autonomous AI means for your security posture cover the capability side of the same story.
The prices in this article were read from Google's, OpenAI's and Anthropic's own pricing pages on 7 September 2026. Given the last three months, check them again before you sign anything. If you would rather have someone else watch this for you, talk to us — or see what we run on it across the full product range.