For the first time, open-weight AI models do most of the actual work flowing through one of the biggest AI routing services. Vercel's figures show them handling 56% of all tokens in August, and AT&T says it wants 65% to 70% of its own AI traffic on them. Yet those same open models earned only 14% of the money spent. If you pay for AI inside your own software (a chatbot, invoice reading, call summaries), that gap is the question: is your business paying premium prices for work a cheaper model could do? This article sets out what the numbers say, what AT&T actually did, what each option costs per job, and the three catches that should decide it for you.
What the Vercel numbers actually say
Vercel runs AI Gateway, a service that sits between thousands of production apps and the AI companies and passes their requests on. Every month it publishes what that traffic looks like. Its September report, published on 17 September and covering data to the end of August, is the clearest public picture of what businesses actually run, as opposed to what they talk about.
- Open models now carry most of the traffic. Open-weight models handled 56% of all tokens on the gateway in August, up from 7% in December 2025 and 13% in April. (A token is roughly three quarters of a word. Input, output, reasoning and cached tokens are all counted.)
- But they collect little of the money. That 56% of the work came to 14% of the spend. Anthropic alone took 64% of all spend in August, and at least 61 cents of every dollar in every month since December.
- Prices are falling fast. The average price per token on the gateway fell 23.2% in August, the third monthly drop in a row.
- Teams move quickly when a better deal appears. Z.ai's GLM-5.3-Flash overtook its predecessor one day after it appeared, and was handling two thirds of Z.ai's tokens by 31 August.
Vercel's own caveats matter. Spend is estimated from list prices, so real bills may differ, and Vercel recently widened what it counts as open-weight. This is one gateway's customers, not the whole market. But the direction matches what large companies are now saying in public. The Financial Times reported on 27 September, as summarised by PYMNTS, that executives mentioned open-weight or open-source models six times as often in August and September as in the same months of 2025.
Vercel's own summary puts it plainly: teams "can now get more inference from the same budget and reserve frontier models only for the tasks that justify the premium."
The saving does not come from choosing a cheap model. It comes from sending each job to the cheapest model that still does that job well.
How AT&T moved most of its AI onto cheaper models
AT&T is the most detailed public example, because one of its AI leaders explained the method on the record. In an interview with Fierce Network on 18 August, Mark Austin, a vice president in AT&T's Data Office, described what the company calls its "tokenomics" strategy.
- The problem was volume. AT&T went from around 8 billion tokens a day last year to upwards of 45 billion. By using open models, smarter routing and fine-tuning, it kept costs "relatively flat", Austin said.
- Old jobs move first. Austin said open models are generally six to ten months behind the best closed models. So AT&T lists every use case still running on a closed model that is more than a year old, on the reasoning that an open model has probably caught up. That was about 25% of its tokens.
- Some jobs do not need a language model at all. AT&T found seven use cases, mostly sorting things into categories, that it rebuilt as traditional machine learning. If the answer can only be A, B or C, you might as well "not burn tokens on that", he said.
- A router decides, job by job. AT&T built its own router that sends work to an open model and escalates to a premium one when needed, scored on "thousands of benchmark tasks".
- It chose American models. AT&T runs Google's Gemma and NVIDIA's Nemotron. Asked about DeepSeek and Kimi, Austin said "we look at them all" but that AT&T has "mostly focused on the American."
His advice for anyone starting is the most useful line in the interview: rank your use cases by how many tokens they use and what they cost, then decide for each one whether to tighten the prompt, move it to an open model, fine-tune, or split it across models. "Just coming up with your plan is probably the most important thing."
What one month of AI work costs on each model
Here is a worked example sized for a small business. Take 10,000 jobs a month, each one about 2,000 tokens in and 500 tokens out. That is roughly a customer message plus its context and a short reply, or one invoice read and summarised. The prices are list prices per million tokens from each company's own pricing page, checked on 28 September 2026, with no caching or discounts applied.
The prices behind that chart, per million tokens:
| Model | Type | Input | Output |
|---|---|---|---|
| Claude Opus 5.5 | Closed | $4.00 | $20.00 |
| GPT-6 Sol | Closed | $2.00 | $10.00 |
| Claude Sonnet 5 | Closed | $2.00 | $10.00 |
| GLM-5.3 | Open | $1.40 | $4.40 |
| DeepSeek V4 Pro, peak | Open | $1.32 | $3.96 |
| DeepSeek V4 Pro, off-peak | Open | $0.66 | $1.98 |
| Claude Haiku 4.5 | Closed | $1.00 | $5.00 |
| DeepSeek Flash, peak | Open | $0.30 | $1.20 |
| DeepSeek Flash, off-peak | Open | $0.15 | $0.60 |
| GLM-5.3-Flash | Open | $0.15 | $0.50 |
| GPT-6 Luna | Closed | $0.10 | $0.50 |
Two things stand out. First, the gap between the top and bottom of the table is about forty times, so the choice is real money even for a small business. Second, "open" does not automatically mean cheapest. At list price, the cheapest model in the table is GPT-6 Luna, a closed OpenAI model. The honest comparison is not open against closed. It is a large, expensive model against a small, cheap one, and every big vendor now sells both.
Where a cheap model is the right call, and where it is not
AT&T's rule works just as well for a small business: move the jobs that are simple, frequent and easy to check. In practice that means:
- Sorting and tagging. Is this email a complaint, an order or spam? Which expense category does this receipt belong to? Answers from a short list are exactly what cheap models do well, and some of them may not need a language model at all.
- Pulling fields out of documents. Invoice number, date, amount, supplier. The output is easy to check against the page.
- First-draft summaries. Call notes, long email threads, daily reports that a person skims before acting on them.
- High-volume, low-stakes replies. Opening hours, delivery status, price lists, where a wrong answer is an irritation rather than a loss.
Keep the expensive model where a mistake costs more than the model does:
- Anything a customer reads as a commitment. Refund decisions, quotes, contract wording.
- Judgement calls. Deciding whether a camera event is a real threat, or whether an angry customer should go straight to a manager.
- Long, multi-step work. Agents that plan, call tools and check their own output. This is where the six-to-ten-month gap Austin described shows most.
The three catches nobody puts in the headline
1. Where your data goes
"Open-weight" describes the model, not the company running it. If you call DeepSeek's own API, its privacy policy says it will "collect, process and store your Personal Data in People's Republic of China." That may be fine for tagging product descriptions. It is a different decision for customer records, payroll or anything covered by a contract with your own clients. The same open models are also offered by other hosting companies and can be run on your own hardware, and that is what AT&T does. You get the model without sending the data to its maker, though usually at a different price.
2. Peak hours that land on your working day
DeepSeek charges double during its peak hours: 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. In Pakistan that is 6am to 9am and 11am to 3pm, and in the UAE 5am to 8am and 10am to 2pm. That covers much of the business day in both places. A customer chat that runs during office hours pays the peak rate. Overnight batch work, like reprocessing yesterday's invoices, gets the half price. Plan with the rate you will actually pay, not the headline one.
3. Switching has a cost the price table hides
Austin made a point that most price comparisons miss. When you switch models in the middle of a long task, you lose the cache, and cached input is about a tenth of the normal price. A conversation that has built up a long history can be cheaper to finish on the expensive model than to restart on a cheap one. Every switch also needs retesting. Different models format answers differently, and a prompt that works on one can quietly fail on another. That is why AT&T scores its router on thousands of tests rather than switching on instinct.
What we have learned from running AI in production
We run AI inside live products, including the analysis behind our AI Cam security cameras and the answers from our AI customer-support agent. The Vercel data matches what running those has taught us:
- Measure cost per finished job, not per token. A model with a lower token price that writes twice as much, or has to be asked twice, is not cheaper. We record what each analysis actually costs, and that number drives the decision, not the rate card.
- Always have a second provider from a different company. Two models from the same vendor usually share the same billing account and the same outage. A real fallback is a different company, and that matters more than a few cents per thousand jobs.
- Test on your own work before you switch. Public leaderboards measure general skill. Your invoices, your customers' Urdu and English messages and your camera angles are specific. Run a few hundred real examples through the cheaper model and compare the answers before you commit.
- Do not rebuild a job that already costs almost nothing. If a task costs you five dollars a month, the engineering time to move it is worth more than the saving. Start with the jobs at the top of your bill, exactly as AT&T did.
Our view: the 56% figure is not a sign that the big closed models are losing. They still take 86% of the money, and the likeliest reason is that the hardest and most valuable work still goes to them. The lesson for a business is to stop using one model for everything. A business that picked one AI a year ago and still sends everything to it is the one most likely to be overpaying today.
Your next step: split your AI bill by job
You do not need AT&T's scale to use AT&T's method. This week, list every place your business uses AI through an API, roughly how many times a month each one runs, and which model it uses. Pick the highest-volume simple job and test it on a small model, whether that is an open one or a cheap closed one like GPT-6 Luna. Keep the expensive model for the jobs where it earns its price.
Prices in this area change monthly. Some are already scheduled to rise, as we set out in what doubles on 1 January 2027. If you are choosing a subscription for staff rather than an API for software, our ChatGPT vs Claude vs Gemini comparison with Pakistan prices covers that choice. And if you would rather have a customer chatbot or AI camera system where the model routing and fallback are already handled for you, talk to us.