On 15 September Google put a price on an AI that can hold a phone conversation: half a cent per minute to listen and 1.8 cents per minute to talk. Five days earlier OpenAI put its own voice model in the API at 5 cents a minute, and xAI's Grok, which an independent leaderboard puts within two points of the top, is 8 cents. All three numbers are real, all three are on the vendors' own pricing pages, and all three are the smallest line on the bill you will actually pay. We run voice agents and call analysis in production, so this article does the arithmetic the launch posts leave out: the voice layer, what sits on top of it, what an independent benchmark says, the small print, and what 300 calls a month works out to for a clinic in Lahore or a trading firm in Dubai.
What Google Shipped on 15 September, and What It Costs
Google's announcement, by Tom Ouyang and Malini Jaganathan, introduces two models. Gemini 3.8 Live is "built for scale and cost efficiency, combining conversational intelligence with fluid dialogue". Gemini 3.8 Live Extended Thinking is "built for high-complexity tasks, with increased intelligence and multi-step reasoning". Both are audio-to-audio: the model hears the caller and speaks back directly, with no separate speech recogniser or voice in between. Both are in the Gemini API and Google AI Studio today, and the model list marks them stable, not preview.
The pricing page lists the two models together, and the paid-tier rows read, verbatim:
- Audio in: $3.00 per million tokens, or $0.005 per minute. That is the caller's side of the line.
- Audio out: $12.00 per million tokens, or $0.018 per minute. That is the AI speaking. Talking costs 3.6 times as much as listening.
- Text in $0.75, text out $4.50 per million tokens. Your instructions, your knowledge base, and, for the Extended Thinking model, the thinking. The output row says "including thinking tokens".
- Image and video in: $1.00 per million tokens, or $0.002 per minute. Only relevant if the agent is watching a camera as well as listening.
Unlike the Gemini 3.8 Flash text rows, which carry a "through December 31, 2026" note and double on 1 January, the Live rows have no dated wording on them as of 16 September. We covered that January step-up in our post on AI prices with an expiry date; the voice rows are not part of it, at least not yet.
Three other facts from Google's own pages matter for a business phone line. The model "automatically detects and transitions between 97 supported languages mid-conversation". It "executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting", which is the difference between an agent that books the appointment and one that only talks about it. And "all audio generated by our AI products is watermarked with SynthID", so a recording of your AI receptionist is detectable as AI-generated.
OpenAI's Answer, Five Days Earlier: GPT-Live-1 at 5 Cents a Minute
On 10 September OpenAI released GPT-Live-1 in its API. It is the same full-duplex model behind ChatGPT's voice mode: it "is capable of listening and speaking at the same time" and "can delegate deeper reasoning and actions to the models and tools it is paired with". OpenAI lists "telephony support" as a headline strength, "for phone calls, from restaurant reservations to customer support", and quotes Yelp using it to answer reservation and food-order calls.
The price is on OpenAI's pricing page in one line: gpt-live-1, $0.05 per minute, and "GPT-Live 1 voice sessions are billed per second, without rounding up to a whole minute. Backend model and tool usage is charged separately." That last sentence is the structural difference between the two products.
- Gemini 3.8 Live thinks and talks in one model. You pay per minute of audio in and out, plus text tokens for instructions and, on the Extended Thinking model, for the reasoning it does in the background.
- GPT-Live-1 is the voice; the brain is a separate bill. OpenAI's own post says developers "might pair GPT-Live-1 with a model like Luna for high-volume tasks like scheduling or order updates, and use a model like Astra for complex customer issues". On the same pricing page, GPT-5.6 Luna is $0.20 in and $1.20 out per million tokens; GPT-6 Astra is $10.00 in and $50.00 out. The voice minute is fixed; what you pair it with decides the rest.
The voice model is now the cheapest thing on an AI phone call. The money is in what sits behind it, what sits in front of it, and what you do with the recording afterwards.
The Comparison Table, Priced From the Vendors' Own Pages
Every figure below is from Google's Gemini API pricing page, OpenAI's API pricing page or xAI's models page, read on 16 September 2026. Per-minute figures are the ones the vendors publish; we have not converted token prices into minutes ourselves.
| What you are paying for | Gemini 3.8 Live | Gemini 3.8 Live Extended Thinking | GPT-Live-1 | Grok Voice Think Fast 2.0 |
|---|---|---|---|---|
| Caller's audio (in) | $0.005 / min | $0.005 / min | $0.05 / min, billed per second | $0.08 / min of audio ($4.80 / hour) |
| AI's audio (out) | $0.018 / min | $0.018 / min | ||
| Instructions and knowledge base (text in) | $0.75 / 1M tokens | $0.75 / 1M tokens | Backend model's rate: Luna $0.20, Terra $2.00, Sol $4.00, Astra $10.00 / 1M | Text input billed separately at xAI's rate card price |
| Reasoning | Included in the audio minute | Thinking billed as text out, $4.50 / 1M tokens | Backend model's output rate: Luna $1.20 up to Astra $50.00 / 1M | Included in the audio minute |
| Tools (booking, lookups) | Your own systems | Your own systems | "Charged separately" | Your own functions, or xAI's built-in web search, X search and file search |
| Languages | 97, switching mid-conversation (Google's list includes Urdu, Arabic, Punjabi, Sindhi, Pashto) | "Expand voice options and language availability over the coming months"; no count published | "20+ languages with native-quality accents", Arabic and Hindi named; custom voices available | |
| Free tier | Yes, and the pricing page says free-tier data is "used to improve our products" | None listed | None listed | |
Read across the first two rows and the gap is large: a minute in which the caller talks half the time and the AI talks half the time costs about 1.4 cents on Gemini 3.8 Live, 5 cents on GPT-Live-1 and 8 cents on Grok before any brain is counted. Read the whole table and the gap narrows, because on Gemini the thinking is either included (3.8 Live) or billed as output tokens (Extended Thinking), and on OpenAI it depends entirely on which backend you choose. Which brings us to the only independent numbers available.
What an Independent Benchmark Says, and the 1.2 Seconds Nobody Mentions
Artificial Analysis runs a Speech to Speech leaderboard that both companies cite. Google's launch post quotes it for the "#1 overall spot" at 82.6. We read the leaderboard ourselves on 16 September rather than trusting the quote, and here is what it shows for the models a business would actually consider. The index is a weighted blend of reasoning, agentic task completion, human preference and task success; "cost per hour" is what Artificial Analysis says it paid to run a fixed 40-question audio task through each model, normalised to an hour of input audio, so it includes whatever reasoning the model did along the way.
Measured cost per hour of input audio, Artificial Analysis fixed task (USD, read 16 Sep 2026)
Source: artificialanalysis.ai/speech-to-speech, "Cost per Hour of Input Audio" column and "Speech to Speech Index". Artificial Analysis notes the two GPT-Live-1 rows are based on a single trial.
Four things stand out, none of them in either launch post.
- The cheap Google model is the one that answers fastest. Artificial Analysis measures time to first audio at 1.18 seconds for Gemini 3.8 Live, 1.35 for Extended Thinking, 1.24 and 1.34 for the two GPT-Live-1 configurations, and 0.70 for Grok Voice Think Fast 2.0. On a phone line a pause over a second is audible; callers fill it with "hello?". Test whatever you deploy on a real handset, not in a browser tab.
- The top four scores sit within 2.5 points; their costs do not. 82.6, 81.5, 81.3, 80.1, and then $3.50, $5.83, $4.80 and $4.47 an hour for the same four. The plain Gemini 3.8 Live scores 76.0 at $0.84 an hour, a seventh of the dearest, with a "task success rate" of 93.2%, the best of the Google models. For scheduling, order status and opening-hours questions, that is the model to price first.
- Agentic completion is where the gap opens. On the τ-Voice customer-service benchmark, Extended Thinking resolves 68.6% of scenarios and plain 3.8 Live resolves 30.1%. If your calls are multi-step work, change a booking, check an account, confirm, the cheap model will cost you in failed calls what it saves in tokens.
- Sample sizes are small. Artificial Analysis states the GPT-Live-1 agentic results rest on one trial each. Treat the ordering as indicative, not settled.
Three Pieces of Small Print That Change the Bill
These come from Google's session-management documentation and its pricing page.
- A phone call is longer than a connection. Google's docs say "the lifetime of a connection is limited as well, to around 10 minutes. When the connection terminates, the session terminates as well." A ten-minute support call will hit that. The fix, "session resumption", is documented, but your integrator has to build it, and the server warns with a "GoAway" message that includes the time left. Separately, "without compression, audio-only sessions are limited to 15 minutes"; context window compression has to be switched on for anything longer. If a vendor demos a two-minute call and quotes you for hour-long ones, ask about minute eleven.
- Thinking is billed as talking. On the Extended Thinking model the output row is "$4.50 (text)" and it says "including thinking tokens". A caller with a complicated request can make the model think for thousands of tokens before it says a word, and every one of those is on the output rate. This is why the measured hourly cost is $3.50 against listed audio rates that add up to $1.38 an hour even if both sides were billed for every minute. Budget from the measured figure, not the rate card.
- The free tier is not free of consequences. The same pricing table says free-tier usage is "used to improve our products" and paid-tier usage is not. For a phone line that hears patient names and order numbers, that is the wrong tier to prototype on, for the reasons we set out in our piece on who can read your AI conversations.
Urdu, Arabic, Punjabi and Sindhi Are on the List
For readers in Pakistan and the Gulf the language question decides everything, and it can be answered from Google's own capabilities page rather than from a marketing line. The Live API's supported-language table, with BCP-47 codes, includes Urdu (ur), Arabic (ar), Punjabi (pa), Sindhi (sd), Pashto (ps), Persian (fa), Hindi (hi) and Bengali (bn), alongside English. Because the model switches languages mid-conversation, a caller who opens in English and drifts into Urdu, which is how most Karachi phone calls actually go, does not have to be routed anywhere.
OpenAI's post lists twelve new voices, one described as "Australian English influenced", and says it "will continue to expand voice options and language availability over the coming months". It publishes no language count for GPT-Live-1. That is the state of the documentation on 16 September, and the first thing to check in a demo if your callers are not English speakers.
A Worked Example: 300 Calls a Month
Take a clinic or a small trading firm that gets 300 calls a month, averaging four minutes, with the caller and the AI each talking about half the time. That is 1,200 minutes of call. Using only the published per-minute rates:
- Gemini 3.8 Live voice layer: about $17 a month. 1,200 minutes of listening at $0.005 is $6.00; 600 minutes of speaking at $0.018 is $10.80. Add text tokens for your instructions and knowledge base at $0.75 per million; how many times those are counted depends on how the session is managed, so treat it as a second line and read it off your first week's usage.
- GPT-Live-1 voice layer: $60 a month. 1,200 minutes at $0.05, billed by the second. Then a backend model: on Luna at $0.20 in and $1.20 out, the reasoning for calls like these is small change; on Astra at $10 and $50, it is not.
- Grok Voice Think Fast 2.0: $96 a month. 1,200 minutes at $0.08, reasoning included. The dearest of the three per minute, and on the independent leaderboard the fastest to answer and the most likely to finish the task.
- Telephony: whatever your carrier already charges per minute. A SIP trunk, a GSM gateway or a cloud number all bill their own minutes, and that line does not shrink because the AI got cheaper. Neither vendor publishes it, so we will not invent it.
- The part that was always the real cost: the work after the call. Transcribing, scoring, flagging the angry caller, catching the missed disclosure, putting the booking in the diary. None of it is in a per-minute voice rate.
At these numbers, the model is not the decision. Seventeen dollars a month, sixty or ninety-six, any of them is less than one missed booking. The decision is whether the thing on the line can actually complete the call in the caller's language, hand a difficult one to a human without dropping it, and leave you a record you can act on.
What to Do If You Are Deciding This Quarter
We have run voice agents on real lines and analysed real recordings long enough to have opinions.
- Price the cheap model first, then prove it fails. Start with the low-cost tier for hours, directions, order status and appointment booking. Move to the reasoning tier only for the call types where the cheap one measurably fails, on your calls, not a leaderboard.
- Ask every vendor the minute-eleven question. How do long calls survive the connection limit, and what happens to the conversation history when the connection is resumed? A vendor who has not hit that wall has not put an agent on a real line.
- Insist on a handover and a record. The agent must be able to pass a call to a person, and every call must come back as a transcript with a score, not an audio file in a folder.
- Do not prototype on a free tier with real callers. The pricing table tells you why.
If you would rather not assemble the layers yourself, SmartLine puts an AI voice agent on your existing phone lines, SIP trunks, GSM gateways and website, speaks Urdu, English and Arabic, and returns every call as a transcript with sentiment, action items and compliance scoring. CallSentinel does the after-the-call part for calls your human team takes: transcription with speaker labels, five-dimension scoring and objection coaching. For the business case, our earlier post on AI voice agents sets out what a missed call costs and where the agent pays off first. Or talk to us, in whichever of the 97 languages you prefer.