Mistral released a public preview of Mistral Large 4 on 6 October. It is a trillion-parameter model from Europe's biggest AI lab, priced at $1.36 per million input tokens and $4.18 per million output tokens, with open weights promised by the end of the month. On price per token it looks like one of the cheapest serious models you can buy. The first independent test tells a different story: per finished task, it costs more than models that score higher. This article covers what Mistral actually shipped, what it costs, what Artificial Analysis measured, how it compares with DeepSeek, OpenAI and Anthropic at the same tier, and who should use it.
What Mistral released on 6 October, and what is still to come
Mistral calls the model ML4 for short, and "le Chonk" for fun. The facts from its own announcement:
- Size. One trillion parameters in total, of which 49 billion are active for any one answer. That design, called mixture of experts, is why a model this large can still be served at a mid-range price.
- What it does. One model for instructions, step-by-step reasoning and agent work. It reads text and images and writes text.
- Who can use it today. Anyone with a Mistral account, through the preview API on Mistral Studio. This is a public preview, not a closed beta.
- What is still coming. The weights, which Mistral says it will release "by the end of the month", along with details of the architecture and more benchmarks. Until then, Mistral says it is red-teaming the model with cybersecurity firms, vetted partners and state authorities.
- Where it was built. Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and served from the same machines. Mistral says the model will be available in several regions, including a European deployment it runs end to end under European law.
- Languages. Mistral says a large share of the training data covered more than 160 languages, including every official EU language. It does not name Urdu or Arabic, so test your own language before you rely on it.
Mistral also says the training run behind this preview "is still in flight" and that it expects "large and rapid improvements in the weeks and months to come". That matters for everything below. You are judging a model that its maker says is not finished.
The price: $1.36 in, $4.18 out, and what that buys you
On the price list, Mistral Large 4 sits well below the mid-tier models most businesses use. Claude Sonnet 5.5 and GPT-6.1 Sol both cost $2 per million input tokens and $10 per million output tokens. Mistral charges about a third less for input and well under half for output. Mistral's pricing page adds that batch jobs cost 50% less and repeated (cached) input up to 90% less, and Artificial Analysis lists a 90% cache discount for this model.
Here is the same tier side by side. Prices are per million tokens. The score is the Artificial Analysis Intelligence Index, an independent average of ten tests of real work, and cost per task is what Artificial Analysis paid, on average, to run one of those tasks. All figures were checked on 7 October 2026.
| Model | Input | Output | Score | Cost per task |
|---|---|---|---|---|
| Mistral Large 4 Preview | $1.36 | $4.18 | 38 | $1.13 |
| DeepSeek V4 Pro 0813 (max) | $1.32 | $3.96 | 36 | $0.67 |
| GLM-5.3 (max) | $1.40 | $4.40 | 45 | $2.01 |
| Kimi K3 (max) | $3.00 | $15.00 | 44 | $2.00 |
| Qwen3.8 Max (0902) | $2.00 | $6.00 | 45 | $5.41 |
| Claude Sonnet 5.5 (medium) | $2.00 | $10.00 | 41 | $0.59 |
| GPT-6.1 Sol (low) | $2.00 | $10.00 | 42 | $0.13 |
The words in brackets are effort settings. Most new models let you choose how hard they think, and the same model scores and costs very differently at each setting. We have picked the setting nearest to Mistral's score for each rival so the comparison is fair.
The independent number: 38 points at $1.13 a task
Artificial Analysis tested the preview within a day of launch. Its results put Mistral Large 4 at 38 on its Intelligence Index, 64th of the 225 models it currently lists. That is above the median of 26 for comparable models, and a big step up for Mistral: its previous flagship, Mistral Large 3, scores 9 on the same index, and Mistral Medium 3.5 scores 14.
The problem is the bill. Artificial Analysis found the model "very verbose": it produced 200 million tokens to complete the index, against a median of 81 million. Every one of those tokens is charged at the output rate. So a low price per token turns into a high price per job.
Read the chart from the bottom up. GPT-6.1 Sol at its highest setting scores 14 points more than Mistral Large 4 and still costs less per task, even though its output tokens cost more than twice as much. At its lowest setting, Sol scores 4 points more for about a ninth of the cost. The gap is almost entirely verbosity. Sol at low effort used 9 million tokens to finish the index, and Claude Sonnet 5.5 at medium used 31 million. Mistral Large 4 used 200 million.
The price per token is the sticker. The number of tokens a model needs to finish your job is the meter, and on day one, Mistral Large 4's meter runs fast.
One more point in fairness to Mistral. It is also fast: Artificial Analysis measured 116 output tokens per second, above the average of 86, and a context window of 524,000 tokens. Speed helps a chat feel responsive. It does not reduce what the job costs.
Mistral Large 4 vs DeepSeek V4 Pro: the head-to-head at the same price
The nearest rival is DeepSeek V4 Pro, and the match is almost exact: $1.32 in and $3.96 out against Mistral's $1.36 and $4.18. Both labs compete on cost, and both are a long way below the American mid-tier price list.
- Score. Mistral leads by 2 points on the independent index, 38 against 36. On Mistral's own coding test, its Coding Agent Index score of 49.8% is ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. That is Mistral's number, not an independent one.
- Cost per task. DeepSeek wins clearly, at $0.67 against $1.13. DeepSeek also used fewer tokens to finish the index: 160 million against Mistral's 200 million.
- Agent work. Mistral says it scores 59.9% on AutomationBench, a set of 657 business workflows across apps such as Gmail, Google Sheets, Slack and Salesforce, ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro. Again, Mistral's figure.
- Where the model runs. This is the real difference for many buyers. Mistral trains and serves from its own European datacenters and offers a European deployment it operates end to end. For a business whose customers, regulator or bank asks where its data goes, that answer is easier to give.
Our view: on the numbers available today, DeepSeek V4 Pro does similar work for about 40% less per task. Mistral earns its place on where it runs and who controls it, not on price.
Where Mistral Large 4 is genuinely strong
An average score hides specialities, and Mistral is making specific claims. These are the ones worth knowing, with whose number each one is.
- Security work. Mistral says ML4 ranks in the top five on the Artificial Analysis Cyber Index, and scores 82% on one of its tests, reproducing a real software flaw and then patching it, the highest of any model. Mistral notes that Claude Opus 5.5 and GPT-6 Astra score near zero on that test because they refuse the task. If your IT partner does vulnerability testing and keeps hitting refusals, this is the model to try.
- Reading images and drawings. Mistral says ML4 slightly beats GPT-6 Astra on Dense 200, a visual grounding test, at 42% against 41%. It shows the model checking parts in engineering drawings and searching large satellite images.
- Finance and legal. Mistral says independent evaluator vals.ai found ML4 ahead of GPT-6 Astra on representative legal and finance tasks, and that it beats every open-source model on Harvey's Legal Agent benchmark. Mistral has not published those scores in the announcement text, so treat this as a claim to test.
- Resisting hijacked instructions. On Lakera's public B3 benchmark, Mistral says ML4 resists 93.3% of attacks. That matters for any agent that reads emails, web pages or documents written by strangers.
- Human judges on code. In a blind evaluation run with Surge AI, professional reviewers ranked ML4 second of five models at 3.74 out of 5, behind only Claude Opus 5 at 4.22.
Open weights sound free. A trillion parameters is not
The open weights are the headline for many readers, and they are a real benefit: you can run the model where you choose, keep data inside your own walls, and nobody can switch it off or change it under you. But be realistic about the size.
- The arithmetic. A trillion parameters stored at 8 bits each is roughly a terabyte of memory just to hold the model, and still about half a terabyte at 4 bits. That is our arithmetic, not a Mistral figure, and Mistral has not yet published hardware requirements. Either way, it is a multi-GPU server, not a box under a desk.
- What that means for most businesses. You will not host ML4 yourself. You will use it through Mistral's API or a cloud provider that hosts the weights. Open weights still help, because they create competition between hosts and let a regulated client insist on a specific location.
- The more useful part may come later. Mistral says ML4 will be the base for "a new generation of specialized and optimized" models. Smaller models distilled from a strong base are what a business can actually run on its own hardware.
- The licence is not out yet. The announcement promises open weights but does not state the licence terms. Check them before you plan on commercial self-hosting.
For Gulf businesses there is a local thread worth watching. In August, Mistral announced a collaboration with Saudi Arabia's HUMAIN "in the hundreds of millions of Euros", covering local datacenters, sovereign AI for regulated industries, and plans for frontier models that perform strongly in Arabic. Initial focus areas are cybersecurity and voice. None of that is a product you can buy today. It does make Mistral a name to know if your clients are banks, government bodies or telecoms in the Kingdom.
Who should use Mistral Large 4, and who should wait
Here is how we would decide, for a small or medium business in Pakistan or the Gulf.
- Try it now if you do security work and other models refuse it, you need a European data location for a client or regulator, or your job is reading technical drawings and dense images. Run your own tasks through it and measure the cost per job, not the price per token.
- Wait for the weights and the re-test if your interest is self-hosting or data control. The licence, the hardware guidance and a re-tested score should all arrive within weeks.
- Stay where you are if you are choosing a general model for customer replies, document summaries or bookkeeping help. On today's independent numbers, GPT-6.1 Sol, Claude Sonnet 5.5 and DeepSeek V4.1 Flash all do comparable or better work for less per task.
- Whatever you pick, cap the output. A verbose model is a billing risk. Set a maximum output length per request and a monthly spending cap in the provider's console, and log the cost of each job. That advice applies to every model in the table above.
We will refresh this comparison when the weights are released and Artificial Analysis re-tests the finished model. Mistral says the model is still improving. If its verbosity drops, the cost per task drops with it, and the verdict could change.
Your next step: choose by cost per job, not by launch day
Every launch week brings a new "cheapest" model. The one that is cheapest for you is the one that finishes your jobs in the fewest tokens at an acceptable quality. We set out how to measure that in the cheapest AI model for a business, by cost per task. If owning your model appeals, read what open-weight models really cost a business before you buy hardware. For the two mid-tier models Mistral is now competing with, see Claude Sonnet 5.5 vs GPT-6.1 Sol.
If you would rather use AI than manage models, that is what our products are for. IO Snack Accounts uses AI to categorise entries and flag anomalies as you post them, and AI Cam uses AI to review what your cameras see and alert you only when something needs attention. If you are weighing a bigger project and want a second opinion on which model fits it, talk to us.