🔍 Read the full analysis: Mistral Large 4’S Global Edge Comes With A Caveat For AI Agents on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index, a sharp rise from its predecessor but below current leading US and Chinese models. The source report says its cost per benchmark task, high output volume and observed confident errors may make it a poor fit for some long-running AI agents. Mistral says the model remains in public preview and is still undergoing reinforcement learning, so results could change.
Mistral has released Large 4, a multimodal model now available through its API in Research Public Preview, but independent benchmark data cited by ThorstenMeyerAI.com puts it below leading US and Chinese models and raises concerns about using it for AI agents. It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2; the report also flags higher benchmark-task costs and unusually high output volume.
Artificial Analysis’s index places Large 4 behind six listed US frontier models, whose scores range from 52.6 to 57.6, and several Chinese models. The source report describes it as the most intelligent model from outside the US and China, a characterization based on the benchmark results and the comparison set. Its score is close to OpenAI’s smaller GPT-6 Luna, listed at about 38, and below Chinese models including GLM-5.3-Flash at 41.8 and DeepSeek V4.1 Flash at 39.5.
The model has one trillion total parameters, with 49 billion active, accepts text and images, produces text, and has a 512,000-token context window, according to the source. Mistral offers it through its API as a Research Public Preview; the company has promised to release model weights by the end of October. The report says the licence had not been published at the time it was written.
At standard rates, the reported price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens. A 50% discount applies for the first two weeks, according to the source. Artificial Analysis data cited in the report puts Large 4 at $1.13 per benchmark task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those models also scored higher on the index. The task-cost figures describe the benchmark, not every customer’s workload.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The trade-offs for AI agents
The index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. That makes the score relevant to buyers considering models for multi-step work, software tasks and business workflows, rather than only short question-and-answer exchanges. A weaker result does not by itself predict how often a particular deployed agent will fail, but it is a reason to test performance on the intended tasks.
The report also says Large 4 generated 200 million output tokens across the benchmark, against a median of 81 million for comparable models. If a workflow produces similarly lengthy outputs, token use could add cost and time across repeated agent steps. Actual consumption will depend on prompts, tools, task design and deployment settings, so the benchmark figure should not be treated as a universal operating cost.
ThorstenMeyerAI.com’s author separately reports seeing confident false statements in hands-on testing. That is an attributed observation, not a result from Artificial Analysis, and the source provides no test protocol or sample size. Still, reliability matters in agent systems because later actions can build on earlier outputs. Buyers need to evaluate factual accuracy and error handling alongside benchmark scores and price.
As an affiliate, we earn on qualifying purchases.
A sharp rise from Mistral’s predecessor
On the same version of the Artificial Analysis index, the source reports that Mistral Large 3 scored 9 and Medium 3.5 scored 14, compared with Large 4’s 38.4. The increase is substantial within that benchmark, even though Large 4 remains well behind the top entries. Scores from different index versions may not be directly comparable; the report says the listed model figures use version 4.3.2.
The source frames Large 4 as a European alternative in a field where the strongest listed models come mainly from the United States and China. Its results support a narrower point: among models outside those two countries in the cited comparison, it has a high score. That does not make it competitive with the overall leaders, nor does it establish how it performs across every product or use case.
Mistral says reinforcement learning is still underway and that scores may change. The preview status, pending weights and unpublished licence described in the report also mean the release is not yet a settled basis for every deployment decision.
“Reinforcement learning is still running, so scores may move.”
— Mistral, as reported by ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
Preview results and reliability limits
Large 4 is still in public preview, and Mistral says further reinforcement learning could change its results. The source does not provide a release date for the promised weights beyond “the end of October,” and says the licence had not yet been published. The report also does not establish whether the benchmark’s token use or task costs will match customers’ own workloads.
The reported hallucination concern is based on the author’s hands-on testing, with no details in the supplied material about the number of prompts, task types or comparison method. Artificial Analysis’s figures and the author’s observation should be treated as separate kinds of evidence. It remains unclear how reliably Large 4 performs in specific agent settings, how it compares on an individual customer’s tasks, and whether later model updates will alter the trade-offs.
As an affiliate, we earn on qualifying purchases.
Weights, licensing and further testing
The next stated milestone is Mistral’s planned release of Large 4’s weights by the end of October. Buyers will also need the licence terms, which the source says had not been published, before judging whether and how the model can be deployed outside Mistral’s API. Mistral has said training is continuing, so updated benchmark results may follow.
For organizations evaluating the preview now, the source’s figures point to practical checks: run the model on representative workflows, track output-token use, and test whether it can identify uncertainty rather than confidently repeating a false premise. Those checks would help establish whether Large 4’s performance and costs suit a particular application; the benchmark alone cannot settle that question.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4?
Mistral Large 4 is a multimodal model available through Mistral’s API in Research Public Preview. The source describes it as a one-trillion-parameter model with 49 billion active parameters, a 512,000-token context window, and text-and-image input.
How did Large 4 score against leading models?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the report. The listed US leaders scored from 52.6 to 57.6, while several Chinese models also scored above Large 4.
Why does the report raise concerns about AI agents?
The index includes agent-focused evaluations, and the report says Large 4 used more output tokens per benchmark run than the comparable-model median. Its author also reports seeing confident false claims in hands-on testing. Those observations warrant task-specific evaluation but do not establish how the model will perform in every deployment.
How much does Mistral Large 4 cost?
The source lists standard API pricing at $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. It reports a 50% discount for the first two weeks; prices and discounts may depend on timing and service terms.
When will the model weights be available?
Mistral has promised to release the weights by the end of October, according to the source. The source does not provide a more specific date and says the licence was unpublished at the time of its report.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
