🔍 Read the full analysis: How Mistral Large 4 Measures Up Outside The US And China And In Agent Work on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index, a sharp rise from its predecessor but below leading US and Chinese models. The source report says it is expensive per benchmark task and unusually verbose, raising concerns about using it for long-running agents; those findings do not establish how it will perform across all deployments.
Mistral released Large 4 as a research preview, and the Artificial Analysis Intelligence Index v4.3.2 gives it a score of 38.4—higher than the cited scores for the company’s earlier models, but below leading US and Chinese systems. The result places France’s Mistral among the strongest model developers outside those two countries, while the benchmark and reported task costs raise questions about its suitability for long-running agent work.
Artificial Analysis lists Large 4 below US models including Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7, as well as Chinese models including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. These figures come from the same version of the index, according to the source report. Mistral’s score is also close to OpenAI’s smaller GPT-6 Luna, listed at about 38.
The model is described as having 1 trillion total parameters, with 49 billion active, text-and-image input, text output and a 512,000-token context window. Mistral offers it through its API as a research preview. The source says the company plans to release model weights at the end of October, but does not provide a year for that date. The licence had not been published at the time covered by the report.
The source report gives standard API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; it says a 50% discount applies for the first two weeks. Its benchmark-cost comparison estimates $1.13 per Intelligence Index task for Large 4, versus $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two Chinese models also score higher on the cited index. Mistral says reinforcement learning is continuing, so its scores may change.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Cost of Agent Tasks
The score matters in part because the Artificial Analysis index includes tasks that resemble multi-step knowledge work, software workflows and coding, not just short question-and-answer exchanges. A lower result does not directly tell buyers how often a particular product will fail, but it is relevant evidence when comparing models for agent systems that must complete sequences of actions.
The source report argues that capability gaps can compound over a long run: an early error may become the basis for later steps. It also reports that Large 4 generated 200 million output tokens across the index, compared with a median of 81 million for comparable models. If that difference carries over to a buyer’s workload, more generated tokens could add latency and cost. The benchmark observation is not a guarantee that every customer task will show the same pattern.
For organisations seeking a model from a European provider, Large 4 offers another option at the frontier end of the market. But the report’s figures do not support treating it as a performance leader: several US and Chinese models score higher, and two lower-cost Chinese models in the comparison outperform it while costing less per benchmark task. Procurement decisions will also depend on licensing, deployment choices, data controls and performance on a company’s own tasks.
As an affiliate, we earn on qualifying purchases.
A Jump From Mistral’s Earlier Scores
The report describes Large 4’s 38.4 index score as a marked improvement over Mistral Large 3, which scored 9 on the same index version, and Medium 3.5, listed at 14. That makes the release a significant move for Mistral in the cited benchmark, even though it does not close the gap to the highest-scoring systems.
The source frames the “outside the US and China” distinction carefully: Mistral is based in France, and the report calls it the strongest model in that geographic grouping. That is a narrower comparison than ranking against the entire market. The source argues that the field of serious competitors outside the US and China is limited; this characterization is the report author’s assessment, not a separate benchmark finding.
Large 4’s current status also matters. The source says it is a proprietary API preview until promised weights are released, with the licence still unpublished. Mistral’s statement that reinforcement learning is ongoing means the published preview results could change. Buyers therefore have benchmark evidence for the version tested, but not yet a settled picture of the eventual weights, terms or performance.
Preview Results and Open Questions
The evidence presented is a snapshot of Artificial Analysis Index v4.3.2 and the source author’s testing. The source does not provide enough detail to establish how representative the benchmark is of any one organisation’s workflows, or how Large 4 compares on a customer’s specific tasks. Its score may also change as Mistral continues reinforcement learning.
The source author reports seeing confident hallucinations in hands-on use, but supplies no measured Large 4 hallucination rate or test protocol. The cited hallucination figures for other models come from a separate Artificial Analysis measure and cannot be used to assign Large 4 a rate. The report’s token-count and cost comparisons likewise describe benchmark performance, not a guaranteed bill for a typical deployment.
Several commercial details remain unsettled in the source: the planned weight release is dated only as “the end of October,” the licence is unpublished, and it is not clear whether the initial preview pricing or discount will persist. The report’s cost comparison should also be read as its own estimate of cost per benchmark task, not as a universal price for equivalent customer work.
Weights, Licensing and Retesting
The next milestones identified in the source are Mistral’s expected release of Large 4’s weights at the end of October and publication of the licence terms. The source does not state a year for the release date, so the timing should be treated as the company’s reported plan rather than a dated guarantee. Until weights and terms are available, customers evaluating deployment options have only the API preview described in the report.
Further Artificial Analysis results could show whether the score changes as reinforcement learning continues. Buyers considering agent use will also need to test the model on their own multi-step tasks, measuring completion rates, factual errors, output volume, latency and total cost. Those tests can clarify whether Large 4’s improvement over Mistral’s earlier models offsets its lower benchmark score against several rivals.
Key Questions
What score did Mistral Large 4 receive?
It scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2, according to the source report. The score is higher than the cited results for Mistral Large 3 and Medium 3.5, but below several listed US and Chinese models.
Is Large 4 the top model in the benchmark?
No. The source’s table lists multiple US and Chinese models above it, including Claude Opus 5.5 at 57.6 and GLM-5.3 at 44.8. The “outside the US and China” description is a narrower geographic comparison.
Why does the report question its use for agents?
The author points to its lower index score than leading models, higher reported output-token use and a personal observation of confident false statements. The hallucination observation is not a measured Large 4 rate, and benchmark performance does not determine how it will behave in every deployed agent.
How much does Mistral Large 4 cost?
The source reports standard pricing of $1.36 per million input tokens and $4.18 per million output tokens, plus $0.14 per million cached input tokens. It also reports a 50% discount for the first two weeks and estimates $1.13 per Artificial Analysis task; actual costs depend on usage and pricing terms.
When will the model weights and licence be available?
The source says Mistral planned to release weights at the end of October, but gives no year. It says the licence was unpublished at the time of reporting, so both the release timing and terms need confirmation.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
