Could Mistral Large 4 Lead Outside The US And China? Its Agent Shortfall
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Could Mistral Large 4 Lead Outside The US And China? Its Agent Shortfall on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a major improvement over Mistral’s earlier models but below current US and several Chinese systems. The source argues that weaker agent-task performance, high output volume and benchmark-task costs make it a difficult choice for sustained workflows; its findings on hallucinations come from the author’s hands-on testing, not the benchmark.

Mistral released Large 4 as a research preview, and the model scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2, according to the cited benchmark data. That result makes it a strong showing for a model from outside the United States and China, but it remains below the listed US frontier models and several Chinese competitors, raising questions about its fit for demanding agent workflows.

The model is described as having 1 trillion total parameters, with 49 billion active, and supporting text and image input with text output. Its 512,000-token context window is another stated specification. Mistral made Large 4 available through its API as a research preview. The source says the company plans to release the weights at the end of October; until then, the model is proprietary, and its licence has not been published.

On Artificial Analysis Index v4.3.2, Large 4 scored 38.4. The same source lists Claude Opus 5.5 at 57.6, Google’s Gemini 4 Argon at 52.6, and Z.ai’s GLM-5.3 at 44.8. It also places two Chinese models, GLM-5.3-Flash at 41.8 and DeepSeek V4.1 Flash at 39.5, ahead of Mistral. The source reports Large 3 at 9 and Medium 3.5 at 14 on the same index version, marking a substantial improvement within Mistral’s lineup.

The listed API price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source reports a 50% discount for the first two weeks. It also says Artificial Analysis measured $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Mistral says reinforcement learning is continuing, so the model’s benchmark performance may change.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral has released Large 4 as a research preview, with independent benchmark results placing it below leading US and several Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost of Long Agent Runs

The result matters most to buyers choosing models for multi-step work, not just short exchanges. The source notes that the Artificial Analysis index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. A lower score across those tests may signal a larger practical gap on lengthy tasks, though the benchmark does not by itself predict how any particular deployment will perform.

The source also reports that Large 4 generated 200 million output tokens across the index tasks, against a median of 81 million for comparable models. If that pattern carries into a customer’s workload, additional output can add cost and latency even when the price per token appears competitive. The reported task-cost comparisons point in the same direction: two listed Chinese models score higher while costing less per benchmark task. These figures make Large 4’s value proposition less clear for organizations weighing price against agent capability.

There is a separate reliability concern, but it should be treated as an observation rather than a measured benchmark result. The source author says hands-on testing found confident false statements. In an agent workflow, an unsupported claim can become an assumption that later steps rely on. The report does not provide a reproducible test protocol or sample size for that observation, so it cannot establish how often the issue occurs.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Gain From Mistral’s Earlier Models

The benchmark comparison suggests Large 4 is a significant step up from Mistral’s recent models: the source gives Large 3 a score of 9 and Medium 3.5 a score of 14 on the same version of the index. That change supports the case that Mistral has made rapid progress, even though the current score still falls short of the leading models listed.

The headline that France now has the highest-scoring model outside the US and China depends on how the field is defined. The source calls it an accurate description, while arguing that few labs outside those two countries compete at this tier. Its comparisons also show that Large 4 would sit behind multiple Chinese open-weight models if its weights are released as planned. Because the weights are not yet available and the licence is unpublished, its open-weight status and practical terms remain unsettled.

The benchmark figures cited here are from Artificial Analysis Intelligence Index v4.3.2, which allows the source to compare the listed models on a common index version. The scores are not a complete measure of model quality: the source’s own cost analysis and hands-on reliability observations address different questions and should not be conflated with the index score.

“The headline from Artificial Analysis is the one Paris wanted: France is home to the most intelligent model from outside the United States and China.”

— ThorstenMeyerAI.com report author

Amazon

large language model API tokens

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results and Reliability Questions

Large 4 remains a research preview, and the source says Mistral is still applying reinforcement learning. It is not clear how much the score may shift, when the weights will be available beyond the stated end-of-October plan, or what licence will govern them. The source provides no independent confirmation of the planned release date.

The index and price figures do not settle how Large 4 will perform in every real-world agent deployment. The report’s hallucination concern is based on the author’s hands-on testing, with no detailed methodology supplied. Nor does the benchmark establish that the reported output-token volume will apply to a buyer’s own tasks. The model’s reliability, total operating cost and comparative value in specific workflows remain to be tested by users.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and New Scores

The next milestones are the promised weight release at the end of October and publication of a licence. Those details will determine whether developers can run the model outside Mistral’s API and under what conditions. Mistral’s continuing reinforcement learning may also produce revised benchmark results, although the company has not supplied a timeline for a new score.

For now, buyers can compare the preview’s published API pricing and index results with alternatives, while testing the model against their own tasks before relying on it for long-running workflows. Updated independent evaluations and clearer information about model access, licensing and reliability would help establish whether Large 4’s rapid improvement translates into a practical advantage.

Amazon

AI model performance analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source report. The source lists several US and Chinese models with higher scores.

Is Large 4 available as an open-weight model?

Not yet, according to the report. It is available through Mistral’s API as a research preview; Mistral has said weights are planned for the end of October, while the licence has not been published.

Why does the report question Large 4 for agent tasks?

The cited index includes agent-oriented evaluations, and Large 4 scores below several alternatives. The report also says it used more output tokens than the comparison median, which could raise cost and latency in some workflows.

Has Large 4 been shown to hallucinate more often?

The report author describes seeing confident false claims during hands-on testing, but gives no test protocol or rate. That observation is not an Artificial Analysis benchmark finding, so the frequency and severity are unclear.

Could the benchmark score change?

Possibly. The source says Mistral is continuing reinforcement learning, which may affect performance. It does not give a schedule for updated results.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The European Bet: How Mistral, Aleph Alpha, and Black Forest Labs Are Playing a Different Game

Mistral, Aleph Alpha, and Black Forest Labs are positioning themselves for the EU AI Act’s enforcement, emphasizing compliance and sovereignty over frontier capabilities.

How The Hugging Face Incident Shapes Our View Of AI Safety Protocols

OpenAI’s recent cybersecurity breach reveals risks of autonomous AI agents and underscores the need for robust safety protocols in AI development.

Fable 5 Is Back. GPT-5.6 Is Next. And Anthropic Reportedly Already Has Something Stronger.

Fable 5 is back after 18 days, GPT-5.6 is in preview, and rumors suggest Anthropic has an even more advanced model. What this means for AI development.

6 Best Desktop Processors for Gaming and Everyday Performance in 2026

Explore the best desktop processors in 2026 for gaming and everyday tasks, including AMD Ryzen 7 9700X, Ryzen 7 9800X3D, and more.