AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What To Know About Mistral Large 4 Before Running AI Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 has entered public preview with a 38.4 score on Artificial Analysis Intelligence Index v4.3.2, a sharp improvement over earlier Mistral models but below leading US and Chinese systems. The source says its measured task costs are higher than two Chinese models that score above it, while its weights and licence are not yet available.

Mistral has released Large 4 in public preview, presenting a substantial jump over its earlier models while the Artificial Analysis Intelligence Index v4.3.2 gives it a score of 38.4. The result makes it the highest-scoring model from outside the United States and China in the cited comparison, but it trails leading US and Chinese systems—an important distinction for teams considering it for AI agents and automated workflows.

Artificial Analysis’ current index places Large 4 below several US frontier models and Chinese models, including Z.ai’s GLM-5.3 at 44.8, Moonshot’s Kimi K3 at 43.6, and DeepSeek V4.1 Flash at 39.5. The leading listed score is Claude Opus 5.5’s 57.6. These figures are from the same index version, according to the source, but are benchmark results rather than guarantees of performance on any particular customer’s work.

The model has 1 trillion total parameters, with 49 billion active, and accepts text and images while producing text. Its context window is 512,000 tokens. Mistral has made it available through its API as a Research Public Preview; the source says the model’s weights are promised for the end of October. Until then, it is proprietary, and its licence has not been published.

Listed pricing is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source reports a 50% discount for the first two weeks. It also says Mistral has told users reinforcement learning is ongoing, meaning benchmark performance may change. Artificial Analysis estimates Large 4’s cost at $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; both of those models score above Large 4 in the cited results.

At a glance
reportWhen: Public Preview announced yesterday in t…
The developmentMistral released Large 4 in public preview, with independent benchmark results showing a major improvement over its predecessors but weaker scores than leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Runs Face Cost and Reliability Tests

The index cited in the source includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. That makes the score relevant to buyers weighing models for multi-step work, though it does not establish how a model will perform in every deployed agent or tool setup.

Long-running agents can magnify errors: a mistaken result early in a workflow may become an assumption for later steps. The source also reports that Large 4 generated 200 million output tokens across the index tasks, compared with a median of 81 million for comparable models. If that pattern carries into a customer’s workload, output volume could add to latency and operating costs, beyond the stated price per token.

The article’s author separately reports seeing confident false statements during hands-on testing. That is an individual observation, not a published Artificial Analysis measurement, and the source does not describe the testing protocol. It is still a reason for prospective users to test factual accuracy, tool calls, and recovery from errors before assigning the model consequential work.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Jump From Mistral’s Earlier Scores

The cited index gives Mistral Large 3 a score of 9 and Medium 3.5 a score of 14, compared with Large 4’s 38.4. Those results suggest a major improvement within Mistral’s own lineup. They do not erase the gap with the top-scoring systems: the source’s comparison puts Large 4 19.2 points below Claude Opus 5.5.

The launch framing cited by the source is that France is home to the most intelligent model outside the US and China. That description fits the listed national comparison, but it does not mean Large 4 leads the global field. The source says that, if its weights are released as promised, the model would rank eighth among open-weight models in this comparison, behind seven Chinese models. The ranking remains conditional on the weights becoming available and on the benchmark results used.

“Reinforcement learning is still running.”

— Mistral, as reported in the supplied source

Weights, Licence and Results Remain Open

The weights have not yet been released, and the source gives only an end-of-October target without specifying a year. Mistral’s licence terms are also unpublished in the material provided, so buyers cannot yet assess the conditions that would apply to using or modifying released weights.

It is not clear how much benchmark scores may move as reinforcement learning continues, or whether the reported token use and task costs will match real-world agent workloads. The source supplies no testing methodology for its hands-on hallucination observations. Nor does the supplied material include the full cost comparison promised at the end of its discussion, so no further price-to-performance conclusion should be inferred from it.

Watch for Weight Release and Retesting

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. Users will need to check whether that release occurs, what licence accompanies it, and whether the available model matches the preview version. Mistral’s ongoing reinforcement learning may also lead to revised results.

For teams evaluating the API now, the practical next step is a controlled trial using representative tasks. Track success rates across full workflows, factual errors, recovery behavior, output-token use and total task cost—not just the headline benchmark score or per-token price. Those results will show whether Large 4’s improvement over earlier Mistral systems suits a particular use case.

Key Questions

Is Mistral Large 4 available now?

According to the supplied source, it is available through Mistral’s API as a Research Public Preview. Its weights have not yet been released.

How did Large 4 score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The cited table places several US and Chinese models above it, including GLM-5.3 at 44.8 and Claude Opus 5.5 at 57.6.

What will Large 4 cost?

The listed standard rates are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source reports a 50% discount for the first two weeks; actual costs depend on usage and output volume.

When are the model weights expected?

Mistral is reported to have promised the weights for the end of October. The supplied material does not specify the year, and the licence has not yet been published.

Should teams use it for AI agents?

The benchmark results alone do not settle that question. Teams should test Large 4 on their own multi-step tasks, including error rates, tool use, output volume and total cost, before relying on it for production workflows.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Preparing for Agentic Browsers: How AI Will Redefine Web Interactio

Explore how AI is transforming web interactions in our case study on Preparing for Agentic Browsers. Get ready for the future of online browsing.

Every Benchmark Launched 2023-2024 Has Fallen — The METR / SWE-Bench / CORE-Bench / MLE-Bench / PostTrainBench Sequence

Every major AI research benchmark launched in 2023-2024 has either saturated or is nearing saturation, signaling accelerated AI capability growth.

Federated Learning Infrastructure: Privacy‑Preserving Patterns

Privacy-preserving patterns in federated learning ensure secure, decentralized model training, but understanding how they balance privacy and accuracy requires further exploration.

Engineering Is Automated. Research Is the Residual.

Recent advances show AI has automated most engineering tasks in AI research, raising questions about the future of scientific innovation.