AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 And AI: A Strong Global Option With An Agent Catch on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a sharp gain from the company’s earlier models and the highest result in the source’s framing for a model outside the United States and China. The supplied analysis says it remains behind major US and Chinese models and flags high task costs, verbosity and reported hallucinations as concerns for agent workflows. Its weights and licence have not yet been released.

Mistral AI has released Mistral Large 4 as a research public preview, scoring 38.4 on the Artificial Analysis Intelligence Index v4.3.2. The result is a substantial jump from Mistral’s previous models and makes the French company a notable option outside the United States and China, but the supplied analysis says leading US and Chinese systems score higher and raises concerns about the model’s cost and reliability on agent tasks.

Artificial Analysis’s current index places Large 4 below the US and Chinese models listed in the source’s comparison. Its score of 38.4 trails the index leader, Claude Opus 5.5, at 57.6, and sits below Chinese models including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. The source says the result is still the highest for a model outside the US and China, a distinction that does not mean Large 4 leads the broader frontier-model field.

The source reports that Large 4 has one trillion total parameters, with 49 billion active, and accepts text and images while producing text. It has a 512,000-token context window. The current release is a proprietary preview through Mistral’s API; the company has promised weights for the end of October, but the licence has not been published. Listed pricing is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source says Mistral is offering a 50% discount for the first two weeks.

The supplied analysis reports a cost of $1.13 per Artificial Analysis benchmark task. It compares that with $0.25 for GLM-5.3-Flash, which scores 41.8, and $0.27 for DeepSeek V4.1 Flash, which scores 39.5. Those figures suggest that, on this benchmark and at the stated standard prices, the two listed alternatives cost less per task while scoring higher. They do not by themselves establish which model is cheaper for a particular company’s workload.

At a glance
analysisWhen: Released yesterday, according to the so…
The developmentMistral has released Large 4 as a research public preview, prompting a new comparison of its benchmark performance, price and suitability for agentic tasks.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Trade-Off for Agent Work

Large 4’s result matters because the Artificial Analysis index includes tasks intended to measure agentic work, including knowledge work, SaaS workflows and coding. The source argues that capability gaps can compound over multiple steps: an error early in a long workflow may shape what an agent does next. That is an interpretation of benchmark performance, not a guarantee that Large 4 will fail on every extended task.

Cost and output length also affect deployment decisions. The source reports that Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. If that pattern carries over to a buyer’s own tasks, greater output volume could add latency and expense beyond the per-token price. Actual results will depend on prompts, task design, usage and provider pricing.

The source author also describes seeing confident hallucinations during hands-on testing. That observation is not an Artificial Analysis benchmark result and is not independently substantiated in the supplied material. It still points to a practical concern: a fabricated claim in an agent workflow can affect later steps. Buyers would need to test the model against their own accuracy, verification and cost requirements before assigning it consequential work.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Jump From Mistral’s Earlier Models

The source compares Large 4 with earlier Mistral scores on the same version of the Artificial Analysis index: Large 3 scored 9, while Medium 3.5 scored 14. Large 4’s 38.4 represents a major improvement in that comparison. The source characterizes it as the largest step by a European lab, but that broader superlative is the source author’s assessment rather than a result established by the figures provided here.

The launch headline, as relayed by the source, describes Mistral as home to the most intelligent model outside the US and China. The analysis cautions that this framing relies on a narrow competitive set. In its ranking, Large 4 trails several US and Chinese models, including newer systems than the Chinese models Mistral reportedly highlighted in its launch comparisons. The distinction is relevant for European buyers seeking another provider, but it should not be confused with a lead over the leading global models.

“Reinforcement learning is still running.”

— Mistral AI, according to the supplied source

Amazon

large language model development tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Limits and Testing Gaps

Large 4 remains a preview, and the source says its weights are not due until the end of October. The licence, which will shape how organizations may use or modify those weights, is unpublished. The exact release date, final licence terms and whether the public weights will match the API preview are not established in the supplied material.

The source also does not provide the full benchmark methodology, task-level results or independent replication of its hands-on hallucination observations. Artificial Analysis scores measure performance on a defined test set; they do not settle how the model will behave in every company’s workflows. Mistral’s statement that reinforcement learning is ongoing means scores may change, but the size and timing of any change are unknown. The claimed two-week introductory discount also needs a precise start date to determine when it ends.

Amazon

AI model cost analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licence and Updated Scores

The next milestones identified in the source are the promised end-of-October weight release and publication of the licence terms. Those details will clarify whether developers can run or adapt the model outside Mistral’s API and under what conditions. Artificial Analysis may also update its comparison if Mistral’s ongoing reinforcement learning changes measured performance.

For prospective users, the immediate next step is practical evaluation rather than reliance on a single index score. Teams considering Large 4 for agents can compare it with alternatives on representative tasks, tracking successful completion, errors that propagate between steps, output-token use, latency and total cost. The source offers no evidence yet that the benchmark cost or the author’s hallucination observations will reproduce in every deployment.

Amazon

AI model performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is Mistral AI’s newly released model, available as a research public preview through the company’s API. The source describes it as a text-and-image input model with text output and a 512,000-token context window.

How did it score against other models?

Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. In the supplied comparison, that is below the listed US frontier models and several Chinese models, including GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash.

Can developers download its weights now?

Not according to the supplied source. The current preview is proprietary and served through Mistral’s API; the company has promised weights for the end of October. The licence has not been published.

Is Large 4 a good choice for AI agents?

The supplied analysis raises concerns about benchmark performance, output volume, per-task cost and reported hallucinations, but it does not establish that the model is unsuitable for every agent workflow. Organizations should test it on their own tasks and compare total cost and reliability with alternatives.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

HBM Ate the Fab

High Bandwidth Memory (HBM) now accounts for nearly half of DRAM revenue, driving a global memory shortage and impacting GPU supply for 2026.

The System Behind AI’s Billion-Dollar Funding Surge

An analysis of how AI companies are raising billions through complex financial structures, revealing the scale and risks of the current funding surge.

The Power Bottleneck: AI Data Centers and the Grid Cliff Approaching 2027-2028

Power availability is constraining AI data center expansion, with grid upgrades lagging behind hyperscaler investments, risking a capacity shortfall by 2028.

The Impact Of Grok Bot On The Future Of Artificial Intelligence

SpaceXAI reportedly launches Grok Bot, an AI agent, raising questions about capabilities, release, and industry implications. Details remain limited.