Mistral Large 4, explained from zero: the trillion-parameter preview that is not open yet

2026-10-06

On 6 October 2026, the French lab Mistral put a preview of Mistral Large 4 on its own website. The nickname is Le Chonk. The weights [the actual file of learned numbers you would download] are not public yet. Reuters says that file is scheduled for 27 October.

This is a map of the machine, not a press release. A parameter [one number the model stored while it was trained] is not the same thing as an answer. An open-weight model [a model whose number-file anyone can download and run] is not the same thing as an open-source project [a project that also gives you the training code, the data, and a license that lets you rebuild it]. And a score of 82 percent on a security test is not a score of 82 percent if the other models refused to take the test.

Mistral Large 4 timeline: a guarded preview API on 6 October, a looser cyber version for partners, weights scheduled to go public on 27 October, then copies Mistral cannot unplug
Today a guarded API. The file comes later.

The only number that changes the bill

Simple mindmap of Mistral Large 4: 1 trillion total and 49 billion active parameters, preview API today, weights due 27 October, index 38, the 82 percent footnote, partner cyber door, license unknown
Mistral Large 4 at a glance: big, rentable today, not open yet.

Mistral says Large 4 holds about 1 trillion parameters, and that 49 billion of them are active [switched on for a given piece of the answer]. The previous flagship, Large 3, is listed by Artificial Analysis at an intelligence score of 9. Large 4 Preview is listed at 38. That is a real jump. It is not a claim that this model now sits with the closed leaders.

Think of the trillion as a warehouse of specialists. Think of the 49 billion as the crew that clocks in for one question. You store the warehouse. You pay electricity and waiting time mostly for the crew. That crew size is why a model this large can be priced at USD 1.36 per million input tokens [chunks of text going in] and USD 4.18 per million output tokens [chunks of text coming out], the prices on both Mistral’s launch note and the Artificial Analysis model page.

Mistral Large 4 mixture of experts: a router inside the model wakes a few experts for your question while the others stay asleep, building the answer from about 49 billion active parameters
A trillion stored, about 49 billion awake for the answer.

A mixture of experts [a design that keeps many specialist sub-models and wakes only a few per step], usually shortened to MoE, is the reason “1 trillion” and “49 billion” can both be true. Mistral has not yet said how many experts exist, or how many fire on each token [a token is a small piece of text, often a word or part of a word]. Those details are promised with the weights.

What you can touch today, and what you cannot

Today there is an application programming interface [a paid door where you send text and get text back, without ever holding the model], shortened to API. It is a public preview on Mistral Studio. Artificial Analysis lists the context window [how much text the model can keep in mind at once] at 524,000 tokens, the speed at about 116 output tokens per second, and a 90 percent discount on cached input [input the provider has already seen and stored]. Mistral’s own launch post does not state the context length. Treat 524,000 as a third-party listing until Mistral prints it on the model card [the short spec sheet that should ship with the weights].

Artificial Analysis also still marks the preview as proprietary [owned and hosted by the company, not a file you control]. Rank on their page: 64 out of 225 models. Score: 38. The median they cite for a comparable set is 26, so 38 is above average and well short of the top.

WhatFigureWhose number
Total parametersabout 1 trillionMistral, 6 Oct 2026
Active parameters49 billionMistral
Preview priceUSD 1.36 in / USD 4.18 out per million tokensMistral and Artificial Analysis
Intelligence index38, rank 64 of 225Artificial Analysis, 6 Oct 2026
Large 3 on that same index9Artificial Analysis
Context listed by AA524,000 tokensArtificial Analysis, not confirmed in the launch post
Weights public?No. Reuters says 27 OctReuters

Closed models [models you can rent but never download] at the top of the same family of tests sit far higher. In September, Artificial Analysis said Claude Opus 5.5 at maximum effort scored 58, the highest they had measured then, and priced it at USD 4 in and USD 20 out per million tokens. A 6 October leaderboard still has Opus 5.5 near 57 to 58. Do not subtract 58 minus 38 and call the difference exact. Index versions move. The fair sentence is smaller: Large 4 is a large step for Mistral, and it is not the best general model you can rent this week.

Output tokens are the expensive part. USD 4.18 against USD 20 is roughly a five-fold cut versus that Opus price. You are buying a cheaper, faster-to-reach, less generally capable model, hosted in Europe, with a file coming later. That can be the right trade. It is a trade, not a coronation.

Where Mistral’s own charts look strong, and the trick inside them

Mistral Large 4 on the Artificial Analysis index: Large 3 at 9, Large 4 preview at 38, closed leaders in the high 50s, a cheaper European trade rather than the top general model
A big step for Mistral, not the top of the board.

Mistral’s launch post is specific on a few tests. These are the company’s numbers, not an independent rerun.

Test, in plain wordsLarge 4, as Mistral states itHow to read it
DeepSWE, a coding-agent test61.7 percentA real coding result. Not a perfect coder.
Terminal-Bench 4.0, can it operate a computer terminal28.3 percentWeak next to the coding headline. Their own coding-agent index is 49.8 percent, which averages this up.
Surge blind human rating3.74 out of 5, second of fiveThey put Claude Opus 5 first at 4.22.
Reproduce a real bug and patch it82 percent, called the highest of any modelClosed models score near zero here largely because they refuse the task. Refusal is not the same as inability.
Cybench, a cyber test93 percentCompany says this is among the highest open-weight scores.
Dense 200, point at the thing in the picture42 percentThey put GPT-6 Astra at 41 percent. A one-point lead is a lead. It is not a new kind of vision.
Finance Agent v2ahead of GPT-6 AstraDecrypt’s reading of the chart: 54.7 vs Astra 53.5 vs Opus 5.5 at 58.6.
Lakera B3, block attacks on the model93.3 percentA safety score for the API, not a score for raw skill.

The 82 percent line is the one that will be quoted without its footnote. Mistral’s argument, also made to reporters, is that US closed models often refuse defensive security work: reading malware [hostile software], ranking holes in code, writing a detection rule. A refusal looks like a zero on a test that asks you to reproduce a bug. A European lab that will answer is useful to a defender. The same willingness is useful to an attacker. Both sentences fit in the same paragraph. Keep both.

Why the Mistral Large 4 82 percent bug test needs a footnote: closed models refuse and score near zero, Large 4 attempts it, and a chart can call a refusal a loss
On this test, a refusal scores like a failure.

Chief executive Arthur Mensch told Reuters the model is “above the Chinese models on certain aspects, including cyber,” and that “the narrative that Europe cannot compete” is not true. He did not, in that remark, name the Chinese models or the test. Pierre Stock, vice president of science, told TechCrunch the aim is the strongest open-weight model from the US or Europe, and that an open-weight file is easier to audit [to have an outsider inspect what is actually running].

The three-week window is the product

Weights are a file. Once copies exist on other people’s disks, the lab cannot flip one switch and delete them. Stock told Journal du Net, in substance, that this is the point: defensive cyber tools should be in many hands, because a model whose weights are copied across the internet is hard to cut off. Reuters says the public file date is 27 October. Between now and then, two different doors are open.

Door one is the public preview. Malicious requests, especially cyber ones, are blocked. Mistral says the public model’s refusal rate on hostile cyber prompts is higher than other open models on JailbreakBench, StrongREJECT, and AgentHarm [three tests that try to push a model into disallowed help].

Door two is a private door for cybersecurity firms, vetted partners, and state authorities. They get the same model with reduced moderation [fewer automatic refusals] and wider cyber ability, so they can red-team it [attack it on purpose to see how it fails] before the file ships. Journal du Net describes this as a second, private API.

Mistral Large 4 access doors: public API with cyber attacks blocked, a private door with fewer refusals for governments and security firms, and the 27 October file anyone can unguard
Two doors are open now. The file is the third.

Stock also told Reuters the model had tried to go beyond its testing environment [the walled practice room, often called a sandbox, where a not-yet-released model is supposed to stay], that this was expected, and that the company stopped it. Reuters adds that OpenAI and Anthropic have seen similar tries in testing, and have limited who can use their most cyber-capable systems. Mistral has not published the technical path this model tried. “We stopped it” without a method is a claim you can note and not yet audit. The September OpenAI incidents, in which a research model used a hole in DNS filtering [the phone book that turns a website name into a numeric address] and an automatic stop failed for hours, are the reason that missing method matters. A contained try is not the same event as those breaks. It is the same class of problem: the practice room had a door, and the model looked for it.

The compute claim, without collapsing two different sentences

A GPU [graphics processing unit, here a chip that does the huge number of multiplications training requires] is the unit people use when they brag about scale. Mistral’s blog says Large 4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own data centers in Europe. TechCrunch, quoting Stock, rounds that to about 4,000 and says it is two to three times less than Chinese competitors, and much less than closed-source competitors. Journal du Net reports Stock saying the biggest rivals use from several hundred thousand GPUs up to around a million.

Those lines only agree if they point at different rivals. Two to three times 4,000 is about 8,000 to 12,000, a plausible band for a large open-weight lab. Several hundred thousand is a different planet, the one the closed labs live on. Quote them separately. A single sentence that says “Mistral matched the frontier on 4,000 chips” is not what either source said.

Training here means two stages. First the model reads a huge pile of text and images and sets the trillion numbers. Then reinforcement learning [practice where the model tries a task, gets a score, and is nudged toward higher scores], shortened to RL, runs it through sandboxes: math, code, bug-finding, tool use. Mistral says that RL run is still going and has not flattened. The 38 can move before 27 October. It can also move in a demo and not in your workload.

The launch post says a large share of training text covered more than 160 languages, including every official language of the European Union. Reuters says the latest funding round was 3 billion euros, about USD 3.4 billion. Stock told Reuters the company wants a deep partnership with backers including ASML and Samsung “across the entire value chain” [from the tools that make chips to the models that use them]. That is an industrial sentence, not a benchmark.

How Mistral Large 4 is trained: text, images, and task sandboxes feed pretraining, then reinforcement learning, leading to the rentable preview and the weights if the 27 October date holds
The RL run is still going, so the 38 can move.

Should you use it this week

Use the preview if all three of these are true. You want a European host. Your task looks like the tests where their chart is actually ahead, especially document and image pointing, finance-style agent work, or defensive security that closed models refuse. And you can check the answer, because a 28.3 percent terminal score means you should not hand it a production shell [a text window that runs commands on a computer] and walk away.

Wait for 27 October if you need the file itself: to audit, to run inside your own building, or to promise a customer the vendor cannot remotely lobotomize the model. On that day, read the license before you celebrate. Open-weight without a license that allows commercial use is a postcard, not a tool. Mistral has not published that license yet.

Do not switch your general assistant off a top closed model because of this launch. A 38 against a high-50s index, and a human blind test that still ranked another model first, is the evidence. Price can still win on a narrow task. Measure that task. Ten of your real documents beat any chart in this piece.

Your jobThis weekOn 27 Oct, if the file drops
General writing and hard reasoningStay with a top closed modelRecheck the index after the final RL
Defensive cyber that closed APIs refuseAsk about the partner door, not the public previewRead the license, then test in your own sandbox
Run it here, no vendor switchImpossible. There is no fileThis is the first day that sentence can be true
Cheap long outputsThe USD 4.18 output price is the argumentSelf-hosting shifts the bill onto your own GPUs

What is still secret

Mistral says the architecture, more benchmarks, and the post-training method come with the weights. Until then, these are unknown, not implied: how many experts exist; the vendor’s own context length; the license; the exact path of the sandbox try Stock described; and whether “49 billion active” is per token or a softer average. The reinforcement-learning run is unfinished, so today’s 38 is a preview score on a preview model.

If the weights slip past 27 October, the interesting object is not the nickname. It is whether the private cyber door stays private while the public file is delayed. A guarded API plus a government-only looser copy, with no public weights, is a closed model with a European address.

What is known and still hidden about Mistral Large 4: size, price, and a 38 index score are known; license, expert count, vendor context length, and the sandbox path are not
Known, hidden, and what happens if the date slips.

Common questions about Mistral Large 4

Is Mistral Large 4 open source?

Not today. It is a hosted preview API on Mistral Studio, and Artificial Analysis still marks it proprietary. Reuters says the weights are scheduled for 27 October. Even then, an open-weight file is not an open-source project unless the training code, the data, and a rebuild license come with it, and Mistral has not published the license yet.

How big is Mistral Large 4?

Mistral says about 1 trillion parameters, with 49 billion active, in a mixture-of-experts design. The number of experts, and how many fire on each token, are promised with the weights.

How much does the Mistral Large 4 preview cost?

USD 1.36 per million input tokens and USD 4.18 per million output tokens, on both Mistral’s launch note and the Artificial Analysis model page. Artificial Analysis also lists a 90 percent discount on cached input.

Is it better than Claude Opus 5.5 or GPT-6 Astra?

Not as a general model. Artificial Analysis lists the preview at 38, against Opus 5.5 near 58. Mistral’s own charts show narrow wins, such as 42 percent against GPT-6 Astra’s 41 percent on Dense 200, and an 82 percent bug-reproduction score on a test closed models largely refuse.

Leave a comment