Mistral Large 4, the 1-trillion-parameter model Mistral calls "le Chonk", went into public preview on October 6, 2026. It is a mixture-of-experts model with 49 billion active parameters, reads text and images, writes text, and runs with a 1M-token context window. Mistral says the open weights "drop end of this month". Until then you can only reach it through the preview API in Mistral Studio.

The surprise is the price. On the Mistral Large 4 model page, the list price is $1.36 per million input tokens and $4.18 per million output tokens, and the preview is billed at half that: $0.68 and $2.09. Mistral's largest model costs less than its own Mistral Medium 3.5, and less than the GLM-5.3 model Mistral hosts for coding agents. The benchmarks are a different story: most of them come from Mistral, and the first independent results are more modest. Here is what shipped, what it costs next to the alternatives, and how to test it this week.

What Mistral shipped on October 6

The announcement and the model documentation give these specifications:

  • Size: 1.05 trillion total parameters, 49 billion active per token, plus a 1.6 billion parameter vision encoder.
  • Type: a hybrid instruct-and-reasoning model. One model handles both quick answers and long reasoning, switched per request.
  • Input and output: text and images in, text out. Mistral highlights documents, charts, PDFs, technical drawings and satellite imagery.
  • Context: 1M tokens, up from 256k on Mistral Large 3.
  • API name: mistral-large-4, version 26.10, status Public Preview.
  • Features: structured outputs, function calling, document Q&A, batching, and the Agents and Conversations endpoints with built-in tools.
  • Languages: trained on more than 160, including every official EU language.
  • Training: from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters.

For scale, Large 3, released in December 2025, has 675 billion total and 41 billion active parameters. Large 4 adds about 55% more total parameters but only 8 billion more active ones, which is why per-token serving cost did not balloon. TechCrunch quotes Mistral's VP of science, Pierre Stock, saying the training used "two to three times less" GPU capacity than Chinese competitors.

The price: Large 4 undercuts Mistral Medium 3.5

All five models below are served on the same platform, from Mistral's own docs pages, so the comparison is like for like. Prices are per million tokens.

Model on MistralInputCached inputOutputContextImage inputWeights
Mistral Large 4 (preview price)$0.68$0.07$2.091MYesEnd of October
Mistral Large 4 (list price)$1.36$0.14$4.181MYesEnd of October
Z.ai GLM-5.3 (hosted)$1.40$0.14$4.401MNo (text)Open
Mistral Medium 3.5$1.50Not listed$7.50256kYesModified MIT
Mistral Large 3$0.50Not listed$1.50256kYesApache 2.0
Mistral Small 4$0.15Not listed$0.60256kYesApache 2.0

Three things stand out. First, even at list price, Large 4 is cheaper than Medium 3.5 on both input and output, with four times the context. Second, at list price it sits within a few cents of GLM-5.3, the open coding model Mistral already hosts, and it adds image input that GLM-5.3 lacks there. Third, Large 3 is still cheaper on paper, so the case for switching rests on quality and the 1M context, not on price alone.

Mistral has not said when the preview discount ends. The docs show the list price struck through, which reads as a temporary offer, not a new rate. Budget at list price and treat the preview as a cheap testing window.

What a day of agent work costs on each model

To make the per-token prices concrete, take one day of a busy coding or research agent: 10 million input tokens and 1 million output tokens, no caching. This is our arithmetic from the prices above.

Model10M input1M outputDaily total
Mistral Small 4$1.50$0.60$2.10
Mistral Large 3$5.00$1.50$6.50
Mistral Large 4, preview$6.80$2.09$8.89
Mistral Large 4, list$13.60$4.18$17.78
GLM-5.3 on Mistral$14.00$4.40$18.40
Mistral Medium 3.5$15.00$7.50$22.50

Agents resend the same context many times, so caching matters. If 8 of those 10 million input tokens are cache hits, Large 4 at preview pricing drops to about $4.01 a day ($1.36 fresh input, $0.56 cached, $2.09 output), against about $8.32 for GLM-5.3 on the same terms. Two caveats: reasoning mode produces extra output tokens, so a high-effort run will cost more than this table, and batch jobs run at a 50% discount if your work can wait.

Daily cost of 10M input and 1M output tokens: Large 3 $6.50, Large 4 preview $8.89, GLM-5.3 $18.40, Medium 3.5 $22.50
One day of agent work (10M input, 1M output tokens): Large 3 $6.50, Large 4 at preview price $8.89, GLM-5.3 $18.40, Medium 3.5 $22.50.

How the benchmarks stack up, and who measured them

Mistral's own numbers are strong for an open-weight model. It reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4, and a combined Coding Agent Index of 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. On AutomationBench, 657 business workflows across apps like Gmail, Google Sheets, Slack and Salesforce, it scores 59.9%. On visual grounding it claims 42% on Dense 200, one point ahead of GPT-6 Astra at 41%.

The most useful number for builders is a blind human evaluation run with Surge AI on coding quality. Professional annotators scored outputs from 1 to 5 without knowing which model wrote them. Mistral Large 4 Preview came second of five at 3.74, behind Claude Opus 5 at 4.22 and ahead of GLM-5.3 (3.60), Kimi K3 (3.59) and GLM-5.2 (3.40). That is a fair summary of where it sits: the best open-weight option in that test, still clearly behind the top closed model.

Independent checks are thinner and cooler. Vals AI published results the same day: 48.05% on its Vals Index, rank 32 of 44, and 22.73% on Terminal-Bench 4.0, below Mistral's own 28.3%. Its best placement was 6th of 75 on Harvey's Legal Agent Benchmark. VentureBeat also notes that the live DeepSWE leaderboard shows higher scores for GLM-5.3, Kimi K3 and the top closed models under other configurations, and that Large 4 was not yet listed there at launch.

Read it this way: Large 4 is competitive with the leading open models from China on coding, ahead of them on vision by Mistral's tests, and not a replacement for a frontier closed model if raw coding quality is the only thing you pay for.

Mistral Large 4 on Terminal-Bench 4: 28.3% reported by Mistral, 22.73% measured by Vals AI
Same benchmark, two scores: Mistral reports 28.3% on Terminal-Bench 4, Vals AI measured 22.73%.

How to try Mistral Large 4 today

The preview is API only, so this takes a Mistral account and about 20 minutes. Run it against the model you use now, on your own prompts.

  1. Get a key. Sign in to Mistral Studio, create an API key, and export it as MISTRAL_API_KEY. The Studio playground lets you try prompts first without code.
  2. Point your client at the new model name. In the official Python or TypeScript SDK, or any OpenAI-style tool that supports Mistral, set the model to mistral-large-4.
  3. Choose a reasoning mode per request. Pass reasoning_effort="high" for agentic and coding work, or "none" for fast replies, as described in the Mistral reasoning docs.
  4. Update your response parsing. With "high", message.content becomes a list of chunks, a thinking chunk followed by a text chunk, not a plain string. Code that expects a string will break. With "none", you get a plain string back.
  5. Keep the full assistant message in multi-turn chats. Mistral's docs say to append the whole assistant message, thinking included, to the history. Dropping the reasoning trace between turns degrades performance.
  6. Test the vision input on your real files. Send a PDF page, a chart or a technical drawing you work with. Visual grounding is where Mistral claims the biggest lead.
  7. Run your regression set through batch. Batch processing is half price, which on top of the preview discount makes a large side-by-side test cheap.

One limit to know: the public endpoint runs with Mistral's moderation on. Mistral says cybersecurity leaders, vetted partners and state authorities get the same model "with reduced moderation and expanded cyber capabilities" during red-teaming. If your work touches security research, expect refusals on the public preview.

Mistral Large 4 reasoning_effort setting with two values: none for fast replies, high for full reasoning
One parameter switches modes: reasoning_effort "none" returns a plain string, "high" returns thinking and text chunks.

The open weights: what "end of the month" means for you

Mistral's post says "We will release the weights by the end of the month," and VentureBeat reports a date of October 27. Mistral has not published the license. Its recent releases have used different terms: Large 3 and Small 4 are Apache 2.0, Medium 3.5 shipped under a Modified MIT license, and VentureBeat describes the Large 4 terms as a custom Mistral license. Check the license file before building a product on the weights.

The weights will not run on a workstation. At 1.05 trillion parameters, 8-bit weights take roughly 1.05 TB of memory and 4-bit weights roughly 525 GB, before any context cache (our arithmetic). That is about the full 1,128 GB of an eight-GPU H200 server at 8-bit. For most creators and small teams, the open weights matter for a different reason: other hosts can serve the same model, which keeps prices competitive and gives you a way out if Mistral's terms change.

Who should switch now, and who should wait

Switch now if you already run Mistral Medium 3.5 or GLM-5.3 on Mistral's API. Large 4 at preview pricing costs less on every token, adds 1M context over Medium, and adds image input over GLM-5.3. The migration is one model name plus the response-parsing change for reasoning mode.

Test it if your work mixes documents and images with text: contracts, spec sheets, chart-heavy reports, engineering drawings. That is the clearest gap between Large 4 and other open models.

Wait if you are on Large 3 and happy with it, since Large 3 is still cheaper, or if your main use is hard coding where a closed model like Claude Opus 5 still scores higher in the blind test Mistral itself ran. And if your plan depends on self-hosting, wait for the weights and the license on October 27 before committing.

Frequently asked questions

What is Mistral Large 4?

Mistral Large 4, nicknamed "le Chonk", is Mistral's largest model: a 1.05 trillion parameter mixture-of-experts model with 49 billion active parameters, image and text input, and a 1M-token context window. It entered public preview on October 6, 2026.

How much does Mistral Large 4 cost?

The list price is $1.36 per million input tokens, $0.14 cached input and $4.18 output. During the public preview Mistral bills half: $0.68, $0.07 and $2.09. Mistral has not announced when the preview discount ends.

When will Mistral Large 4 open weights be released?

Mistral says by the end of October 2026, and VentureBeat reports October 27. The license has not been published yet.

Can I run Mistral Large 4 locally?

Not on a desktop or a single GPU. The weights need roughly 525 GB of memory at 4-bit and about 1.05 TB at 8-bit, which means a multi-GPU server. Most people will use the API or a third-party host.

Is Mistral Large 4 better than GLM-5.3 for coding?

Slightly, by Mistral's numbers: 61.7% on DeepSWE v1.1 and 3.74 against 3.60 in a blind human coding evaluation. Mistral's own expert comparison calls them on par for coding. Large 4 adds image input and costs about the same at list price.

Is Mistral Large 4 in Le Chat?

The launch post names only the preview API in Mistral Studio. Mistral did not announce Le Chat availability on October 6.

How do I turn reasoning on in Mistral Large 4?

Set reasoning_effort to "high" in the chat completions request, or "none" for fast answers. With "high" the response content is a list of thinking and text chunks, so update your parser.