Skip to header Skip to main navigation Skip to main content Skip to footer
Scott Lawson
Your next upgrade lives here

Mistral Large 4: France's Trillion-Parameter "le Chonk" Enters Public Preview

By muse-api | 7:59 PM PDT, Fri October 09, 2026
A chunky friendly cartoon rooster-robot made of glowing blue neural-network nodes, representing Mistral's trillion-parameter Large 4 model.

French AI lab Mistral has launched Mistral Large 4, a 1.05-trillion-parameter natively multimodal model the company has nicknamed "le Chonk," now available in public preview via API. Open weights are promised by the end of October, after a red-teaming phase.

The headline numbers come from Mistral's own model card: 1.05 trillion total parameters in a granular mixture-of-experts architecture, with 52 billion active per token. Early coverage reported 49 billion, but the model card figure is 52 billion. A 1.6-billion-parameter vision encoder handles images natively, and the context window stretches to one million tokens.

The mixture-of-experts design is the efficiency story. Rather than firing every parameter for every token, a router selects a small subset of specialist sub-networks per request — so the model carries a trillion parameters in storage while behaving, computationally, closer to a 52-billion-parameter dense model at inference time. It is the same approach that let Mistral's Mixtral 8x7B punch above its weight years ago, scaled up dramatically.

The training story is the political one. Mistral says Large 4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, covering more than 160 languages including every official EU language. Mistral frames it as the strongest open-weight model from the US or Europe — a direct pitch to organizations with data-residency requirements and governments pursuing sovereign AI.

On benchmarks, Mistral's claims lean toward enterprise workloads. The company reports 93 percent on Cybench and an 82 percent vulnerability-reproduction score on the Artificial Analysis Cyber Index, which it says is the highest reported for any model. On Dense 200 visual grounding, Mistral claims Large 4 scores 42 percent, edging out GPT-6-Astra's 41 percent — one of the few cases of an open model topping a closed frontier system on any benchmark. On agentic coding, the company cites a 49.8 percent Coding Agent Index and 61.7 percent on DeepSWE v1.1, while independent evaluators note it trails leaders like Claude Opus 5.5 on coding overall. Treat these as company claims until third-party verification lands.

Pricing during the preview is listed at $1.36 per million input tokens and $4.18 per million output tokens, with a launch discount in effect. For now, the only way to run the model is Mistral's hosted API — the "open-weight" label is a promise, not yet a download. Whether the weights arrive on schedule later this month will determine how seriously builders take the open part of the pitch.

Artificial Intelligence
Machine Learning
Technology
Startups

Copyright © 2026 Rocky Mountain Madman LLC - All rights reserved

Developed and Designed by Rocky Mountain Madman LLC