What is Mistral Large 4?
Mistral AI launched Mistral Large 4, also called ML4 or Le Chonk, as a public preview on October 6, 2026. Its positioning covers coding, tool-using agents and image understanding; these are provider claims, not performance measurements by this site. Official announcement.
The launch gives a rounded total of 1T parameters with 52B active. Mistral trained it in its own European datacenters, where the preview is also served. Officially confirmed. The current model reference specifies 1.05T total; the memory arithmetic below uses that more precise documented count.
The preview API is available, but downloadable weights have not shipped as of this review. Mistral targets the end of October 2026 for weights and fuller technical details. Release timeline.
Mistral Large 4 VRAM Requirements
A consumer single GPU cannot hold the full Mistral Large 4 weights in conventional FP16 or INT4: the official launch count is 1T total parameters, not just the 52B active parameters. This is a capacity conclusion from parameter arithmetic, not a measured deployment result. Official parameter source. Today, use the preview API to try ML4.
What is confirmed, and what is still missing?
Officially confirmed — launch announcement:
- 1T total parameters (rounded); 52B active parameters.
- Native multimodal model with text and image input.
- Training used 3,800 NVIDIA Grace Blackwell GPUs.
- Le Chonk / ML4; weights are scheduled for the end of October 2026.
Not yet confirmed for the downloadable release:
- Full architecture details: Not publicly confirmed. The docs identify MoE, but the complete release specification is still pending.
- Weight license: Not publicly confirmed.
- Released-weight context length and runtime support: Not publicly confirmed. The preview API docs now list 1M tokens; this does not establish a local memory requirement.
- Measured VRAM after quantization: Not publicly confirmed.
Mistral says more architecture information will accompany the weight release. Official release plan.
The memory budget: three practical situations
Official Large 4 weights: arithmetic only
Estimated from parameter count; verify after weights ship. Using the official documented 1.05T total, weight-only storage is approximately 2.1 TB at FP16, 1.05 TB at INT8, and 525 GB at INT4. These are this site's arithmetic estimates: 1.05 trillion parameters × 2, 1 or 0.5 bytes, using decimal units. They are not measured VRAM requirements.
This excludes KV cache, quantization metadata, activations, runtime buffers and other overhead. It is neither a peak-VRAM measurement nor a recommended GPU configuration. Active parameters describe computation; inactive experts still need storage somewhere. CPU offloading would move part of that storage to system RAM, with unverified support and speed for this unreleased model.
Large 3 reference: keep the generations separate
These figures refer to Mistral Large 3, not Mistral Large 4. We inspected ApXML’s Large 3 memory page on October 8, 2026. Its retrieved HTML selects an FP16 panel, but the displayed memory does not reconcile with full-weight storage at the official parameter count. We could not establish reliable values across the interactive quantization states, so we do not reproduce those numbers or use them as GPU buying advice.
For a sourced deployment reference, Mistral documents Large 3 with 675B total parameters and NVFP4 deployment on a node of 8× A100 or 8× H100. This is an official configuration reference for the previous generation, not a Large 4 recommendation. Hardware capacity alone does not establish runtime compatibility, performance or sufficient cache headroom.
A smaller Mistral model for your own machine
For local experiments, start with Ministral 3 3B, which Mistral designs for edge deployment. The official Ministral 3 family also has 8B and 14B size labels. These are separate models, not smaller Large 4 checkpoints. ApXML's reviewed Large 3 page currently lists no related models, so these alternatives come from Mistral's own sources.
Choose a supported quantized checkpoint, start with short prompts and low concurrency, then measure your actual workload. Allow room for the vision encoder, cache and OS rather than treating the model's size label as a complete memory budget. For API-based development, see our developer guides; compare other uses in text models.
When will real Large 4 VRAM numbers be available?
Mistral targets the end of October 2026 for the weights. Reliable local requirements need a downloadable checkpoint and a reproducible run recording quantization, context length, batch size, hardware, runtime and peak memory. This section will be updated after weights ship and verifiable results become available. Until then, no measured Large 4 VRAM result is reported here.
Model facts, with the gaps left visible
| Detail | Reviewed information |
|---|---|
| Total parameters | Official launch: 1T (rounded); current model documentation: 1.05T. See source notes. · Source |
| Active parameters | 52B — officially confirmed · Source |
| Modalities | Text + image input; native multimodal · Source |
| Weights | Not released yet; Mistral states weights ship by end of October 2026 · Source |
| Training hardware | 3,800 NVIDIA Grace Blackwell GPUs — officially confirmed · Source |
| Codename | Le Chonk / ML4 · Source |
| Context length | Preview API: 1M tokens in current official docs. Released-weight runtime limits: Not publicly confirmed. · Source |
| License | Not publicly confirmed · Source |
| Third-party benchmarks | Not measured by this site · Source |
Launch figures and the more detailed preview documentation are attributed separately. The preview context is public; local runtime limits and the weight license remain unconfirmed.
Use cases to evaluate
- Evaluate repository understanding and coding assistance through the preview API.
- Prototype tool-using agents with application-level validation.
- Explore document, chart and image analysis on your own evaluation set.
- Assess future private deployment after weights and license terms are available.
These are evaluation ideas, not benchmark results. Validate outputs and tool actions on your own tasks.
Model Weights and License Status
Weights are not released yet as of Oct 8, 2026. A promised release, an API, a public SDK or a GitHub organization is not a downloadable model checkpoint. The license is Not publicly confirmed; do not assume the previous generation's license applies. Official weight timeline.
Official resources
Preview pricing note
Pricing is no longer wholly unpublished: the launch page lists USD $1.36 input / $4.18 output per million tokens, while the current official docs display preview sale rates of $0.68 input / $2.09 output per million tokens. These are provider-listed rates, not independently verified bills. Post-preview pricing: Not publicly confirmed. Check current account terms before use.
Frequently asked questions
How much VRAM does Mistral Large 4 need?
No measured requirement for released official weights is available. The sourced, estimated weight-only storage calculation above excludes cache and runtime overhead; it is not a deployment measurement.
Can I run Mistral Large 4 on an RTX 4090?
A single RTX 4090 cannot hold the full model at the conventional precisions estimated above. We have not tested offloading, and official downloadable weights are still pending. Use the preview API or evaluate a smaller Ministral checkpoint.
Can Mistral Large 4 run on 24GB VRAM?
Not with the complete conventional FP16, INT8 or INT4 weights resident in that memory. Offloading is a different configuration and has no verified Large 4 result here.
Are Mistral Large 4 weights available?
Not as of October 8, 2026. Mistral targets the end of October for their release; the linked announcement is a timeline, not a released checkpoint.
Can I run Mistral Large 4 on a consumer GPU?
Not with all conventional FP16 or INT4 weights resident on a consumer single GPU. The weight-only estimates above are far beyond that capacity, and downloadable weights are not available yet. Try the preview API or a smaller Ministral model.
Is 16GB of VRAM enough for AI development?
It can be enough for selected smaller quantized models and some experiments, depending on the checkpoint, context and workload. API-based development does not require loading model weights onto your GPU. It is not enough to hold Large 4 weights; no Large 4 inference was tested here.
How much VRAM do you need in 2026?
There is no universal requirement. API use, local inference and training have different budgets. Start from the complete checkpoint size and precision, then add cache and runtime headroom for your context and concurrency. Use model-specific, reproducible measurements before buying hardware.
Are the Large 3 figures also Large 4 VRAM requirements?
No. The table above is explicitly a third-party Large 3 reference with unresolved storage assumptions. Large 4 figures on this page are arithmetic estimates only, pending the weight release.
Go straight to the source
Reviewed Oct 8, 2026. Official statements, third-party reference figures and this site's arithmetic estimates are distinguished above. No Large 4 local inference or third-party benchmark was measured by this site. Unknown values reflect the sources reviewed on this date.