Skip to main content
A 1T-parameter MoE model with open weights coming doesn't make self-hosting the right call for most teams.

Mistral Large 4 'Le Chonk': Self-Host or Use the API?

KU
Kiril Urbonas
2 days ago • 6 min read•1 view

A 1T-parameter MoE model with open weights coming doesn't make self-hosting the right call for most teams.

Key takeaways

A 1T-parameter MoE model with open weights coming doesn't make self-hosting the right call for most teams.

Mistral shipped Large 4 on October 6, 2026, nicknamed "Le Chonk" after the internet's own joke about the company, and the headline spec is a mixture-of-experts model with roughly 1 trillion total parameters and about 49 billion active per request; the preview is live on Mistral's API now, the weights are promised later in October, and the honest read is that this changes almost nothing for the vast majority of teams who should keep calling an API, open or closed.

What actually shipped#

Mistral's own announcement, corroborated by The Register and several other outlets, puts Large 4 at around 1 trillion total parameters with about 49 billion active at inference, a mixture-of-experts design that keeps per-token compute far below what a dense trillion-parameter model would cost. Some coverage cites slightly different figures (1.05T total, 52B active); treat the numbers as "roughly 1T, roughly 50B active" rather than exact. The model is natively multimodal, reportedly trained on several thousand Nvidia Grace Blackwell GPUs (reports range from about 3,800 to 4,000 chips) over roughly two months in Mistral's own European data centers. Mistral says it beats open and closed competitors on enterprise workloads like cybersecurity, finance, and legal tasks. Independent testing from Artificial Analysis tells a less flattering story, placing the preview behind other current open models on general intelligence benchmarks. Mistral's own benchmarks and the third-party ones disagree, which is normal for launch week and a good reason to wait for your own eval before switching anything.

Right now you can only reach Large 4 through Mistral's hosted API preview. The open weights, the part that matters for self-hosting, aren't out yet. Mistral says they're coming by the end of October; one outlet names October 27, but that date isn't corroborated widely enough to treat as fixed, so plan around "later in October" and confirm the day closer to release.

The self-hosting reality check#

A ~1 trillion parameter MoE model, even with ~49B active parameters, is not a weekend self-host. At FP8 you're looking at roughly 1TB+ of weights to hold in GPU memory before you even add KV cache headroom, which means multi-node clusters of H100/H200/Blackwell-class GPUs, NVLink or fast interconnect between them, and someone on staff who has run tensor-parallel and expert-parallel inference before. This is not "docker run" territory. The active-parameter count helps with compute per token, not with the memory footprint of holding the whole model in VRAM, and MoE routing adds its own operational quirks around load balancing across experts that dense-model operators haven't had to think about.

yaml.yaml
# vllm-large4-moe.yaml: illustrative multi-node serving config, not a copy-paste guarantee
model: mistralai/Large-4-Le-Chonk
tensor_parallel_size: 8
pipeline_parallel_size: 2       # spans 2 nodes of 8 GPUs each
enable_expert_parallel: true
max_model_len: 131072
gpu_memory_utilization: 0.92
quantization: fp8               # halves memory vs bf16, still needs ~500GB+ across the cluster
served_model_name: le-chonk
bash.bash
# rough monthly cost sanity check: self-hosted GPU-hours vs API tokens
$ echo "16x H100 (2 nodes) on-demand @ ~$2.50/GPU-hr = ~$28,800/mo running 24/7"
$ echo "vs. API: at $1.36/$4.18 per 1M in/out tokens, that buys ~5.7B input"
$ echo "tokens or ~4.1M requests of ~1.4K tokens each, per month, with zero ops burden"

That second block is rough, token pricing and GPU spot pricing both move weekly, but the order of magnitude holds: unless you're running the cluster near saturation with a team to babysit it, the API is cheaper and faster to ship.

Who open weights actually help#

Open weights on a model this size aren't a self-hosting invitation for most companies; they're an option for three groups. Teams with data-residency or sovereignty constraints who legally cannot send prompts to a US or French-hosted API get a real alternative: run it inside your own EU data center and keep data in-region, which is exactly the pitch Mistral is making to European enterprises. Teams that already operate serious GPU infrastructure, the ones running their own training or fine-tuning clusters for other reasons, can slot Large 4 in without standing up new capacity from scratch. And researchers who need to inspect weights, run interpretability work, or fine-tune on proprietary data without sending it anywhere benefit regardless of inference cost, because for them the alternative isn't "cheaper," it's "impossible."

Everyone else, including most startups and mid-size engineering teams reading this, should treat the open weights as a future inference-provider SKU, not a self-hosting project. Expect Fireworks, Together, Groq, or similar providers to host Large 4 within days of weights landing, at which point you get the open-weights pricing and licensing benefits without the multi-node cluster.

The closed-vs-open question this actually answers#

The interesting decision isn't "Mistral vs OpenAI vs Anthropic," which we've covered in detail in our closed-API comparison. It's whether "frontier-capable" now has a credible open-weights path at all, separate from which closed vendor you'd pick. Large 4 is Mistral's answer to that question, and the answer is "yes, if you can afford the infrastructure or wait for a host to do it for you." That's a narrower claim than "open models have caught up," and the Artificial Analysis numbers are a useful reminder that Mistral's own benchmarks and independent ones don't yet agree.

The decision, concretely#

  • Need EU data residency and can't send prompts to a US API? Wait for weights, budget for a multi-node GPU cluster or a sovereign-cloud inference provider, and treat this as infrastructure, not a quick swap.
  • Already running a GPU cluster for training or other inference? Worth evaluating once weights land and your own benchmarks confirm Mistral's claims, since the marginal cost of adding one more model is low.
  • Just want a good model at a fair price, no infra team? Skip self-hosting entirely and pick from the best LLM APIs once a provider hosts Large 4, or keep using what you already have.
  • Evaluating it today, before weights ship? Use the API preview for a quick eval, but don't make procurement decisions off Mistral's own benchmarks alone, run your own tasks.

The call we'd make#

We'd use the API preview to see if Large 4 actually beats what we're running today on our own evals, and we'd wait for an inference provider to host the open weights rather than standing up a multi-node cluster ourselves, unless data residency or an existing GPU fleet makes the math different. A trillion-parameter MoE model being open is genuinely useful for the handful of teams who need it badly enough to run it themselves; for everyone else it's one more good option on someone else's API, not a self-hosting project.

Explore topics:AI
React

Get the DevOps Troubleshooting Cheat Sheet

Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.

Share this post
KU

Kiril Urbonas

AI Engineer

567 articles
View all articles by Kiril Urbonas

You might have missed

Evergreen posts worth revisiting.