Mistral previews trillion-parameter model, promises October weights
On 6 October, Mistral AI launched a public preview of Mistral Large 4, a natively multimodal mixture-of-experts model it says has one trillion total parameters and 49 billion active parameters. The company says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its European data centers, and plans to release the weights by the end of October. Those are company-reported specifications; the weights and further architecture details were still forthcoming in the announcement.
The record ends here. Everything below this line is invented.
What if only 49 billion parameters worked per token, but a trillion still had to be on hand?
Part one
The Idle Rack
At 18:40, the model refused to finish a word. Eira's console showed 49 billion active parameters—the slice working on each token—and a red bar for one trillion total; the cooling row behind her hummed as if those were the same number.
Her brother Olan waited beside a scanner with a page from their grandmother's paper ledger. The archive stayed on the municipal network by agreement; no image left the room. One word in the margin could settle which of two orchard plots belonged on the family map.
"It's only using forty-nine billion at a time," Eira said. "I'll bring the hot rack back. The rest can sleep."
The dispatcher lit a yellow request for an expert stored in the dark row. Olan tilted the paper toward the lamp; its ink faded into the grain.
Eira had thought active meant needed. She opened the activity trace. The counter stayed near 49 billion as different experts lit for different parts of the same sentence; the full model's list still stretched across the dark racks. Every idle rack remained part of the model, waiting for its turn.
The word came through in pieces: a place name, then a dash, then nothing. Active parameters counted the token's work; total parameters counted the model's full inventory. The missing syllable blinked on the local screen.
What’s real
- Mistral reports one trillion total and 49 billion active parameters.
- Large 4 is a mixture-of-experts model.
- Mistral announced training on 3,800 Grace Blackwell GPUs.
What’s invented
- Eira, Olan, and the municipal compute cooperative are fictional.
- Their local archive deployment and cooling failure are invented.
- The model's use for private family handwriting is speculative.
Part two: The Local Share
Part two isn’t on sale yet. Check back soon.
Filed under technology, mixture of experts, open weights, data centers.