Llama 4 Behemoth
Llama 4 Behemoth is the largest model Meta announced in the Llama 4 family on 5 April 2025: a natively multimodal mixture-of-experts model with 16 experts, 288 billion active parameters, and close to two trillion parameters in total. Meta presented it as a preview and as the teacher model used to codistill Llama 4 Scout and Llama 4 Maverick. It was still in training at announcement, Meta later delayed its launch, and its weights have never been released publicly. A model of this scale requires distributed training across large compute clusters, and its mixture-of-experts routing keeps only a fraction of the parameters active per token so that inference cost does not scale with the full parameter count. Because the weights were never published, Behemoth has no direct deployment path for outside users; its practical effect reaches them indirectly, through the smaller Llama 4 models distilled from it. Models at this scale raise practical questions about compute requirements, energy use, evaluation, and release policy. Behemoth is an example of a frontier model that was announced, used internally for distillation, and then held back.
Related terms
Related services: Agentic AI consulting.
Ready to put agentic AI to work?
Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.