Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

About

This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd context including the previewed Behemoth teacher model, (ii) architectural characteristics beyond a high-level MoE description covering routed/shared-expert structure, early-fusion multimodality, and long-context design elements reported for Scout (iRoPE and length generalization strategies), (iii) training disclosures spanning pre-training, mid-training for long-context extension, and post-training methodology (lightweight SFT, online RL, and lightweight DPO) as described in release materials, (iv) developer-reported benchmark results for both base and instruction-tuned checkpoints, and (v) practical deployment constraints observed across major serving environments, including provider-specific context limits and quantization packaging. The manuscript also summarizes licensing obligations relevant to redistribution and derivative naming, and reviews publicly described safeguards and evaluation practices. The goal is to provide a compact technical reference for researchers and practitioners who need precise, source-backed facts about Llama 4.

Redacted by arXiv• 2026

Related benchmarks

TaskDatasetResultRank
GeolocationWanderBench Street level
Accuracy19.8
21
GeolocationWanderBench City level
Accuracy37.2
21
GeolocationWanderBench Country level
Accuracy69.6
21
GeolocationWanderBench Overall
Distance Error (km)1.73e+3
21
Dynamic Cross-View Spatial IntelligenceLinkS2Bench (tiny)
Target Verification (TV)79.6
10
Showing 5 of 5 rows

Other info

Follow for update