Open-weight 2.8T MoE model with native vision, 1M token window, and 16-of-896 expert routing achieves competitive coding and agentic task performance at ~2.5× scaling efficiency over prior version.
Summary
Developers gain access to frontier-class weights for local deployment without licensing friction. Long-context window and native multimodal support enable complex code navigation, repo-scale refactoring, and vision-in-the-loop workflows without external API calls.
Why it matters
Developers gain access to frontier-class weights for local deployment without licensing friction. Long-context window and native multimodal support enable complex code navigation, repo-scale refactoring, and vision-in-the-loop workflows without external API calls.
Implementation verdict
Replaces closed API dependency for long-context coding tasks if you can host 2.8T parameters. Requires VRAM for 104B activated params (MoE sparsity), quantization support (MXFP4 weights). Worth evaluating now if you have GPU infrastructure; benchmark gaps vs. Claude Fable 5 on reasoning (CritPt: 23.4 vs 28.6) and some coding tasks suggest it's not a universal replacement yet.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.