Meta Superintelligence Lab has launched Muse Glimmer, a 30-billion parameter agentic AI model under the Apache 2.0 license. Distilled from the larger Muse Spark model using logit distillation, Glimmer is optimized for local execution on consumer hardware. It features a dedicated perception encoder for multimodal understanding and utilizes 4-bit quantization to reduce its memory footprint to under 20 GB, fitting within 24 GB VRAM GPUs. The model employs speculative decoding via a "DFlash" drafter model, achieving up to a 3x speedup on high-end consumer cards like the RTX 5090. Benchmarks suggest Glimmer outperforms Google's Gemma 4 and rivals Qwen 3.6 in agentic coding tasks. Mark Zuckerberg’s accompanying manifesto advocates for open-source AI to prevent power concentration among a few labs. Despite recent shifts toward closed models, Meta plans to release open weights for Muse Spark 1.2. The model's design emphasizes autonomous task completion, long-context memory, and instruction following without requiring cloud connectivity.
Muse Glimmer is a 30B parameter model released under the Apache 2.0 open-source license. The model was distilled from Meta's larger, closed Muse Spark model using logit distillation.
Quantization techniques compress the model to under 20 GB, allowing it to run on consumer GPUs. Speculative decoding with the DFlash drafter model provides up to a 3x performance increase.
Benchmarks indicate strong performance in agentic coding compared to Gemma 4 and Qwen 3.6. Meta intends to release open weights for the Muse Spark 1.2 foundation model soon.
Chapter guide
Worth noting
- The video contains a paid sponsorship for OpenRouter.
- Performance benchmarks are based on specific community tests that may vary in real-world use.
- The release of Muse Spark 1.2 open weights is a stated intention and not yet fulfilled.
- The video uses a futuristic setting (2026) and references unreleased versions of competitor models.