Meta Launches Muse Glimmer, a 30B Open Source Multimodal Agent
Meta’s Apache 2.0 Muse Glimmer packs multimodal agents into 30B parameters, with local inference, broad day-one support and optional DFlash acceleration.
Summary
Meta released Muse Glimmer on August 10, 2026, a 30B parameter multimodal model distilled from Muse for private, lower-cost local agents spanning coding, documents, assistants, tool use, object detection, images and silent video. Apache 2.0 licensed and hosted on Hugging Face, it combines a 2B ViT-style Perception Encoder with a 28B text decoder. The decoder repeats three 2,048-token sliding-window attention layers and one full NoPE layer across 52 layers; grouped-query attention shares each key-value head across 16 query heads, cutting KV-cache memory 16 times. Its 50-layer vision tower uses 14 by 14 patches, while 2 by 2 pixel shuffle reduces image tokens fourfold; video processing targets 2 frames per second and caps clips at 96 frames. Optional DFlash speculative decoding trades memory for faster generation, using blocks of 16 with up to 15 proposed tokens.
Published results place Muse Glimmer first among Gemma4-31B Thinking Mode and Qwen3.6-27B Thinking Mode on 12 of 22 capability benchmarks, including MCP Atlas at 75.5, DeepSearch QA at 74.6, SWE-Bench Pro at 51.2, SciCode at 43.6, Charxiv Reasoning at 78.8, IFBench at 77.0 and AIME 2026 at 94.7. It trailed leaders on GDPval-AA, SkillsBench, OSWorld-Verified, SWE-Bench Verified, TerminalBench 2.1, ScreenSpot Pro, OmniDocBench v1.5, MMMU Pro, GPQA Diamond and Humanity’s Last Exam. Safety remained uneven: CI Memories recorded 26.4 violation and 64.8 coverage, while Siren AgentDojo showed 28.4 attack success and 94.2 utility. Day-zero support covers Transformers, llama.cpp, vLLM and Inference Endpoints across NVIDIA CUDA, AMD ROCm and Intel XPU. BF16 evaluation needs one 80GB H100, full SFT needs eight, and eight GPUs are usually insufficient for full-finetune GRPO; TRL supports SFT through Async GRPO.
Positives
- Apache 2.0 licensing permits local deployment, modification and commercial use of Meta’s 30B parameter Muse Glimmer.
- Published scores place Muse Glimmer first among the three compared models on 12 of 22 capability benchmarks, including MCP Atlas at 75.5.
- Grouped-query attention reduces KV-cache memory 16 times by sharing each key-value head across 16 query heads.
- Day-zero support spans Transformers, llama.cpp, vLLM and Hugging Face Inference Endpoints on NVIDIA, AMD and Intel accelerators.
- DFlash speculative decoding accelerates generation with an optional block-diffusion drafter, particularly for structured output such as code.
Risks & concerns
- Muse Glimmer’s CI Memories violation rate was 26.4, above Gemma4-31B’s 12.1, although below Qwen3.6-27B’s 53.4.
- Siren AgentDojo recorded a 28.4 attack success rate, worse than Gemma4-31B’s 25.6 despite Muse Glimmer’s leading 94.2 utility.
- Qwen3.6-27B led Muse Glimmer on OSWorld-Verified, 75.6 to 65.9, and TerminalBench 2.1, 60.7 to 51.7.
- Full BF16 supervised fine-tuning requires eight 80GB H100 GPUs, while eight GPUs are usually insufficient for full-finetune GRPO.
- Video inference excludes audio, targets 2 frames per second and limits clips to 96 evenly sampled frames.