Timeline

Meta releases an open-weight model small enough to run on one consumer GPU

Muse Glimmer is a 30-billion-parameter Apache-licensed model, under 20GB at 4-bit quantisation, distilled from Meta's larger Muse Spark and released alongside Zuckerberg's essay on AI strategy.

  • Open weights & ecosystem
  • Models & capabilities
  • Notable

Meta released Muse Glimmer, a 30-billion-parameter dense multimodal model under the permissive Apache 2.0 licence, compact enough to run at under 20GB using 4-bit quantisation — small enough for a single consumer graphics card rather than a datacentre cluster. Meta positioned it for “always-on local agent workflows” such as scheduling, drafting messages and organising files, and said it supports multimodal input, tool use and multi-step reasoning while running fast enough for real-time conversational use. The company listed day-one support in llama.cpp, MLX and ExecuTorch, with integrations for Ollama, LM Studio, vLLM and SGLang following shortly after.

Meta said Glimmer was trained through logit distillation from the outputs of Muse Spark, its larger sibling model, making it a compressed, locally deployable version of Meta’s flagship line rather than an independently trained architecture.

The release landed the same day Mark Zuckerberg published a lengthy essay arguing for individually controlled “personal superintelligence” over centrally hosted systems, and announcing that Meta would resume open-weight releases after a period of tighter control over its most capable models. Glimmer’s small footprint, permissive licence and day-one support across local-inference tooling gave that pledge an immediate, concrete example — a model genuinely runnable by an individual rather than only by a well-resourced lab.