Meta
Muse Glimmer 30B
Meta's open-weight multimodal agentic model distilled from Muse Spark for local consumer hardware (~24–32GB class with 4-bit).
Catalog record checked August 10, 2026; individual provider fields may change.
| Specification | Muse Glimmer 30B |
|---|---|
| Provider | Meta |
| Tier | Balanced |
| Context window | 131K |
| Max output | Not verified |
| Input / 1M tokens | $0 |
| Output / 1M tokens | $0 |
| Weights | Open |
| Parameters | ~29.6B dense (incl. ~1.8B perception encoder) |
| Modalities | text, image |
| Released | August 10, 2026 |
Pricing tiers: Apache 2.0 open weights (BF16 + official GGUF); hosted inference billed by provider.
Best for
- Local agents
- On-device coding + tool use
- Privacy-sensitive multimodal work
Watch out
Common questions
Muse Glimmer 30B
Answered from the verified figures on this page rather than general guidance.
How much does Muse Glimmer 30B cost per million tokens?
Muse Glimmer 30B is listed at $0 per million input tokens and $0 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Apache 2.0 open weights (BF16 + official GGUF); hosted inference billed by provider.
What is Muse Glimmer 30B's context window?
Muse Glimmer 30B accepts about 131K tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is Muse Glimmer 30B best for?
Muse Glimmer 30B is a balanced tier from Meta. It suits local agents, on-device coding + tool use, privacy-sensitive multimodal work. Full BF16 needs far more than 24GB — plan on official GGUF / 4-bit packs (and mmproj for vision). Agentic quality ≠ frontier Muse Spark API; validate on your workflows.
Can I self-host Muse Glimmer 30B?
Muse Glimmer 30B publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.
Compare it
Head-to-head model comparisons
These are the published pairings that put this model against a plausible alternative.