Open weights are a deployment decision, not a magic label
Compare the releases that matter for local inference, private deployment, and provider independence. The licence, memory footprint, serving stack, and workload all count.
A compact on-device agent and tool-calling model for latency-sensitive local workflows.
Licence
LFM Open License v1.0
Scale
2.6B
Practical fit: A better candidate for edge and private inference than the frontier-sized releases in this list.
Deployment: Commercial use is free below $10M annual organisational revenue under the LFM Open License v1.0; measure tool-call reliability and instruction following at your target memory and quantisation level.
Meta’s open-weight multimodal agentic model distilled from Muse Spark for local consumer hardware.
Licence
Apache 2.0
Scale
~29.6B dense (incl. ~1.8B perception encoder)
Practical fit: The practical local-agent path when Muse Spark API is too heavy or closed for your deployment.
Deployment: Full BF16 needs far more than 24GB — plan on official GGUF / 4-bit packs (and mmproj for vision). Validate agentic quality on your workflows.
Official Hugging Face checkpoint for the Qwen3.8 Max-class open weights — text-only, thinking required, native 262K context.
Licence
Qwen3.8-Max License
Scale
2.4T total / 95B active MoE
Practical fit: The self-host path for Qwen3.8 flagship weights. It is not a drop-in for the hosted Qwen3.8-Max API's vision, tools, or 1M default context.
Deployment: Datacentre-class memory. YaRN can extend toward ~1,010,000 tokens. MaaS / AI Work Assistant businesses over $50M trailing 12-month revenue need a separate commercial licence.
Apache 2.0 Qwen3.8 dense VLM with native 262K context and image/video input — the practical local sibling of hosted Qwen3.8-Max.
Licence
Apache 2.0
Scale
27B dense
Practical fit: The first Qwen3.8 local target when 2.4T-A95B is too large and you can accept a dense 27B serving envelope.
Deployment: Hosted 1M-context API is coming soon. Vendor DeepSWE/SWE-bench Pro numbers used Claude Code, not the public Datacurve or swebench.com boards.
These release signals stay visible without being presented as deployable open-weight comparisons. We will promote them when public checkpoints, licences, and practical access details are clear.
Qwen
Qwen3.8-Max / Qwen3.8-2.4T-A95B / Qwen3.8-27B
Hosted API plus two Hugging Face checkpoints are published — keep them as separate identities
Qwen3.8-Max (hosted API), Qwen3.8-2.4T-A95B (Qwen3.8-Max License), and Qwen3.8-27B (Apache 2.0 dense VLM) are all catalogued. Do not merge them: the API adds vision, non-thinking, 1M context, and tools; 2.4T is text-only with thinking required; 27B is the dense local VLM.
Live since 2026-08-02; watermark grace to 2026-12-02; high-risk deferred under Omnibus
Regulatory context for UK SMEs shipping chatbots / synthetic content — not a model SKU. Retire any copy that treats 2 Aug 2026 as a universal high-risk deadline.
Replacement path to Gemini 3.x; EOL mid-Oct 2026 (Google docs conflict Oct 16 vs Oct 20)
Do not recommend 2.5 as a long-term default. Confirm the live Google Cloud lifecycle date before migration day; prefer Gemini 3.x Flash/Pro paths already in the catalog.
Quality Mode GA; grok-imagine-image-2.0 is $0.04/image on the xAI API
Creative image product — not a text/coding frontier row. Do not conflate with Grok 4.6 text. Public API pricing is live; a dedicated image-model profile is still deferred until the comparison table has an image-model axis.
Free / Premium $20 sunsetting / Plus ~$30 / Business ~$100
Browser UI builder sits beside Lovable/Bolt rather than IDE harnesses. Prefer Plus/Business framing — Premium $20 is sunsetting. Full ACE harness row deferred until destination + setup depth match catalog standards; see Lovable vs Bolt guide shortlist.
Gated Daybreak / cyber-defense access, not a general ChatGPT or API shortlist substitute. OpenAI lists GPT-5.6 Cyber at $12.50 input / $75 output per million tokens; enrollment is still required.