Skip to main content

Open weights

Open weights are a deployment decision, not a magic label

Compare the releases that matter for local inference, private deployment, and provider independence. The licence, memory footprint, serving stack, and workload all count.

Moonshot AI

Kimi K3

A frontier-scale open-weight model with a million-token context and long-horizon agent ambitions.

Licence
Kimi K3 License
Scale
2.8T total / 104B active MoE

Practical fit: Best for teams evaluating provider independence and long-context work, not casual local inference.

Deployment: The checkpoint is roughly a datacentre deployment; review the custom licence before redistribution.

An official post-trained V4-Flash checkpoint focused on coding, agents, reasoning, and tool use.

Licence
MIT
Scale
284B total / 13B active MoE

Practical fit: A compelling hosted API price-performance option with open deployment for teams that can carry the hardware.

Deployment: Only a small fraction activates per token; the full expert set still has to be served.

Thinking Machines

Inkling

An open multimodal mixture-of-experts release designed for broad agent and reasoning experiments.

Licence
Apache 2.0
Scale
975B total / 41B active MoE

Practical fit: Useful for teams building an open-model evaluation lane, especially where modality breadth matters.

Deployment: Total parameter count makes quantisation, serving, and memory planning central to the decision.

Thinking Machines

Inkling Small

A smaller Inkling checkpoint aimed at bringing much of the family’s capability into a more practical serving envelope.

Licence
Apache 2.0
Scale
276B total / 12B active MoE

Practical fit: The better first experiment for teams that want to test the family without starting with the full checkpoint.

Deployment: Smaller does not mean laptop-sized; estimate memory and throughput from the actual release artefacts.

LG AI Research

K-EXAONE 2.0

A large sovereign-model release from LG AI Research for teams evaluating Korean and enterprise-focused open infrastructure.

Licence
Apache 2.0
Scale
750B foundation model

Practical fit: Interesting for regional language, sovereignty, and research use cases where a broad open ecosystem matters.

Deployment: Confirm the actual checkpoint, serving recipes, and hardware requirements before treating it as a production option.

A smaller policy-adaptive safety model designed to sit around a generation system rather than replace it.

Licence
Apache 2.0
Scale
3B classifier

Practical fit: Useful for moderation, routing, and safety layers where a compact model can reduce latency and cost.

Deployment: Evaluate false positives and policy fit on your own content; a classifier is not a complete safety programme.

Liquid AI

LFM2.5-2.6B

A compact on-device agent and tool-calling model for latency-sensitive local workflows.

Licence
LFM Open License v1.0
Scale
2.6B

Practical fit: A better candidate for edge and private inference than the frontier-sized releases in this list.

Deployment: Commercial use is free below $10M annual organisational revenue under the LFM Open License v1.0; measure tool-call reliability and instruction following at your target memory and quantisation level.

An open vision-language-action model aimed at embodied and autonomous systems rather than general chat.

Licence
OpenMDW-1.1
Scale
34B total / 32B VLM backbone + action expert

Practical fit: Relevant to robotics and physical-world evaluation, not a drop-in chatbot replacement.

Deployment: Hardware, simulator, sensor, and safety validation are part of the product decision.

Meta’s open-weight multimodal agentic model distilled from Muse Spark for local consumer hardware.

Licence
Apache 2.0
Scale
~29.6B dense (incl. ~1.8B perception encoder)

Practical fit: The practical local-agent path when Muse Spark API is too heavy or closed for your deployment.

Deployment: Full BF16 needs far more than 24GB — plan on official GGUF / 4-bit packs (and mmproj for vision). Validate agentic quality on your workflows.

InclusionAI

Ling 3.0 Flash

InclusionAI’s cost-focused Flash MoE tier for high-frequency hybrid reasoning and agent loops.

Licence
MIT
Scale
124B total / 5.1B active MoE

Practical fit: A budget open-weight option when you need MoE efficiency and can host the expert set.

Deployment: ~262K context is smaller than many frontier peers — confirm serving provider and licence before volume commit.

Official Hugging Face checkpoint for the Qwen3.8 Max-class open weights — text-only, thinking required, native 262K context.

Licence
Qwen3.8-Max License
Scale
2.4T total / 95B active MoE

Practical fit: The self-host path for Qwen3.8 flagship weights. It is not a drop-in for the hosted Qwen3.8-Max API's vision, tools, or 1M default context.

Deployment: Datacentre-class memory. YaRN can extend toward ~1,010,000 tokens. MaaS / AI Work Assistant businesses over $50M trailing 12-month revenue need a separate commercial licence.

Apache 2.0 Qwen3.8 dense VLM with native 262K context and image/video input — the practical local sibling of hosted Qwen3.8-Max.

Licence
Apache 2.0
Scale
27B dense

Practical fit: The first Qwen3.8 local target when 2.4T-A95B is too large and you can accept a dense 27B serving envelope.

Deployment: Hosted 1M-context API is coming soon. Vendor DeepSWE/SWE-bench Pro numbers used Claude Code, not the public Datacurve or swebench.com boards.

Watchlist

Interesting, but not verified for the catalog yet

These release signals stay visible without being presented as deployable open-weight comparisons. We will promote them when public checkpoints, licences, and practical access details are clear.

Qwen

Qwen3.8-Max / Qwen3.8-2.4T-A95B / Qwen3.8-27B

Hosted API plus two Hugging Face checkpoints are published — keep them as separate identities

Qwen3.8-Max (hosted API), Qwen3.8-2.4T-A95B (Qwen3.8-Max License), and Qwen3.8-27B (Apache 2.0 dense VLM) are all catalogued. Do not merge them: the API adds vision, non-thinking, 1M context, and tools; 2.4T is text-only with thinking required; 27B is the dense local VLM.

Qwen3.8-2.4T-A95B on Hugging Face

Amazon

Amazon Nova Premier / Canvas / Reel (Bedrock)

Legacy — EOL 2026-09-14 (Premier) / 2026-09-30 (Canvas/Reel)

Not a buyable long-term shortlist: Bedrock lifecycle lists hard EOL; migrate before those dates rather than recommending as a stable AWS buy.

AWS Bedrock model lifecycle

Cohere

Cohere Command R / Command R+ (Bedrock)

EOL on Bedrock 2026-08-19

Enterprise AWS paths should migrate this week; Command A+ is the newer Cohere story — not yet a verified ACE comparison row.

AWS Bedrock model lifecycle

EU / UK buyers

EU AI Act Article 50 transparency

Live since 2026-08-02; watermark grace to 2026-12-02; high-risk deferred under Omnibus

Regulatory context for UK SMEs shipping chatbots / synthetic content — not a model SKU. Retire any copy that treats 2 Aug 2026 as a universal high-risk deadline.

EU AI transparency guidelines

Relay

Relay.app shutdown

EOL — free wipe 2026-08-15 23:59 PT; paid 2026-09-14

Do not recommend. Export and migrate to Zapier / Make / n8n. See /ai-guides/relay-app-exit.

Relay.app shutdown notice

Google

Gemini 2.5 Pro / Flash retirement

Replacement path to Gemini 3.x; EOL mid-Oct 2026 (Google docs conflict Oct 16 vs Oct 20)

Do not recommend 2.5 as a long-term default. Confirm the live Google Cloud lifecycle date before migration day; prefer Gemini 3.x Flash/Pro paths already in the catalog.

Google Cloud model versions & lifecycle

xAI

xAI Imagine Image 2.0

Quality Mode GA; grok-imagine-image-2.0 is $0.04/image on the xAI API

Creative image product — not a text/coding frontier row. Do not conflate with Grok 4.6 text. Public API pricing is live; a dedicated image-model profile is still deferred until the comparison table has an image-model axis.

xAI models docs

Vercel

v0 (Vercel) plan ladder

Free / Premium $20 sunsetting / Plus ~$30 / Business ~$100

Browser UI builder sits beside Lovable/Bolt rather than IDE harnesses. Prefer Plus/Business framing — Premium $20 is sunsetting. Full ACE harness row deferred until destination + setup depth match catalog standards; see Lovable vs Bolt guide shortlist.

v0 pricing docs

Google

Gemini 3.5 Flash Cyber

Limited or preview access

The release signal is useful, but public pricing and stable API details were not sufficient for the verified comparison table.

Google model announcement

OpenAI

OpenAI Daybreak / GPT-5.6-Cyber

Defender-focused; HSK path from Sep 2026

Gated Daybreak / cyber-defense access, not a general ChatGPT or API shortlist substitute. OpenAI lists GPT-5.6 Cyber at $12.50 input / $75 output per million tokens; enrollment is still required.

OpenAI Daybreak

The three questions

Can I use it?

Read the licence, model card, and commercial conditions. Open weights do not make every downstream use automatically permissible.

Can I run it?

Total parameters, active parameters, quantisation, memory, throughput, and serving software determine feasibility.

Should I run it?

Compare privacy, cost, latency, maintenance, and workload quality against a hosted API before buying hardware.