# AI Choice Engine model catalog

> Token prices, context windows, and attributed benchmarks for models we maintain. Unverified fields say so. HTML: https://aichoiceengine.com/ai-models

Do not treat vendor-only scores as public-board scores. DeepSWE figures are from https://deepswe.datacurve.ai/ unless a page says otherwise.

## GPT-5.6 Sol

- Profile: https://aichoiceengine.com/ai-models/gpt-5-6-sol
- Provider: OpenAI
- API id: not published
- Open weights: no
- Context / max output: 1,050,000 / 128,000
- Price per million tokens: $5 input, $30 output
- Reasoning levels: none, low, medium, high, xhigh, max
- Modalities: text, image
- Best for: Complex reasoning; Agentic workflows; Hard coding tasks
- Watch out: 25x Luna's input price — overspecified for routine generation. Sol Fast mode (where offered) is roughly ~2× price for ~2.5× speed — only enable when latency is the constraint. `gpt-5.2-chat-latest` / `gpt-5.3-chat-latest` shut down 2026-08-10 — migrate API chat traffic to Sol. ChatGPT Plus/Pro Chat got an Aug 6 Sol refresh + effort slider; Work/Codex/API July builds are unchanged.
- Verified: 2026-08-10
- Benchmarks: Artificial Analysis Intelligence Index 61 (max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 73% at $8.39/task

## GPT-5.6 Terra

- Profile: https://aichoiceengine.com/ai-models/gpt-5-6-terra
- Provider: OpenAI
- API id: not published
- Open weights: no
- Context / max output: 1,050,000 / 128,000
- Price per million tokens: $2 input, $12 output
- Reasoning levels: none, low, medium, high, xhigh, max
- Modalities: text, image
- Best for: Mixed workloads; Teams standardising on one model
- Watch out: 10x Luna's input price; check whether Luna already suffices for the workload.
- Verified: 2026-08-12
- Benchmarks: Artificial Analysis Intelligence Index 57 (max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 70% at $4.95/task

## GPT-5.6 Luna

- Profile: https://aichoiceengine.com/ai-models/gpt-5-6-luna
- Provider: OpenAI
- API id: not published
- Open weights: no
- Context / max output: 1,050,000 / 128,000
- Price per million tokens: $0.20 input, $1.20 output
- Reasoning levels: none, low, medium, high, xhigh, max
- Modalities: text, image
- Best for: Classification; Summarisation; High-volume chat; ChatGPT Free/Go default
- Watch out: The bare `gpt-5.6` alias routes to Sol at 25x the input price — specify the full id. ChatGPT Free→Luna is consumer Chat only; Work/Codex/API Luna builds did not change with the Aug 6 Chat refresh.
- Verified: 2026-08-12
- Benchmarks: Artificial Analysis Intelligence Index 52 (max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 67% at $3.03/task

## Claude Opus 5

- Profile: https://aichoiceengine.com/ai-models/claude-opus-5
- Provider: Anthropic
- API id: not published
- Open weights: no
- Context / max output: 1,000,000 / 128,000
- Price per million tokens: $5 input, $25 output
- Reasoning levels: low, medium, high, xhigh, max
- Modalities: text, image
- Best for: Agentic coding; Long-context analysis; Multi-step engineering work
- Watch out: Output tokens cost five times input — long generations dominate the bill.
- Verified: 2026-08-07
- Benchmarks: Artificial Analysis Intelligence Index 63 (max) (2026-08-14, Artificial Analysis); Artificial Analysis Intelligence Index 63 (xhigh) (2026-08-14, Artificial Analysis); Artificial Analysis Intelligence Index 61 (high) (2026-08-14, Artificial Analysis)
- DeepSWE Pass@1 74% at $11.84/task

## Claude Fable 5

- Profile: https://aichoiceengine.com/ai-models/claude-fable-5
- Provider: Anthropic
- API id: not published
- Open weights: no
- Context / max output: 1,000,000 / 128,000
- Price per million tokens: $10 input, $50 output
- Reasoning levels: low, medium, high, xhigh, max
- Modalities: text, image
- Best for: Work where capability outweighs cost entirely
- Watch out: Opus 5 scored higher on the Artificial Analysis Intelligence Index at half the price — justify the premium.
- Verified: 2026-08-07
- Benchmarks: Artificial Analysis Intelligence Index 62 (max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 70% at $21.63/task

## Claude Sonnet 5

- Profile: https://aichoiceengine.com/ai-models/claude-sonnet-5
- Provider: Anthropic
- API id: not published
- Open weights: no
- Context / max output: 1,000,000 / 128,000
- Price per million tokens: $2 input, $10 output
- Reasoning levels: low, medium, high, xhigh, max
- Modalities: text, image
- Best for: Everyday generation; Short, well-scoped tasks
- Watch out: Cheaper per token than Opus 5 but $26.40 per DeepSWE task against Opus 5's $11.84 — it used 214K output tokens over 268 steps.
- Verified: 2026-08-12
- Benchmarks: Artificial Analysis Intelligence Index 55 (max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 54% at $26.4/task

## Gemini 3.1 Pro

- Profile: https://aichoiceengine.com/ai-models/gemini-3-1-pro
- Provider: Google
- API id: not published
- Open weights: no
- Context / max output: 1,048,576 / 65,536
- Price per million tokens: $2 input, $12 output
- Reasoning levels: not published
- Modalities: text, image, video, audio, pdf
- Best for: Very long documents; Multimodal input; Google Workspace integration
- Watch out: Scored 12% on DeepSWE, far below every other model here including budget tiers — weak on agentic coding specifically. Pricing also steps up to $4/$18 above 200K tokens. Wire calls to `gemini-3.1-pro-preview`.
- Verified: 2026-08-09
- Benchmarks: Artificial Analysis Intelligence Index 48 (2026-08-09, Artificial Analysis)
- DeepSWE Pass@1 12% at $2.14/task

## Gemini 3.6 Flash

- Profile: https://aichoiceengine.com/ai-models/gemini-3-6-flash
- Provider: Google
- API id: not published
- Open weights: no
- Context / max output: 1,048,576 / 65,536
- Price per million tokens: $0.75 input, $3.75 output
- Reasoning levels: minimal, low, medium, high
- Modalities: text, image, video, audio, pdf
- Best for: Multimodal pipelines; High-volume processing; Video and audio input
- Watch out: Introductory $0.75 input is 3.75× Luna; the $1.50 rate returns on 2027-01-01. 66K max output (65,536 tokens) is roughly half Luna's. Prefer 3.7 Flash for new coding-agent work unless you are pinned to 3.6.
- Verified: 2026-08-14
- Benchmarks: Artificial Analysis Intelligence Index 52 (high) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 47% at $4.42/task

## Gemini 3.7 Flash

- Profile: https://aichoiceengine.com/ai-models/gemini-3-7-flash
- Provider: Google
- API id: gemini-3.7-flash
- Open weights: no
- Context / max output: 1,048,576 / 65,536
- Price per million tokens: $0.75 input, $3.75 output
- Reasoning levels: low, medium, high
- Modalities: text, image, video, audio, pdf
- Best for: Coding agents; Multimodal agent loops; High-volume Google API work
- Watch out: Intro price doubles on 2027-01-01. DeepSWE 1.1 best-effort (high) is 65.3% at $2.18/task; medium scores 65.5% a few cents cheaper. 66K max output (65,536 tokens). Knowledge cutoff is March 2026 on some domains (January 2025 on others).
- Verified: 2026-08-14
- Benchmarks: Artificial Analysis Intelligence Index 56 (high) (2026-08-14, Artificial Analysis)
- DeepSWE Pass@1 65% at $2.18/task

## Gemini 3.5 Flash-Lite

- Profile: https://aichoiceengine.com/ai-models/gemini-3-5-flash-lite
- Provider: Google
- API id: not published
- Open weights: no
- Context / max output: 1,048,576 / 65,536
- Price per million tokens: $0.30 input, $2.50 output
- Reasoning levels: not published
- Modalities: text, image, video, audio, pdf
- Best for: High-volume agents; Document pipelines; Cost-sensitive multimodal work
- Watch out: Cheaper than 3.6 and 3.7 Flash, not stronger on hard reasoning — pick it for throughput, not frontier coding.
- Verified: 2026-08-06
- Benchmarks: Artificial Analysis Intelligence Index 37 (2026-08-08, Artificial Analysis)

## Kimi K3

- Profile: https://aichoiceengine.com/ai-models/kimi-k3
- Provider: Moonshot AI
- API id: not published
- Open weights: yes
- Context / max output: 1,048,576 / 1,000,000
- Price per million tokens: $3 input, $15 output
- Reasoning levels: low, high, max
- Modalities: text, image, video
- Best for: Agentic coding without vendor lock-in; Self-hosting at frontier quality; Long-context work
- Watch out: 2.8T parameters means self-hosting is a datacentre exercise, not a workstation one — open weights here mean provider choice, not local inference.
- Verified: 2026-08-04
- Benchmarks: Artificial Analysis Intelligence Index 60 (max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 69% at $4.65/task

## Kimi K2.6

- Profile: https://aichoiceengine.com/ai-models/kimi-k2-6
- Provider: Moonshot AI
- API id: not published
- Open weights: yes
- Context / max output: 262,144 / not verified
- Price per million tokens: $0.95 input, $4 output
- Reasoning levels: not published
- Modalities: text, image, video
- Best for: Cost-sensitive volume; Self-hosting; Avoiding vendor lock-in
- Watch out: 262K context is among the smaller windows here — a real constraint on long-document work.
- Verified: 2026-08-12
- Benchmarks: Artificial Analysis Intelligence Index 45 (2026-08-08, Artificial Analysis)

## GLM 5.2

- Profile: https://aichoiceengine.com/ai-models/glm-5-2
- Provider: Z.ai
- API id: not published
- Open weights: yes
- Context / max output: 1,000,000 / 128,000
- Price per million tokens: $1.40 input, $4.40 output
- Reasoning levels: high, max
- Modalities: text
- Best for: Open-weight deployments; Long-horizon coding agents; Long context on a budget
- Watch out: Ecosystem tooling is thinner than the closed frontier labs — budget integration time. Parameter count is not published on the current Z.ai model card, so it stays unverified here. GLM Coding Plan now defaults to GLM 5.3; requests for 5.2/5.1 on that plan are routed to 5.3. This row remains the open-weight / token-list API identity at $1.40/$4.40.
- Verified: 2026-08-14
- Benchmarks: Artificial Analysis Intelligence Index 53 (max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 44% at $3.92/task

## GLM 5.3

- Profile: https://aichoiceengine.com/ai-models/glm-5-3
- Provider: Z.ai
- API id: glm-5.3
- Open weights: no
- Context / max output: 1,000,000 / 128,000
- Price per million tokens: Not verified input, Not verified output
- Reasoning levels: low, high, max
- Modalities: text
- Best for: GLM Coding Plan agent work; Long-horizon coding in Claude Code / OpenCode / Cline; Teams already on Z.ai subscriptions
- Watch out: Thinking cannot be disabled (`thinking.type: disabled` fails); `reasoning_effort` is low / high / max (default max). Open weights are promised about two weeks after launch pending safety review — not a downloadable checkpoint today. Vendor DeepSWE v1.1 66.9 used mini-swe-agent at 400K context, not the public Datacurve v1.1 board (snapshot generated 2026-08-13, before this launch).
- Verified: 2026-08-14

## Grok 4.5

- Profile: https://aichoiceengine.com/ai-models/grok-4-5
- Provider: SpaceXAI
- API id: not published
- Open weights: no
- Context / max output: 500,000 / not verified
- Price per million tokens: $2 input, $6 output
- Reasoning levels: low, medium, high
- Modalities: text, image
- Best for: Real-time data queries; Technical tasks; Cost-sensitive mid-tier work
- Watch out: Token rates double above a 200K-token prompt, and web/X search tool calls bill separately. xAI's current code/chat default is Grok 4.6 at the same $2/$6 band.
- Verified: 2026-08-12
- Benchmarks: Artificial Analysis Intelligence Index 56 (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 54% at $2.42/task

## Grok 4.6

- Profile: https://aichoiceengine.com/ai-models/grok-4-6
- Provider: SpaceXAI
- API id: not published
- Open weights: no
- Context / max output: 500,000 / not verified
- Price per million tokens: $2 input, $6 output
- Reasoning levels: low, medium, high, xhigh
- Modalities: text, image
- Best for: Coding agents; Chat and knowledge work; Cost-sensitive frontier work
- Watch out: Token rates double above a 200K-token prompt. DeepSWE 1.1 best-effort for this family is the xhigh row at 66.7% (the catalog table rounds Pass@1 to 67%), not the higher-scoring medium config.
- Verified: 2026-08-12
- Benchmarks: Artificial Analysis Intelligence Index 61 (high) (2026-08-12, Artificial Analysis)
- DeepSWE Pass@1 67% at $5.5/task

## Claude Haiku 4.5

- Profile: https://aichoiceengine.com/ai-models/claude-haiku-4-5
- Provider: Anthropic
- API id: not published
- Open weights: no
- Context / max output: 200,000 / 64,000
- Price per million tokens: $1 input, $5 output
- Reasoning levels: not published
- Modalities: text, image
- Best for: High-volume classification; Routine extraction; Cost-sensitive Anthropic stacks
- Watch out: A generation behind the 5-series models it sits alongside in the lineup.
- Verified: 2026-08-03
- Benchmarks: Artificial Analysis Intelligence Index 24 (Non-reasoning) (2026-08-08, Artificial Analysis)

## DeepSeek V4-Flash

- Profile: https://aichoiceengine.com/ai-models/deepseek-v4-flash
- Provider: DeepSeek
- API id: deepseek-v4-flash
- Open weights: yes
- Context / max output: 1,000,000 / 384,000
- Price per million tokens: $0.14 input, $0.28 output
- Reasoning levels: low, high, max
- Modalities: text
- Best for: Cost-sensitive hosted agents; High-volume coding assist; Open-weight deployments with cluster VRAM
- Watch out: API list prices move to peak/off-peak from 2026-08-16 16:00 UTC (see pricing note). Self-hosting the ~284B MoE still needs roughly 90–170GB+ class VRAM depending on quant — not a laptop budget build. DeepSWE snapshot also uses the shared mini-swe-agent harness — verify cost, effort, and serving conditions before treating the result as a forecast.
- Verified: 2026-08-14
- Benchmarks: Artificial Analysis Intelligence Index 52 (DeepSeek V4-Flash-0731, max) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 53% at $0.1/task

## DeepSeek V4-Pro

- Profile: https://aichoiceengine.com/ai-models/deepseek-v4-pro
- Provider: DeepSeek
- API id: deepseek-v4-pro
- Open weights: yes
- Context / max output: 1,000,000 / 384,000
- Price per million tokens: $0.435 input, $0.87 output
- Reasoning levels: low, high, max
- Modalities: text
- Best for: Harder reasoning on a budget; Open-weight deployments; Agentic coding when Flash is not enough
- Watch out: The April preview and 2026-08-13 GA share one API id — pin the date of any score you cite. Artificial Analysis currently evaluates the 0813 (max) checkpoint at Intelligence Index 53, Terminal-Bench v2.1 78.7%, and GPQA Diamond 92.8%. Peak/off-peak API rates take effect 2026-08-16 16:00 UTC.
- Verified: 2026-08-14
- Benchmarks: Artificial Analysis Intelligence Index 53 (DeepSeek V4-Pro-0813, max) (2026-08-14, Artificial Analysis)
- DeepSWE Pass@1 63% at $0.24/task

## Ling 3.0 Flash

- Profile: https://aichoiceengine.com/ai-models/ling-3-0-flash
- Provider: InclusionAI
- API id: not published
- Open weights: yes
- Context / max output: 262,144 / not verified
- Price per million tokens: $0.075 input, $0.22 output
- Reasoning levels: not published
- Modalities: text
- Best for: High-volume agent loops; Cost-sensitive chat and extraction; Fast hybrid reasoning
- Watch out: Native context is 256Ki tokens (262,144). Confirm serving provider and current list before committing volume.
- Verified: 2026-08-13
- Benchmarks: Artificial Analysis Intelligence Index 38 (2026-08-08, Artificial Analysis)

## Qwen3.8-Max

- Profile: https://aichoiceengine.com/ai-models/qwen-3-8-max
- Provider: Qwen
- API id: not published
- Open weights: no
- Context / max output: 1,000,000 / 131,072
- Price per million tokens: $2 input, $6 output
- Reasoning levels: not published
- Modalities: text, image, video
- Best for: Agentic coding; Multimodal knowledge work; Long-context delivery
- Watch out: This row is the hosted Qwen3.8-Max API, not the downloadable checkpoint. The open-weight Qwen3.8-2.4T-A95B card is catalogued separately and does not include the API's vision, non-thinking, or built-in tool extras.
- Verified: 2026-08-13
- Benchmarks: Artificial Analysis Intelligence Index 58 (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 57% at $3.73/task

## Qwen3.8-2.4T-A95B

- Profile: https://aichoiceengine.com/ai-models/qwen-3-8-2-4t-a95b
- Provider: Qwen
- API id: not published
- Open weights: yes
- Context / max output: 262,144 / 131,072
- Price per million tokens: Not verified input, Not verified output
- Reasoning levels: low, medium, xhigh
- Modalities: text
- Best for: Self-hosting the Qwen3.8 flagship; vLLM / SGLang serving; Evaluating the open Max-class weights
- Watch out: Native context is 262,144 tokens (YaRN extends toward ~1,010,000). Hugging Face best-practice caps: 262,144 reasoning tokens and 131,072 final-response tokens. Thinking cannot be disabled. Qwen3.8-Max License: MaaS / AI Work Assistant businesses over $50M trailing 12-month revenue need a separate commercial licence. Self-hosting still needs datacentre-class memory. Do not treat this checkpoint as the multimodal 1M-context API.
- Verified: 2026-08-14

## Qwen3.8-27B

- Profile: https://aichoiceengine.com/ai-models/qwen-3-8-27b
- Provider: Qwen
- API id: not published
- Open weights: yes
- Context / max output: 262,144 / 131,072
- Price per million tokens: Not verified input, Not verified output
- Reasoning levels: low, medium, xhigh
- Modalities: text, image, video
- Best for: Local multimodal agents; Self-hosting a dense 27B VLM; Apache 2.0 deployments
- Watch out: The Hugging Face repo was created 2026-08-05; this row uses the 2026-08-14 card revision. Hosted 1M-context API is documented as coming soon. Vendor SWE-bench Pro 61.7 and DeepSWE 42.2 used a Claude Code harness, not swebench.com or the public Datacurve DeepSWE board. YaRN can extend context toward 1M.
- Verified: 2026-08-14

## Muse Spark 1.2

- Profile: https://aichoiceengine.com/ai-models/muse-spark-1-2
- Provider: Meta
- API id: not published
- Open weights: no
- Context / max output: 1,048,576 / not verified
- Price per million tokens: $1.25 input, $4.25 output
- Reasoning levels: minimal, low, medium, high, xhigh
- Modalities: text, image, video, pdf
- Best for: Terminal coding agents; Long-horizon refactors; Meta Model API stacks
- Watch out: A cheaper contributor tier (`muse-spark-1.2-contributor`) permits training use of prompts and completions; standard-tier `muse-spark-1.2` does not. `reasoning_effort: none` is unsupported (HTTP 400). Weights are not marked open until an HF/license card is verified. For local agentic now, prefer Muse Glimmer 30B.
- Verified: 2026-08-13
- Benchmarks: Artificial Analysis Intelligence Index 57 (xhigh) (2026-08-08, Artificial Analysis)
- DeepSWE Pass@1 55% at $3.7/task

## Muse Glimmer 30B

- Profile: https://aichoiceengine.com/ai-models/muse-glimmer-30b
- Provider: Meta
- API id: not published
- Open weights: yes
- Context / max output: 131,072 / not verified
- Price per million tokens: $0 input, $0 output
- Reasoning levels: not published
- Modalities: text, image
- Best for: Local agents; On-device coding + tool use; Privacy-sensitive multimodal work
- Watch out: Full BF16 needs far more than 24GB — plan on official GGUF / 4-bit packs (and mmproj for vision). Agentic quality ≠ frontier Muse Spark API; validate on your workflows.
- Verified: 2026-08-10

