Skip to main content

Category crowns

The leading pick depends on the job

These are editorial category leaders from the current catalog, not universal rankings. Each crown includes the reason it leads and the caveat most likely to change your decision.

Hard coding and reasoning

Claude Opus 5

The current frontier shortlist combines long context, strong agentic positioning, and a dedicated coding workflow ecosystem.

Caveat: Verify on your repository; frontier output pricing can dominate the cost. Grok 4.6 Extra High leads CursorBench correctness in this catalog, so IDE-specific evals can disagree with this crown.

See the evidence

Best cheap OpenAI tier

GPT-5.6 Luna

Inside OpenAI's family, Luna is the volume tier for repeatable workloads where the cheaper input rate matters.

Caveat: DeepSeek V4-Flash undercuts Luna on API price and is MIT-licensed — pick Luna when you specifically need the OpenAI stack, not as the overall budget winner.

See the evidence

Best open price-performance

DeepSeek V4-Flash

MIT weights, a long context, and hosted rates far below closed budget tiers make it the catalog's open price-performance pick.

Caveat: Text-only today; self-hosting still needs substantial serving capacity and independent validation.

See the evidence

Cheapest multimodal API tier

Gemini 3.5 Flash-Lite

Among catalogued multimodal models, Flash-Lite is the lowest-priced Google Flash-class API option for high-volume document and media pipelines.

Caveat: Cheapest multimodal is not strongest multimodal — pick 3.6 Flash or a frontier model when quality matters more than throughput.

See the evidence

Frontier open-weight evaluation

Kimi K3

It is a useful open-weight frontier test for teams evaluating provider independence and long-context work.

Caveat: The licence and datacentre-scale serving footprint are part of the decision.

See the evidence

Quality versus price

The best headline score can still lose on total cost.

Compare successful outputs, retries, latency, and human review alongside token price. A lower-priced model that needs twice the supervision is not automatically the better value.

Read the benchmark glossary