Decision guide
Should you switch volume work from GPT-5.6 Luna to Gemini 3.7 Flash?
3.7 Flash is Google's current coding/agent Flash row. Catalog list rates, DeepSWE, and context/output caps — not a headline ranking — decide whether it beats Luna on your mix.
Switch if
- Your pipelines send image, video, audio, or PDF input (catalogued 3.7 Flash modalities) that Luna's text+image row does not cover.
- You want Google's current Flash workhorse: catalog DeepSWE 65% Pass@1 at $2.18/task versus Luna's 67% at $3.03/task.
- You are already on Google Cloud or Workspace and the introductory $0.75/$3.75 list (through 2026-12-31) fits a controlled batch test.
Stay if
- Luna already meets your pass rate. Catalog standard-tier list is $0.20/$1.20 versus 3.7 Flash $0.75/$3.75 (then $1.50/$7.50 from 2027-01-01).
- Your OpenAI integrations, tool schemas, and billing are not worth revisiting for a few DeepSWE points.
- You need 128K max output (Luna) rather than 3.7 Flash's 65,536-token cap.
Check before production traffic moves
- 01Run the same high-volume batch on both models with identical validators.
- 02Include multimodal inputs only where your production workload actually sends them.
- 03Compare p95 latency, rate limits, and retry behaviour — quote catalog list rates, not a guessed blend.
This is a budget-tier swap between providers. Catalogued 3.7 Flash is Google's 2026-08-13 Flash workhorse at introductory $0.75/$3.75, DeepSWE 65% at $2.18/task, and Artificial Analysis Intelligence Index 56 (high). Luna is OpenAI's volume tier at $0.20/$1.20, DeepSWE 67% at $3.03/task, and Intelligence Index 52 (max). Do not treat 3.6 Flash as the current Google default.
Run a controlled trial first
A second option in a labelled pilot is cheaper than a production cutover. Score the trial, then migrate or roll back on evidence.
Pilot plan
- Select a high-volume slice with repeatable inputs and a known error budget.
- Run Luna and 3.7 Flash on the same batch, including structured outputs and tool calls.
- Track cost per accepted task, latency p95, and reviewer minutes on a human sample.
Score the trial
- Modality fit
- Pass rate on the inputs your pipeline actually sends.
- Cost per success
- Tokens, retries, and review divided by accepted tasks.
- Throughput
- Rate limits, queue time, and p95 latency under load.
- Integration cost
- Adapter changes, auth, logging, and support path.
Migration sequence
- Add 3.7 Flash behind a feature flag with Luna as the default fallback.
- Route only the pilot cohort first; cap spend and monitor validator failures.
- Expand only when cost per success and latency stay within tolerance.
Rollback plan
- Keep Luna routing and prompt templates provider-specific for instant switch-back.
- Stop 3.7 Flash traffic if validation or latency breaches thresholds; replay safe failures.
- Archive batch logs so a later re-test can separate release changes from integration bugs.
Hidden costs to price
- Introductory Flash rates double on 2027-01-01 — forecast both the $0.75/$3.75 and $1.50/$7.50 catalog bands.
- A second cloud provider adds billing, policy review, and on-call complexity.
- Search grounding and tools on Gemini can add separate line items.
Evidence to check