Skip to main content
← All switch guides

Decision guide

Should you switch volume work from GPT-5.6 Luna to Gemini 3.7 Flash?

3.7 Flash is Google's current coding/agent Flash row. Catalog list rates, DeepSWE, and context/output caps — not a headline ranking — decide whether it beats Luna on your mix.

Switch if

  • Your pipelines send image, video, audio, or PDF input (catalogued 3.7 Flash modalities) that Luna's text+image row does not cover.
  • You want Google's current Flash workhorse: catalog DeepSWE 65% Pass@1 at $2.18/task versus Luna's 67% at $3.03/task.
  • You are already on Google Cloud or Workspace and the introductory $0.75/$3.75 list (through 2026-12-31) fits a controlled batch test.

Stay if

  • Luna already meets your pass rate. Catalog standard-tier list is $0.20/$1.20 versus 3.7 Flash $0.75/$3.75 (then $1.50/$7.50 from 2027-01-01).
  • Your OpenAI integrations, tool schemas, and billing are not worth revisiting for a few DeepSWE points.
  • You need 128K max output (Luna) rather than 3.7 Flash's 65,536-token cap.

Check before production traffic moves

  1. 01Run the same high-volume batch on both models with identical validators.
  2. 02Include multimodal inputs only where your production workload actually sends them.
  3. 03Compare p95 latency, rate limits, and retry behaviour — quote catalog list rates, not a guessed blend.

This is a budget-tier swap between providers. Catalogued 3.7 Flash is Google's 2026-08-13 Flash workhorse at introductory $0.75/$3.75, DeepSWE 65% at $2.18/task, and Artificial Analysis Intelligence Index 56 (high). Luna is OpenAI's volume tier at $0.20/$1.20, DeepSWE 67% at $3.03/task, and Intelligence Index 52 (max). Do not treat 3.6 Flash as the current Google default.

Run a controlled trial first

A second option in a labelled pilot is cheaper than a production cutover. Score the trial, then migrate or roll back on evidence.

Pilot plan

  • Select a high-volume slice with repeatable inputs and a known error budget.
  • Run Luna and 3.7 Flash on the same batch, including structured outputs and tool calls.
  • Track cost per accepted task, latency p95, and reviewer minutes on a human sample.

Score the trial

Modality fit
Pass rate on the inputs your pipeline actually sends.
Cost per success
Tokens, retries, and review divided by accepted tasks.
Throughput
Rate limits, queue time, and p95 latency under load.
Integration cost
Adapter changes, auth, logging, and support path.

Migration sequence

  • Add 3.7 Flash behind a feature flag with Luna as the default fallback.
  • Route only the pilot cohort first; cap spend and monitor validator failures.
  • Expand only when cost per success and latency stay within tolerance.

Rollback plan

  • Keep Luna routing and prompt templates provider-specific for instant switch-back.
  • Stop 3.7 Flash traffic if validation or latency breaches thresholds; replay safe failures.
  • Archive batch logs so a later re-test can separate release changes from integration bugs.

Hidden costs to price

  • Introductory Flash rates double on 2027-01-01 — forecast both the $0.75/$3.75 and $1.50/$7.50 catalog bands.
  • A second cloud provider adds billing, policy review, and on-call complexity.
  • Search grounding and tools on Gemini can add separate line items.