posted an update

Gemini 3.7 Flash is now the only model in production

We promoted gemini-3.7-flash on 2026-08-14. Not because a new model shipped, but because a pre-registered gate told us to.

The test: unpc:en-fr dev split, 158 segments of real UN legal text. Prompts byte-identical between the two arms, the model as the only variable, and the decision metric named before the run, so that a favourable move on one of four metrics could not be selected after the fact.

What moved:

  • chrF++ 65.16 to 66.34, at p=0.0020 on a paired bootstrap
  • BLEU +1.55, TER improved
  • 25.9% fewer output tokens
  • 36.7% less wall clock

What we are not claiming. Launch coverage widely described 3.7 Flash as half the price of its predecessor. Relative to what we were actually paying, it is not: both sit at $0.75 per million input tokens, because 3.6 was discounted to match. Our cost per segment fell because the model emits fewer output tokens at the same rate, not because the rate changed. We measured instead of repeating the launch post.

One honest wrinkle: 3.7 Flash ignores our requested temperature, so this system is no longer bit-reproducible between runs. That is written into the run record rather than quietly dropped from our reproducibility claim.

3.6 Flash stays in the price book. Every figure we published before August 14 was measured on it, and those rows have to stay priceable.

Full run record: benchmarks/journal/2026-08-14/

Log in or sign up for Devpost to join the conversation.