In short
- Gemini 3.7 Flash constructed a playable browser recreation from a single immediate in 2 minutes and 13 seconds, a job Gemini 3.6 Flash failed outright three weeks earlier.
- It failed our bridge logic puzzle with the identical mistaken reply as Claude Fable 5, and stopped wanting truly calculating the maths drawback it appropriately arrange.
- The mannequin runs at 75 cents per million enter tokens by means of December 31, half of three.6 Flash’s price, earlier than doubling to $1.50 on January 1.
Google shipped Gemini 3.7 Flash on August 13, usually out there in additional than 160 international locations on day one. It takes as much as one million enter tokens, returns 64,000, reads photographs, video, audio and PDFs, and may name instruments and drive a pc.
Flash has by no means been the mannequin you attain for when an issue is tough. It is the one you employ to type textual content, compact agent periods earlier than they collapse underneath their very own context, and summarize paperwork you do not need to pay a flagship to learn.

Judged in opposition to these sorts of jobs, 3.7 Flash is an actual improve. Judged in opposition to every thing else, it is a competent mannequin that will get outwritten by software program you may obtain totally free.
Google’s personal benchmark sheet places 3.7 Flash forward of Claude Sonnet 5 and GPT-5.6 Terra on 11 of 18 examined classes. The headline numbers are 1,588 Elo on Code Area’s internet improvement board and 30.4% on AutomationBench. Each come from Google’s methodology, so deal with the lead as the corporate’s declare moderately than settled reality.
We examined the mannequin to see if it lives as much as Google’s claims. These are our outcomes.
Coding: Can it construct one thing that runs on the primary strive?
This take a look at measures zero-shot code technology—whether or not a mannequin turns one instruction into working software program with no examples to repeat and no likelihood to repair itself. We hand over a single immediate for a browser recreation and ship no matter comes again, bugs included. No follow-ups, no error studies, no second try.
Gemini 3.7 Flash handed in 2 minutes and 13 seconds. The sport was playable on the primary run, the syntax was clear, the collision and scoring logic held, and the visible high quality sat above what the worth tier suggests.
The comparability that issues right here is not a flagship. It is Gemini 3.6 Flash, launched July 21, which couldn’t produce a working file in any respect. Its HTML was malformed, parts did not render, and follow-up prompts asking it to restore its personal output went nowhere.
We ended up handing that wreckage to DeepSeek, which discovered 11 bugs and shipped 8 fixes to make it playable. Three weeks later the identical product line wants no rescue, and the consequence sits near what GPT-5.6 Sol produced in our July assessment.
Gemini 3.7 Flash wins this one outright, and it is the only strongest motive to modify. The caveat is that it executes specs moderately than inventing them, so a imprecise immediate will get you a imprecise recreation.
You’ll be able to strive Gemini 3.7 Flash’s recreation right here.
Inventive writing: Can it maintain a paradox and write a sentence?
This part assessments two issues without delay: literary high quality, and whether or not a mannequin can obey a structural rule throughout 1000’s of phrases. The immediate sends Jose Lanz from 2150 again to the yr 1000 and calls for a closed causal loop—his intervention have to be the factor that creates the long run he got here to stop.
The rule that decides the take a look at is the final clause: He can’t perceive what he did till he’s residence.
Gemini 3.7 Flash generated an honest consequence. Jose fires an entropic cannon right into a Pyrenean fissure, by chance forges an obelisk that enslaves Twenty second-century Iberia, and grasps the entire loop whereas nonetheless standing within the mud a thousand years early: “It was the bottom of the Cinder Spire.”
The plot equipment is definitely fairly sound. The story mentions a falling star that historical monks witnessed and make clear it was the flash of Jose’s personal arrival, and the weapon he dropped at erase the anomaly is what forges it. Its closing line—”It had merely been ready for him to finish it”—lands the determinism the immediate requested for.
However for these used to it, the story screams “AI.” Virtually each noun arrives with two adjectives bolted on: “hyper-luminescent towers,” “damp, moss-choked earth,” “thick, obsidian hair.” That’s the texture of a mannequin selecting essentially the most possible subsequent phrase as an alternative of selecting one, and it produces collisions like a monolith “buzzing with a low-frequency hum.”
We in contrast it in opposition to Qwopus3.5-27B-v3, a neighborhood fine-tune of Qwen3.5-27B that distills Claude Opus-style reasoning and runs on a single shopper GPU for nothing per question. It obeyed the rule Gemini broke.
Jose kills a monk at San Millán de la Cogolla, an actual La Rioja monastery that truly mattered across the yr 1000, and solely understands what he did after returning to 2150 and discovering his personal DNA in a wax-sealed codex.
Qwopus just isn’t completely clear both. It dumped its whole planning scratchpad above the story, typos included, and its ultimate part breaks the closed loop it spent eight sections constructing by letting Jose return and make things better.
However all issues thought of, Qwopus takes it. Gemini delivered the tidier package deal and the extra disciplined ending, but it surely failed the one instruction the immediate was constructed round, and a free mannequin working on a gaming GPU wrote the higher story.
Associative pondering: Can a metaphor carry an argument?
This take a look at measures associative reasoning—whether or not a mannequin can generate hyperlinks between unrelated ideas with out having to elucidate itself. The immediate asks for an outline of a twig, makes use of that description to elucidate employee exploitation and the worship of the wealthy, then requires the argument to dissolve into an outline of a lettuce.
Signposting is the failure mode. Naming the metaphor kills it.
Gemini names it within the opening line of its second paragraph: “That is the exact mechanics of the fashionable proletariat.” Every thing earlier than that was working.
Among the imagery earns its place. The employee receives “simply sufficient bark to remain inflexible for one more week of output,” and the fallen twigs are conditioned to consider that with sufficient rigidity any considered one of them would possibly grow to be a trunk. The paragraph containing the primary of these additionally comprises a employee “certain to an huge, top-heavy company hierarchy.”
Not unhealthy by way of logic and construction.
The dissolve is the true collapse. Gemini narrates the transition moderately than performing it—the hierarchies “crumble, dissolving into the quiet, humble actuality of the natural world beneath”—after which a lettuce merely seems, unconnected to something earlier than it.
GPT-5.6 Sol rots the twig into soil and grows the lettuce out of it: “Rain enters the grain. Fibers loosen, darken.” The argument arrives buried within the object too, with wealth reframed as a language of advantage the place “The mansion signifies intelligence.”
GPT-5.6 Sol wins by so much. Gemini produced a pleasant particular person line, but it surely defined its personal metaphor after which skipped the transition the immediate was particularly testing.
Logic: Does it learn the immediate or acknowledge the puzzle?
This take a look at measures non-math reasoning, and particularly whether or not a mannequin reads the query in entrance of it or pattern-matches to a model it memorized. Our bridge immediate offers 4 individuals one torch and crossing instances of 1, 2, 5 and 10 minutes, then asks how briskly they’ll all get throughout.
The trick is what the immediate leaves out. It by no means says solely two individuals may be on the bridge without delay, so the reply is 10 minutes—everybody walks over collectively on the slowest individual’s tempo.
Gemini answered 17 minutes, working the memorized five-step shuffle from the textbook model of the puzzle. It acknowledged the constraint as reality with out ever checking whether or not we had written it.
Its seen reasoning is worse than its reply. The hint argues that sending the 2 slowest throughout collectively could be inefficient as a result of somebody must stroll the torch again—after which the ultimate reply sends them throughout collectively anyway. It contradicts itself inside a single response and studies the consequence with complete confidence.
Claude Fable 5 landed on the identical mistaken quantity again in July. It opened by declaring what it was assuming, “assuming the traditional constraint that the bridge holds solely two individuals at a time,” which is the distinction between a mistaken reply you may catch and one you may’t.
No one wins. Fable takes it on transparency alone, and the false-confidence throughout Gemini agent runs exhibits up right here in a puzzle you may test by hand.
Math: Does it end the job?
This take a look at measures symbolic arithmetic nicely past shopper use, plus one thing less complicated—whether or not the mannequin does what it was requested. The immediate requires a degree-19 odd monic polynomial with actual coefficients and linear coefficient -19, whose curve splits into at the very least three irreducible elements, after which asks for p(19).
Each fashions discovered the identical door. Gemini and Qwen 3.7 Max Preview each recognized the Dickson polynomial, solved the constraint to repair its parameter at 1, and derived the closed type appropriately.
Then Gemini stopped. It printed p(19) as an unevaluated expression involving the nineteenth energy of a sq. root, by no means produced the quantity, and by no means demonstrated the part depend the immediate additionally demanded. It delivered all of this inside a styled HTML web page with CSS and a drop shadow that no person requested.
Qwen completed. It gave the total factorization into 10 elements—one linear, 9 quadratic—ran the recurrence out to 1,876,572,071,974,094,803,391,179, and cross-checked the consequence modularly. We verified that determine independently in SymPy and it holds.
Qwen wins on the one criterion that mattered. Gemini began advantageous and determined to skip the arithmetic, which is a wierd place to cease.
Conclusion
Gemini 3.7 Flash is definitely worth the swap if you’re already inside Google’s ecosystem. It’s dramatically higher at code than the mannequin it replaces, quick sufficient to matter for agent work, and low-cost sufficient that working it at quantity is a rounding error.
Its strengths are execution and construction. Give it an in depth spec and it’ll construct the factor, maintain a plot collectively, and hold the causal logic coherent throughout 1000’s of phrases.
Its weaknesses are creativity and reasoning. The writing is predictable sufficient to establish as machine-made on sight, and the mannequin asserts mistaken solutions with out flagging the belief that made them mistaken.
The worth is the strongest argument for it. At 75 cents per million enter tokens and $3.75 output, it undercuts GPT-5.6 Sol’s $5 enter price by 85% and prices half what 3.6 Flash did at launch.
The argument in opposition to it’s a free 27B mannequin on a gaming GPU that wrote a greater story and charged nothing to do it. Google’s introductory price expires December 31, when enter doubles to $1.50 and output to $7.50.
Each day Debrief Publication
Begin day by day with the highest information tales proper now, plus authentic options, a podcast, movies and extra.
