Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper
Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?
At this point, I think they're mostly targeting Google One and Workspace subscribers, except doing worse compared to Microsoft because they don't have Microsoft's huge enterprise moat built from their DOS and Windows days.
Per-token cost isn't a great metric given that some use way more tokens than others.
They compare it to 5.6 Terra, however https://cognition.com/frontiercode puts Terra at about 1/2 the price
Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper
Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?
Somewhere in the same neighborhood as GPT 5.6 Tera and Sonnet 5, depending on the bench.
It can regularly cost more than Fable, take longer, and deliver far far lower quality.
I'm much more interested how this compares to Luna - which on price is terribly - but at least on quality the benchmarks make this look competitive / usable.
If Google continues monthly Flash releases like Sundar said they would, and they continue to have this much of an improvement in cost/quality - then in a few months this could reasonably be very competitive with the best of the best.
It is not there yet, but at least it's super fast, I guess.
Less than half its price.
More than 50% discount.
I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal.
Luna is similar, and also 8x cheaper. Source: artificialanalysis
The only benefit I can see is the speed, that looks to be outstanding, probably thanks to their TPUs.
That's why DS4 already had a huge price hike announcement.
Deepseek as a company can just increase prices for the crazily cheap cache they have, that's their only lever.
Well, compared to 2 months ago, it's no longer 100x more expensive for similar levels of quality...
If they continue monthly-ish releases by 3.9 - by Halloween - they should be close to the best in terms of what you get for what you pay for.
In 2 months, they've gone from basically the bottom of the pack to at least being somewhat usable and competitive.
OpenAI and Anthropic release in a month, and change things. OpenAI is claiming to be close to an Astra release - but that seems like a Fable type release - where they're just releasing a better more expensive model, not more cost effective models.
this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time.
sure we used to cling to gemini models in the past, demanding 2.5 models to continue to serve, but since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.
heck, even now I'm not sure I even care if they cut the pricing even lower. there are too many models with cheaper price and similar performance now.
They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5'
> since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.
This is my first hand experience. I spent at least $3000 on gemini-3-flash-preview. And exactly $0 total on (3.5+3.6+3.7)
> Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
> Coding and agentic tasks: Significantly higher quality on real-world software engineering and agentic benchmarks, improving issue resolution and reducing failed agent loops.
> Web development and stronger design parity: Generates higher-fidelity desktop and web application code directly from design mocks, with strong gains in design adherence and in auditing existing codebases against mocks to verify 1:1 design parity.
> Promotional pricing: Gemini 3.7 Flash will be available at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens. We’re also applying this new rate to 3.6 Flash. Introductory pricing expires on December 31, 2026; after, $1.50/1M input tokens and $7.50/1M output tokens will apply.
Still no sign of 3.5 Pro. Will have to test it, low expectations given every other model from the Gemini 3 lineage, but one can hope. Just struggle to understand the promotional pricing being temporary for four months. Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing?
It was probably to placate some kind of general internal pricing/revenue benchmark that doesn't account for new model releases. Politicians do shit like this incessantly and it reeks of bureaucracy.
"Hey, we are not in a race to the bottom. This is our usual pricing, but this now is a promotion because we know we're coming from behind and need to entice users."
They're drawing a line in the sand on monetisation and signalling that to everyone, while in reality offering it a deep discount (no idea if profitable or not) knowing that this model will probably be obsolete before then.
For Google, this is still gemini-3.1-pro-preview, right?
Flash is better than Pro for now.
Path A: Deprecated, do not dare use
Path B: Beta, do not rely
Unfortunately, it's often not strong enough for heavy refactoring and long running development loops.
but luna is hard to beat @ capability / cost
And it's arguably not crazy, at least if SemiAnalysis's estimates are to be believed:
* 20% of all TPU shipments from Q3 2026 through Q4 2027 are sold to SPVs serving Anthropic ($150B of contracted revenue); vs
* ~$12B ARR for Gemini.
https://newsletter.semianalysis.com/p/gemini-is-cooked-but-g...Because they compete for the same scarce resource, the result is a resource crunch for the group that's lost: https://www.latimes.com/business/story/2026-05-18/inside-ai-...
Citation needed.
also, why can't a massive company do two things?
The selling point for gemini continues to be speed and particularly end-to-end response time.
Edit: and Sol medium actually has the same AA intelligence score as Gemini 3.7, and has >7x fewer tokens, actually making it faster
I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost.
[edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge...
more of a Terra than Luna competitor which is an interesting positioning. I feel like differentiation at the mid-tier of models is pretty difficult.]
I wonder if this counts as evidence against that hypothesis? That multi-modal is struggling to keep up with SotA and the best they can offer is competent and fast?
Google probably crunches more tokens daily than the other labs combined, just because basically the entire global population uses Google (sans china) and Google has shoved Gemini into everything.
This feels absurd to me (my gut is "I want the SMARTEST model I can get!!"), but often I find that my experience of using a flash/sonnet model for every-day workhorse coding they are better.
Its not the same thing, but when I think of that I am reminded of working with some engineers in the past who are incredibly smart and have PhDs (or to put it another way, over-qualified) and they were crap engineers because they'd just not be able to focus on the task and ONLY the task at hand and would get easily distracted by the "why" or "more interesting" things when I just asked them to fix a simple bug or whatever. Again, its not the same thing at all, but it certainly comes to mind when I think of this or experience a pro/opus model suggesting we make huge refactors when a tactical fix is all that is required etc.
Of course, the opus-sized models are great when it comes to huge comprehension/research/debugging efforts where the deeper reasoning is actually useful.
In my evals 3.6 Flash (pre price change) was usually a bit more token efficient than 3 Flash, so I‘m expecting same or even lower cost-per-task on 3.7.
Maybe a play by Google to deprecate 3 Flash soon.
Google's Opus competitor is 3.1 Pro Preview which is essentially obsolete (competed with Opus 4.6). They do not have a Fable/Sol competitor.
I guess Google's betting on consumers being price-elastic (preferring to tradeoff intelligence for significant cost savings)
I hope so. It seems mind boggling to me that an user needs to surf around different sections (plural) of google cloud console, then this Vertex and do a dozen clicks to issue a simple key.
each new checkpoint can benefit from better reasoning training, RL on specific tasks and more synthetic data
So why do they seem to release around the same time ? my guess is because they time major releases around quarterly earnings, investor meetings and other important business milestones. Once one company announces a major update, the others also have an incentive to ship their latest checkpoint rather than look like they r falling behind.
So it's better than 3.6 Flash, at half the price. I've been pretty excited about Gemini models recently, they just feel so fast after spending most of the day at work waiting for Opus 5.
I think its the same price..
1. How the new model performs against the other top models in the same category.
2. The pricing of the new model against the other top models in the same category.
Original images: https://image.non.io/neonRamenDesigns.webp
Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7
Opus 5 build for comparison: https://html.non.io/neonRamen
Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.
The grok test was ran through the cursor cli agent however.
It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.
Same training dataset, same software, same hardware, same architecture...
I'm wondering what they changed actually for the model to be more powerful if the benchmark results are real and relevant.
Maybe just tweak settings or the reasoning prompts and called it a new version of their model?
We still don't have a 3.5 Pro, and along comes 3.7 Flash?!
> Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
Compare this to Luna which is at $0.2/1M input ($0.02 cached) and $1.2/1M output.
I also think Google is still the best at fitting the most overall intelligences into their models, but for some reason it seems like the model architecture is just bad.
Personally whenever I use Gemini I've just been using 3.1 Pro because I've had insane trouble with them getting things incorrect like this. Hopefully they'll fix it soon / they've fixed it with 3.7 Flash.
It did fine on my usual benchmark about configuring old Sparc hardware, maybe output slightly faster than before. Even included something new to check in the firmware.
GPT-5.6 or Claude models haven't delivered to me non-running code in ages.
Whenever I have Gemini in the flow, it's fast, but mistake riddled. I have low confidence in the output.
I've had some success with Opus driving Gemini models. It's pointless for GPT family since Sol is cheap enough or can drive terra/luna for arguably better performance, same speed, and better outcome.
As for all of the talk in this thread about modalities. Every SOTA model takes screenshots and verifies work now. Grok-4.6 does this, Luna does it, etc. They can also all work _from_ a screen shot or mockup provided.
I don't think it's a major selling point when every model can do it well and reasonably fast.
That said, eagerly awaiting "pro" and improvements to antigravity.
It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago!
(An ambitious pelican, let down by a flawed bicycle: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...)
Theoretically there is some difference between Fable and Opus or Grok and GPT, but at the end of the day I'd look at the bottom left of my screen and to my amusement find out that for the past 3-4 hours I've been using model ______.
If the results are semi-decent, I'd keep it on, if not - I'd randomly switch the model and try again.
Actual thing that would affect my selection would be a number of unused tokens I have left for a model ____ for this week.
Maybe it's cause I'm using those for programming and log parsing and all of them are decent enough, but other than that - there are no leaps I see.
I could probably do text only for my workflow (feature development/debugging for web microservices) but sometimes it is easier to just toss a screenshot into the Claude prompt, so that gives it an edge.
If your workflow is 100%, certifiably never ever going to involve an image, then yeah, this isn't going to be huge.
a decent model with a decent harness will determine when the knowledge base is lacking and attempt to fill the holes; thus the good general models can be very easily brought up to speed on niche domains.
Think like a regulator.
Some model are more aggressive by default, some are more verbose by default. To get the result you want for your specific application, you run experiment with prompts and parameters.
and now we have ai agents to automatic migrate the system with new models. in the past we would need to spend hours to design the prompts, then test the output, then write codes to babysitting it. nowadays any ai agent can do it effortlessly.
Personally, I feel like Google blundered on their pricing because while I was using the free version of the Gemini harness, they took away most of the free limits and made people move over to their Anti-Gravity harness for no apparent reason. I was about to splurge for a Pro sub since I already used Google for extra storage but putting up limits like they did made me not want to trust they wouldn't do more price shenanigans. Now their models are behind and it seems like they're scrambling.
Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though...
Good catch! You're right to point that out. My previous marketing copy missed that specific detail. Thank you for bringing it up!
Luna way cheaper. DeepSeek used to be, but I think it's somewhere on Sol's curve after the price hike.
If it implements something simple like a file export, it just knows that the file should have a meaningful name. Vibecoded feature beats most software's lazy "untitled.png".
So, yes, I want the smartest model even for simple stuff. Maybe especially for simple stuff because the tokens burned will be trivial so the cost doesn't give lower models a comparative advantage.
Consistently, lower intelligence models provide worse results in my own work. But I don't have evals on my side, just vibes.
I guess I need to try harder. :)
Opus can't generate images since A\ doesn't have a diffusion model.
spelk•1h ago
Introductory pricing until December 2026 implies no significant Gemini Flash developments until the next year.
randomblock1•1h ago
nateb2022•58m ago
re-thc•58m ago
Gemini 4 is apparently just around the corner so unless there's a 3 month delay... there's at least a new Flash update.
eis•30m ago