9 ms·
Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direc
by GodelNumbering 5mo ago
Per million input/output tokens:
Gemini 2.5 flash: $0.30/$2.50
Gemini 3.0 flash preview: $0.50/$3.00
Gemini 3.5 flash: $1.50/$9.00
Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview).
3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10
- dbbk 5mo agoI don't think they're really comparable. Seems they created the Flash-Lite tier to take the spot of the old Flash models.
- GodelNumbering 5mo agoNo, 2.5 had both flash and flash lite.
- mlmonkey 5mo agoIt is Google, after all ....
- rudedogg 5mo agoIf Google is actually getting cheaper inference than everyone else with their TPUs, this smells like trouble to me. Maybe serving LLMs at a profit is proving difficult. Or maybe they think because their benchmarks are good they can ramp up the prices. Seems like they don’t have the market share to justify a move like that yet to me.
- IncreasePosts 5mo agoMaybe the margins are just very large for Google because they predict so much demand for 3.5?
- GodelNumbering 5mo agoThis combined with locally runnable models getting pretty good recently (e.g. Qwen 3.6) tells me that it's time to seriously consider local dev setup again
- tempaccount420 5mo agoThis is not priced at inference cost. My guess: it's the price at which they make more money than if they rent the TPUs to other companies. The Gemini team has had trouble securing enough TPUs for their user's needs. They struggle with load and their rate limits are really bad. Maybe at a higher price, they have a better chance at getting more TPUs assigned?
- gpm 5mo agoThe cost at such they could rent out the TPUs, i.e. the market rate, is the inference cost. Just because you are vertically integrated doesn't mean you get to discount the one business units products to the other. Doing so discounts the opportunity cost you pay and is just bad accounting.
- HDThoreaun 5mo agoDepends on if you have spare capacity I think. They have minimal competition so they might be maximizing profit by charging prices higher than what clears all their supply.
- dash2 5mo agoLook up “double marginalisation”.
- KoolKat23 5mo agoBasic business principle, you charge what people are willing to pay not what it costs.
- sumedh 5mo ago> doesn't mean you get to discount the one business units products to the other That depends, if all developers get used to Claude and Codex it will become harder for Google to attract them in the future. They might lose devs in the long term.
- gpm 5mo agoPredatory pricing is a great business strategy and all (particularly when countering the competitors predatory pricing - what could go wrong), but that doesn't mean that the gemini-team should account for it as if they're getting the compute cheaper, it just means that they should run a loss.
- spyckie2 5mo agoIts probably that in 1 or 2 years local (free) models will completely take the place of cheap models so cheap models need to move up the quality chain. You have free local models for most tasks, $20 subscriptions for near-frontier intelligence, and API per token costs for frontier intelligence. Flash seems to be targeting the near-frontier category.
- TurdF3rguson 5mo agoThat might work if it wasn't for FOMO. Are you ok with only $20 of frontier usage a month?
- rohansood15 5mo agoSubjective, but if we compare to compute not everyone needs the most expensive laptops or super computers for their work. I think frontier models will be invaluable for scientific research, defense, financial analysis and such. But the average person probably would be reasonably well-served with a local model. If you're in sales, customer service, product management and such - the leading open models at the 30B mark are already good enough.
- TurdF3rguson 5mo agoI mean customer service maybe, but how much longer will humans even be doing that job at this point?
- booty 5mo agoPrevailing wisdom is that serving LLMs at a profit is achievable... it's when you factor in the cost of training them that prices get astronomical real fast. Open-source model inference providers (who do not have to bear the cost of training) seem able to do it at much lower prices. https://www.together.ai/pricing https://www.together.ai/pricing https://fireworks.ai/pricing#serverless-pricing https://fireworks.ai/pricing#serverless-pricing (scroll down to headline models) Of course, it's possible that they are burning through investor cash as well, and apples-to-apples comparisons are not possible because AFAIK Google does not mention the size/paramcount for 3.5 Flash. But if the prevailing wisdom is true, I think it's actually encouraging. It suggests that OpenAI and Anthropic could perhaps, if they need to, achieve profitability if they slow down model development and focus on tooling etc. instead. If true that's probably good news for everybody w.r.t. preventing a bursting of this economic bubble. ...my opinions here are of course, conjecture built on top of conjecture....
- HDBaseT 5mo agoNot to discredit you, because you are 100% correct but tangential note about together.ai, they seem fairly unreliable with constant outages or higher than normal latency.
- eklitzke 5mo agoMost of the training cost is not in the final training run, it's in all of the R&D (including salaries, equity, etc.) that it takes to get to the final training run. The actual cost of all of the TPUs (or GPUs), power, networking, storage, etc. for the final training run is significant, but it's even more expensive to have this huge R&D team doing frontier model development and using a lot of those same resources during development. I think you're right that releasing models at a slower cadence would bring down costs to some degree, but it's not clear how much. All of these companies could significantly reduce their opex but at the risk of falling behind in terms of being at the frontier.
- BoorishBears 5mo agoThis is trouble if you're not Google/OpenAI/Anthropic: they're all shifting towards pricing for the economic value of the knowledge work they're aiding. The economic value increases non-linearly as models get more intelligent: being 10% more capable unlocks way more than 10% in downstream value. That's trouble because the non-linear component means at some point their margins will stop primarily defined by the cost of compute, and start being dominated by how intelligent the model is. At that point you can expect compute prices to skyrocket and free capacity to plummet, so even if you have a model that's "good enough", you can't afford to deploy it at scale. (and in terms of timing, I think they're all well under the curve for pricing by economic value. Everyone is talking about Uber spending millions on tokens, but how much payroll did they pay while devs scrolled their phones and waited for CC to do their job?)
- tskj 5mo agoThank you, this is obviously where we're heading. People who think in terms of "will it ever be profitable to sell tokens" are thinking in the wrong framework entirely. The correct framework is "will it be profitable to sell knowledge work", and the answer will clearly be "yes".
- fnordsensei 5mo ago3.5 flash is listed as stable rather than preview, or am I misreading? https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flas...
- GodelNumbering 5mo agoah I mistakenly wrote preview
- dr_dshiv 5mo ago3.1 flash lite — $0.25/$1.50 — plus insanely fast. 3.1 flash lite isn’t quite as good as 3 flash preview (which is the most incredible cheap model… I really love it) — but 3.1 is half the price and the insane speed opens up different use cases. For comparison, Opus models are $5/$25
- SwellJoe 5mo agoOpus 4.7 is smarter than even Gemini 3.1 Pro on nearly every metric, though. You're comparing apples to oranges. Gemini 3.1 Flash is somewhere in the neighborhood between current Haiku and Sonnet, I think? Still a better value than the Anthropic models, I guess, which are quite pricey. Since Gemini 3.5 Flash is raising the price to $1.50/$9.00, it's priced between Haiku and Sonnet. If it outperforms Sonnet, it remains a good value, I guess. Though DeepSeek V4 Flash is much cheaper than all of them, and seemingly competitive.
- WarmWash 5mo ago>Opus 4.7 is smarter than even Gemini 3.1 Pro on nearly every metric, Outside of coding, claude models are pretty meh. GPT and Gemini are the workhorses of science/math/finance.
- robwwilliams 5mo agoNot in my fields of science: Genetics and neuroscience. The combination of Opus 4.7 Adaptive used with well structure project folders is amazingly useful.
- epolanski 5mo agoAnd even on coding, they are mostly good at generating new code. They sure are not at thorough analysis or debugging, etc.
- dr_dshiv 5mo agoDefinitely apples to oranges, sorry I wasn’t clear. I only included opus pricing for comparison—it is vastly superior. But even 3.1 flash lite is really useful. Of course, if I manage to reach my limits every week on my Claude $200 sub, opus 4.7 is probably priced closer to flash!
- doginasuit 5mo agoThey probably never intended to keep serving cheap models. This is a natural way to introduce the squeeze, now that they have people who built services on their API. It makes a lot of sense to have an abstraction layer where the provider doesn't matter. If you are working in Kotlin, Koog is excellent.
- hnarn 5mo ago> now that they have people who built services on their API People really can’t wait to be the next Zynga
- lanthissa 5mo agoswitching models is insanely cheap compared to token cost on anything signficant, this is a take so cynical it misses the reality
- Clueed 5mo agoin any corporate or half compliance-relevant setting switching isn't trivial. new DPA, subprocessor notifications, TIA, procurement review, security questionnaires, plus re-running your evals because prompts don't transfer 1:1. token cost is just one of the line items.
- deleted 5mo ago[deleted]
- lanthissa 5mo agono it really not, even the soggiest bank has multiple api vendors atm.
- alexandre_m 5mo agoI agree with parent. I'm not sure where your stance is coming from. From what I hear, most enterprise AI deployments are seat-based subscriptions with annual commitments.
- ilia-a 5mo agoYeah, it is a massive jump in price, hardly a "Flash" model anymore... I wonder if they'll release flash lite or something with a bit more affordable price point.
- OakNinja 5mo agoThere’s already a flash lite tier since 2.5. Latest is 3.1 currently.
- LetsGetTechnicl 5mo agoGen AI is unprofitable, especially at the insanely cheap rates they've been offering to get people in the door. So expect more increases in the future.
- GaggiX 5mo agoIf you don't need SOTA or near SOTA there are plenty of dirt cheap models, just look at Gemma 4 31B on Openrouter.
- ai_fry_ur_brain 5mo ago[flagged]
- Gigachad 5mo agoFor all of the use cases being hyped you really do, and you actually need something much better than the SOTA models to do what we are being told can be done. The small models are useful for small things like summarizing text or search but not much else.
- LetsGetTechnicl 5mo agoYeah a lot of AI hype is look at the amazing new thing our new model can do! Like Google at this event. But when pressed about its pricing reality the answer is “use a worse cheaper model”?? Real convincing argument there
- GaggiX 5mo agoIf don't want to spend 1.5$/9$ for the lastest model then yes use a cheaper model, DeepSeek V4 Flash is 0.11$/0.22$ on OpenRouter and it's more capable than the most expensive model a year ago. Models have never been so cheap given their capabilities unless you want to follow the SOTA (where the hype is).
- scrollop 5mo agoYou mean Kimi or qwen
- hei-lima 5mo agoWe need another "Deepseek moment" or else it will become impossible for the regular dude to use AI. It will become something that only big companies can afford.
- segmondy 5mo agoYou can use lots of open weight models today.
- hei-lima 5mo agoThat's one solution to the problem. But it still needs some good computational capabilities. Either we optimize the hell out of those models, or we wait for the hardware to become good enough for them.
- Gigachad 5mo agoThe real problem is the hardware to run them is still very expensive.
- squidbeak 5mo agoDeepseek had another moment a few weeks ago. V4 isn't far behind the US frontier, and so far its flash variant seems a very reliable coder and costs a pittance.
- ai_fry_ur_brain 5mo agoDeepseek V4 (not flash) trippled in price too by the way (from Deepseek). Get used to this pattern. This is what you get for relying on the generosity of billionaires. Keep offshoring your thinking ability to a machine and let me know how competitive you. Hint, you wont be. There's nothing special about being able to use an LLM.
- npn 5mo agoUnlike other providers, Deepseek does promise that they will lower the price when their Huawei cards arrive in a few more months.
- irthomasthomas 5mo agoAnd they are using this to power search answers?
- CooCooCaCha 5mo agoI bet the API pricing helps pay for search users
- photonair 5mo agoIn general, Gemini flash is still relatively cheaper compared to the "mini" version of the other big 2. However, I agree that newer version seem to have multiple X price increase (similar to the new ChatGPT) and we certainly need competition from the open source models to keep these guys in check with pricing.
- llm_nerd 5mo agoIt might be temporary pricing given that 3.5 Flash is actually superior to the existing 3.1 Pro in almost all regards, so they're in a bit of a lurch as 3.1 Pro really doesn't make sense given that 3.5 Pro has been delayed a bit.
- bjoli 5mo agoI let it loose on a f# codebase that I know was pretty optimized but with a few low hanging fruit changes that would have a big impact. 3.1 Pro did NOT find them. 3.5 flash did. Plus one I hadn't thought of that may or may not work (which it also pointed out). I'm pretty impressed.
- SwellJoe 5mo agoThat's a lot. DeepSeek v4 Flash is just over a tenth the price, and DeepSeek v4 Pro is roughly the same price (currently heavily discounted, but will be $1.74). I mean, the benchmarks for Gemini 3.5 Flash are very strong, but at those prices it has to be. I guess the time of subsidized tokens from the big guys is slowly coming to an end.
- copperx 5mo agoThey have said AI will be priced like a utility, meaning $100-300 per month or so.
- WhitneyLand 5mo agoTheir rationale might be that it’s size and intelligence are growing relative to the market. Fwiw it’s beating Claude Sonnet in most benchmarking (benchmaxxing?), yet they’ve priced it almost half off on a per token basis. Question is are you going to persuade anyone with this argument? Are there many devs at Google who legit prefer Gemini over Claude and Codex? Would love to hear about that.
- SyneRyder 5mo ago> Are there many devs at Google who legit prefer Gemini over Claude and Codex? Would love to hear about that. A few weeks ago, Steve Yegge claimed he'd heard that Google employees are banned from using Claude & Codex. https://x.com/Steve_Yegge/status/2046260541912707471 https://x.com/Steve_Yegge/status/2046260541912707471 A number of Googlers replied to say that was totally false, including Demis Hassabis, but they were all on the DeepMind team. https://x.com/demishassabis/status/2043867486320222333 https://x.com/demishassabis/status/2043867486320222333 This person here claims they left Google because of the ban, and because the ban applied outside of Google work as well: https://x.com/mihaimaruseac/status/2046272726881693960 https://x.com/mihaimaruseac/status/2046272726881693960
- myko 5mo ago> and because the ban applied outside of Google work as well I think false (or hasn't filtered to everyone lol)
- NitiX 5mo ago[flagged]
- m3kw9 5mo agojust subscribe to the plan, cheaper
- verdverm 5mo agoAt the same time, it is supposedly Gemini 3.1 Pro level at 3/4 the price and far cheaper than comparable models, Gemini Pro is cheaper than Claude Sonnet (Anthropic still gets to charge a brand premium)
- throwa356262 5mo agoGemini 2.5 flash was the best Gemini model. Not the most intelligent but perfect balance of cheap, fast and not-too-dumb.
- npn 5mo agoThe 09-2025 preview was awesome.
- __jl__ 5mo agoThis understates the cost increase. 3.5 Flash also uses more tokens. artificialanalysis.ai shows these difference to run the whole eval, which I think is more realistic pricing: Gemini 2.5 flash (27 score): $172 (1.0x) Gemini 2.5 pro (35 score): $649 (3.8x) Gemini 3.0 Flash (46 score): $278 (1.6x) Gemini 3.5 Flash (55 score): $1,552 (9.0x or 2.4x compared to 2.5 pro) This is a massive price increase... 5.6x compared to Gemini 3.0 Flash
- xdertz 5mo agothe era of subsidised ai is ending
- ashirviskas 5mo agoGemini 2.0 Flash: $19
- ahknight 5mo ago... and you get what you pay for. Or less.
- ahknight 5mo agoSonnet-level performance at Haiku prices. They know what they have and who the audience is they want.
- joshmlewis 5mo agoIt's interesting they use output tokens as an eval because all tokens are not made equal. Even from model to model (like Opus 4.6 to Opus 4.7) the tokenizer can be different and it's no longer an apples to apples comparison. No one really talks about this but it directly affects stats like usage limits. Certainly comparing models between providers on an apples to apples comparison token wise is not a good test.
- OakNinja 5mo agoTo be fair, Gemini 3.1 flash _lite_ supports structured output (guaranteed json), it’s super fast, runs circles around 2.5 flash and costs $0.25/$1.50. I use it _a lot_ and it’s very capable if you just plan correctly. I actually almost exclusively use 3.1 flash lite and 2.5 flash lite (even cheaper) and we have 99.5% accuracy in what we do. That said, I think we’ll see the lite/flash models and the pro models will diverge more price wise. The pro models will become more and more expensive.
- drob518 5mo agoI think that’s true on divergence. Basically, the only most is living in the frontier, and even that is only temporary. At some point, the frontier advances such that 99% of tasks can use something short of a frontier model and only a very few tasks actually demand frontier performance.
- dzhiurgis 5mo agoI use Gemini models in Junie daily. When I need accuracy I switch to Gemini 3.1 Pro Preview (why it is still in preview?), but it burns thru credits leaving me topping up $5 every day. 3.1 Flash lite is just not accurate enough. 3 Flash is sweet spot just as Jetbrains suggests it is. Maybe I'll look at Opus again, but it just was slower, much more expensive and worst at all - wasn't listening to you instructions.
- malloryerik 5mo agoTo me this is almost like a tone-deaf naming change. Empty Slot (new Pro as Mythos competitor?) Old Pro -> now Flash Old Flash -> now Flash Lite Old Flash Lite -> now Gemma (and not served by Google) I say "almost" because the situation is more fluid and unstable than a normal naming change. If Apple were to do this with laptops, maybe it'd be like, Air gets better and pricier and becomes Pro-level model, Neo same way becomes Air-level model, etc. But Apple's too design oriented to do something like that. Google, well... This change has made me decide to move to a multi-provider situation like through OpenRouter for consumer-facing LLM api in a service I'm building. I just can't trust Google to not constantly rearrange everything under our feet. Doesn't mean I won't use Gemini, but it clearly means I need to have others in the mix ready to go. In fact I used to use lots of Flash Lite, which is now Gemma territory, and I can't get that served by Google anymore and don't want to run my own hardware. But in any case, I'd compare this "Flash" model with previous "Pro" on all metrics. It's kinda like if in clothes a Small suddenly became what was a Large, or at Starbucks a Grande became the new de facto Venti. And only for the new! drinks. And if we think this way, it's possible that prices are actually falling?
- baq 5mo agoDemis is on record saying they need small models on edge devices and if it’s on the edge the weights may as well be public officially.
- deaux 5mo ago> Old Flash Lite -> now Gemma (and not served by Google) > which is now Gemma territory, and I can't get that served by Google anymore Gemma is served by Google. They're serving Gemma 4 26B A4B at $0.15/$0.60. https://console.cloud.google.com/agent-platform/publishers/google/model-garden/gemma-4-26b-a4b-it-maas https://console.cloud.google.com/agent-platform/publishers/g... https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing#gemma-model https://cloud.google.com/gemini-enterprise-agent-platform/ge...
- malloryerik 5mo agoAh, thanks!
- davedx 5mo agoI use Gemini for heavy web scraping-adjacent API work. Web grounding has been super useful for the project. I will definitely not be updating to this new model, and I think once 2.5 Flash is deprecated I'll have to re-architect so Gemini is only used for web grounding requests. This is an insane price increase.
- harrouet 5mo agoIf you look at the benchmark, the model is not particularly good at coding, and as you point out it costs 3x the price of the previous flash models. So what is the market for it? I think that they might have reached the latency sweetspot where voice applications become more natural. Natural speech is <100 tokens per second (after STT), so $9 for a million token takes you to roughly 3 hours of speech. That's totally competitive compared to human costs.
- ashirviskas 5mo agodon't forget Gemini 2.0 flash at $0.10/$0.40
- jstummbillig 5mo ago> Interesting pricing direction. Is it? More capability, more demand, higher price. Seems relatively uninteresting. The naming structure complicates it: 3.5 Flash is less comparable to 3.0 Flash than it is to 3.0 Pro. More generally, $/token + naming scheme comparisons are just confusing: I am not looking for a wordy idiot and I doubt most people are (at least not with what I would consider worthwhile business ambitions). In fact wordy idiots are fairly costly, because we have to consider the large amounts of cheap garbage that they are producing, and if you price your own time somewhat competitively then fairly quickly that's the bigger lever. Even if we don't consider the last part: How do we price the better model, that can one shot a task without having to go back and forth and spending more tokens or having to fix more bugs later? It is definitely worth something and I think it's quite undervalued right now. What seems to be missing is a better measurement of capability per token. I don't know how that could look like. Maybe something like how we try and measure inflation, some basket of tasks (which then ends up being part of the training data so idk).