16 ms·
Gemini 3.5 Flash
https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flas...
- brikym 5mo agoHow is this progress? The token cost just keeps going up and up. Flash is the new Pro? Do the models actually cost more to run or is it fattening margins?
- deleted 5mo ago[deleted]
- f311a 5mo ago$9/1M output
- explosion-s 5mo agoI wonder if this is because it's a larger model or maybe just because they can? Although with the latest Deepseek it's really tough to compete pricing wise. Inference speed and integration (e.g. Antigravity) might be their only hope here
- hydra-f 5mo agoIt has to be a larger model, wouldn't make much sense otherwise. That isn't to say the price isn't artificially increased as well The Antigravity harness is really well done, so I do agree it's their strong suit. Can't say the same about gemini-cli (though it has a really nice interface) Would still choose Deepseek for the price
- alexdns 5mo agoIts Gemini 3.5 Flash
- nerdalytics 5mo agoYeah, Google chose a misleading title for the blog post.
- jader201 5mo ago> Today, we’re introducing Gemini 3.5, our latest family of models combining frontier intelligence with action. This represents a major leap forward in building more capable, intelligent agents. We’re kicking off the series by releasing 3.5 Flash.
- nerdalytics 5mo agoparagraph vs title
- swe_dima 5mo agoFlash family but costs like a Pro. $9 vs $12 for output.
- deleted 5mo ago[deleted]
- asar 5mo ago$1.5/m input tokens $9/m output tokens 6x the price of 3.1 flash lite
- himata4113 5mo agoI don't think input/output pricing matters, 90% of the cost is cache. $0.15 is pretty good, but still very expensive.
- minimaxir 5mo ago10% of input pricing is standard especially compared to competition.
- himata4113 5mo agoyah, which means that the input cost is the only value that should be paid attention to at the end + the cache discount (x10). If google would start offering x20 discount it would make it twice as cheap while input and output stayed the same.
- wolttam 5mo agoIt depends on the use-case. yes, 90% of cost is cache in agentic coding scenarios (actually 95% in my experience). But not when the model reasons for 200k+ tokens before answering a complex problem.
- himata4113 5mo agogemini models solve a problem in 80% less tokens so that's something to think about.
- johaugum 5mo agoSource?
- 5mo ago
- himata4113 5mo agoEngineers at google have publically stated that the models are too big and are far from their potencial. Glad they're being proven right with every release. They continue to focus on smaller models while openai and anthropic are increasing compute requirements for their SOTA models.
- stri8ed 5mo agoGiven the cost increase associated with this model, and previous model releases, I think the size is trending upwards, not down.
- himata4113 5mo agoThe speed says otherwise. I think they're increasing costs since they want to start seeing ROI.
- JanSt 5mo agoThose are (mostly) new, faster TPU
- himata4113 5mo agolatest TPU's appear to reach 800tok/s rather than the advertised 300tok/s.
- mgambati 5mo agoThey demoed today 8i running ate 1300 to 1600ish tokens per second. I imagine that is caused by having a single rack serving the model just for the demo.
- himata4113 5mo agoThere's a limit to how much you can "scale" this process, it's linear, but if we did napkin math based on vllm parallel batched streams only lose around ~50% performance compared to single-stream output so doesn't explain the ridicioulusly fast numbers here. I wish google just came out and told us how large their flash model is, because if it's as big or smaller than gpt-5.4-nano that's the real headline here.
- deleted 5mo ago[deleted]
- golfer 5mo agoHere's the benchmark scoreboard they published: https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-3-5__benchmarks__light.gif https://storage.googleapis.com/gweb-uniblog-publish-prod/ori...
- mugivarra69 5mo ago[dead]
- mixtureoftakes 5mo agobenchmarks look REALLY good, the price hike is big but it also beats sonnet 4.6 in every discipline?
- benjiro3000 5mo ago[dead]
- deleted 5mo ago[deleted]
- SXX 5mo ago> Create animated SVG of a frog on a boat rowing through jungle river. Single page self contained HTML page with SVG 3.5 Flash: Thinking Medium - 7516 tokens https://gistpreview.github.io/?5c9858fd2057e678b55d563d9bff0b2e https://gistpreview.github.io/?5c9858fd2057e678b55d563d9bff0... 3.5 Flash: Thinking High - 7280 tokens https://gistpreview.github.io/?1cab3d70064349d08cf5952cdc1658ca https://gistpreview.github.io/?1cab3d70064349d08cf5952cdc165... 3.1 Pro - 28,258 tokens https://gistpreview.github.io/?6bf3da2f80487608b9525bce530183af https://gistpreview.github.io/?6bf3da2f80487608b9525bce53018... Though 3.1 took 3 minutes of thinking to generate, but it only one that got animated movement.
- abi 5mo agoYour links are broken FYI.
- John7878781 5mo agoThey work for me.
- TacticalCoder 5mo agoThey do work here too.
- captn3m0 5mo agoAll three links animate for me.
- NitpickLawyer 5mo agoI think they mean the boat is moving. In the flash ones the paddles are animated but the boat is stationary for me.
- codazoda 5mo agoThe boat moves in all three for me
- cesarvarela 5mo agoAdd Flash to the title, please.
- meetpateltech 5mo agoedited it.
- benbencodes 5mo agoPricing is now live on ai.google.dev/pricing: Gemini 3.5 Flash: $0.75 input / $4.50 output per 1M tokens, 1M context window. The output price explicitly "includes thinking tokens" — which is why it's higher than a typical flash-class model. For comparison within the Gemini lineup: - Gemini 2.5 Flash: $0.30 / $2.50 - Gemini 3.1 Flash-Lite: $0.25 / $1.50 - Gemini 3.1 Pro Preview: $2.00 / $12.00 So 3.5 Flash is ~2.5x more expensive input vs 2.5 Flash. The pricing and "including thinking tokens" framing position it as a reasoning-capable flash model rather than just a pure speed optimization.
- conorh 5mo agoI think you have your pricing wrong there, Gemini 3.5 flash is $1.50 input and $9 output.
- mchusma 5mo agoOkay, it's kind of somewhere between haiku and sonnet level pricing, at somewhere between sonnet and opus level performance. Its a great option. I was hoping to see opus class intelligence at haiku level pricing out of google, and this is close to that!
- mchusma 5mo agoNever mind, after looking at more benchmarks, seems closer to sonnet level intelligence at slightly lower cost. Speed is great for latency sensitive applications, but if this was 1/2 the cost it would have been priced to win. If this is the big model release out of google, its a disappointent.
- jpau 5mo agoStandard pricing is showing for me as $1.50 / $9. (I suspect you're viewing the "flex" pricing).
- lyjackal 5mo agoYou’re quoting the batch pricing. On demand is 1.5 per input and 9 per M output. This is effectively comparable cost to Gemini 2.5 Pro in a flash tier model
- aliljet 5mo agoIs there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
- Sevii 5mo agoI haven't been bothered by hallucinations in premier models since early last year. Still see it in smaller local models though.
- aliljet 5mo agoI'm really running into this deep at the edges of content creation. Take, for example, a need to general some kind of legal work. The cost of painstakingly checking and rechecking each case cited is reducing the value of these frontier models immensely. Coding, however, is solved like magic. Easier to add tests, to be fair.
- throawayonthe 5mo agowell there is https://artificialanalysis.ai/evaluations/omniscience https://artificialanalysis.ai/evaluations/omniscience
- goldenarm 5mo agoIt's a gibberish input detection benchmark, and does not measure output hallucinations.
- yieldcrv 5mo agoif last year's models were the ones people got familiar with in late 2022, hallucinations would be an underrepresented rumor, there would be no articles about it because its so rare. overconfident lawyers wouldn't have messed up dockets in court with fake case law, in other domains that move faster, sources would be only partially outdated with agentic search and mcp servers filling in the gaps AI psychosis would be the problem people talk about more, not just outright agreement but subtle ways of making you feel confident in your ideas. "yes, buy that domain name buy these other ones for defensibility" (the domain name is dumb and completely unmarketable)
- bakugo 5mo agoTriple the price of the last Flash model ($3 -> $9 per 1M output). Quickly approaching Sonnet prices. Feels like the AI pricing noose is tightening sooner rather than later.
- eis 5mo ago3.5 Flash was more expensive than 3.1 Pro to run the Artifical Analysis test suite. $1551 for 3.5 Flash [0] vs $892 for 3.1 Pro [1]. That's 74% more cost while ranking lower. It's 2.5x as fast but I don't think the bang for the buck is there anymore like it was with 3.0 Flash. I'm a bit bummed out to be honest. I did not expect such a huge (3x) price increase from 3.0 Flash and I bet many people will not just blindly upgrade as the value proposition is widely different. One interesting point to note is that Google marked the model as Stable in contrast to nearly everything else being perpetually set as Preview. [0] https://artificialanalysis.ai/models/gemini-3-5-flash https://artificialanalysis.ai/models/gemini-3-5-flash [1] https://artificialanalysis.ai/models/gemini-3-1-pro-preview https://artificialanalysis.ai/models/gemini-3-1-pro-preview
- ls_stats 5mo ago>3.5 Flash was more expensive than 3.1 Pro to run the Artifical Analysis test suite That's everything I needed to know.
- ekojs 5mo agoSeems like the only good thing about 3.5 Flash is its speed. Not cost-competitive or benchmark-leading by any means.
- mijoharas 5mo agoThat's what I came here to check. Last model release they only put it into preview[0] at first. Does that mean this model is production ready? [0] https://news.ycombinator.com/item?id=47076484 https://news.ycombinator.com/item?id=47076484
- pingou 5mo agoHow do they calculate that? 3.1 has 57M output tokens from Intelligence Index, 3.5 Flash has 73M, so not a lot more, and 3.5 is a bit cheaper, I don't get how 3.5 can be 74% more expensive.
- knollimar 5mo agoOnly speculation but cache maybe?
- nightski 5mo agoAI being a product is not the future. It's more like an operating system that deserves to be open and free (aka Linux). Unless that happens we are in for a very dystopian future. I wish I had the intelligence, resources and/or connections to try and make that happen.
- lugu 5mo agoWhat we need today is a standard local API (think of it as a POSIX extension). So that each desktop app that needs AI to enhance a feature can simply call that. This way, those apps will need to handle the case where AI is not availabile. This will empower users.
- charcircuit 5mo agoAll major operating systems Windows, macOS, iOS, and Android have local APIs for using AI.
- hedora 5mo agoWhy would I use those instead of just grabbing a model from hugging face? Are they as good as qwen 30B?
- charcircuit 5mo agoBecause it is simpler as an application developer to just use an OS API then trying to figure out some 3rd party thing and setting that up. Each platform has several different models for different things so I can't give a comparison.
- HardCodedBias 5mo agoOh boy. GDM is making (or has been backed into a corner into making) the bet that high throughput, low latency, low capability models are the path forward. That probably works for vibe coded apps by non-practitioners. I suspect that practitioners/professionals will wait longer for better results.
- brokencode 5mo agoWhere do you see that it’s low capability? And Google is trying to make something affordable enough for a mass market, ad-supported audience. They aren’t hyper focused on enterprise like Anthropic is. And that’s okay. There’s room for different players in different markets.
- hedora 5mo agoPrice up (cost up?), benchmarks down. Latency down. So, who is this for? People that want more ads and worse output, but want it faster? Sounds pretty awful to me.
- OsrsNeedsf2P 5mo agoBeats 3.1 Pro for price per token, but artificial analysis is showing it's dumber per token and costs more overall
- sauwan 5mo agoYeah, bummer. I was very excited for this release, but this killed it.
- droidjj 5mo agoThe pricing is an issue.
- golfer 5mo agoArena.ai is saying "Gemini 3.5 Flash’s pricing shifts the Pareto frontier in Text. 8 models from GoogleDeepMind dominate the Text Arena Pareto curve where only 4 labs are represented for top performance in their price tiers." https://x.com/arena/status/2056793180998361233 https://x.com/arena/status/2056793180998361233
- nicce 5mo agoNot sure what to think about this. There is no even GPT 5.5
- s3p 5mo agoYikes. I think the concept of a 'flash' model is changing, no? Google used to market this as its lower-intelligence, faster, cheaper option. I appreciate that they are delivering on both of those, but personally I would appreciate if they could create an incremental knowledge improvement while holding price steady. Fortune 500 companies have to make their money I guess.
- 2001zhaozhao 5mo agoI think flash just means "fast" now
- likium 5mo agoMy guess is Gemini Pro coming later will be 2x more, bringing it comparable to Opus’s pricing.
- deleted 5mo ago[deleted]
- toraway 5mo agoThat would be Flash Lite now, and I'm also interested in the cheaper end of things so kinda disappointed they didn't release 3.5 Flash Lite at the same time...
- kilpikaarna 5mo agoReal smart. I’ve come to associate ”Flash” with ”useless make-shit-up”, and always look for Thinking/Pro when I see it set. Now, suddenly, there is only Flash?
- noelsusman 5mo agoThe Artificial Analysis benchmark results are pretty underwhelming. Roughly the same "intelligence" as MiMo-V2.5-Pro for over 3x the cost. We'll have to see how that translates to actual usage but it's not a great sign.
- hydra-f 5mo agoThat really depends on whether they have similar parameter counts, doesn't it? Unless you know that, the comparison is just strange
- halJordan 5mo agoBad look to tell people they're not allowed to compare things just because we need to respect Google's privacy
- hydra-f 5mo agoI didn't take the price into consideration when writing that. I meant to point out that even if they have similar scores, the Flash model might be smaller than MiMo or Kimi, which would by itself be a win That said, haste makes waste as the price point completely invalidates that
- deleted 5mo ago[deleted]
- noelsusman 5mo agoI don't know why a user should care at all about parameter counts. All that matters is performance and cost.
- merb 5mo agoStil no new processor version for document ai https://docs.cloud.google.com/document-ai/docs/release-notes https://docs.cloud.google.com/document-ai/docs/release-notes that is so weird. (Customer extractor) It’s not possible to uptrain on preview releases and it did not get that much love for a while.
- warthog 5mo agoGPT-5.5 on the benchmarks still seem to perform better than this Plus the vibe of the gemini models are so weird particularly when it comes to tool calling At this point I kinda need them to shock me to make the switch
- simianwords 5mo agoNo one talking about how this flash Beats Pro? Imagine what 3.5 pro looks like? Also concerned about Gemini models being benchmaxxed generally
- NitpickLawyer 5mo ago> concerned about Gemini models being benchmaxxed generally I would say they are the least benchmaxxed out of all the top labs, for coding. They've always been behind opus/gpt-xhigh for agentic stuff (mostly because of poor tool use), but in raw coding tasks and ability to take a paper/blog/idea and implement it, they've been punching above their benchmarks ever since 2.5. I would still take 2.5 over all the "chinese model beats opus" if I could run that locally, tbh.
- computerex 5mo agoI have never had good experience with any Google models in coding. Particularly for coding hard stuff, there is a night and day difference between Opus/Gemini in my experience.
- hubraumhugo 5mo agoJust updated my HN Wrapped project with it and it does well on my totally unscientific LLM humor benchmark: https://hn-wrapped.kadoa.com https://hn-wrapped.kadoa.com
- npn 5mo agoThe price is crazy. And I guess Gemini 3.5 pro will have the pricing increment, too. 12 x 5 = 60? It seems like google does want us to use Chinese models.
- brianwawok 5mo agoWhat exactly are you doing with this that you can’t generate $1.50 of value per million tokens?
- bel8 5mo agoGenerate 5x more value for the same amount of money.
- s3p 5mo agoWrong question. Right question: What exactly is Google's plan for the long term pricing of these models, and are we all going to be priced out in a year?
- npn 5mo agoI sell service. Imagine my users have to pay 4x more for marginal increment just 'cause. They are more willing to wait though, so Chinese models are pretty attractive right now.
- GodelNumbering 5mo agoPer million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview). 3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10
- dbbk 5mo agoI don't think they're really comparable. Seems they created the Flash-Lite tier to take the spot of the old Flash models.
- GodelNumbering 5mo agoNo, 2.5 had both flash and flash lite.
- mlmonkey 5mo agoIt is Google, after all ....
- rudedogg 5mo agoIf Google is actually getting cheaper inference than everyone else with their TPUs, this smells like trouble to me. Maybe serving LLMs at a profit is proving difficult. Or maybe they think because their benchmarks are good they can ramp up the prices. Seems like they don’t have the market share to justify a move like that yet to me.
- IncreasePosts 5mo agoMaybe the margins are just very large for Google because they predict so much demand for 3.5?
- GodelNumbering 5mo agoThis combined with locally runnable models getting pretty good recently (e.g. Qwen 3.6) tells me that it's time to seriously consider local dev setup again
- llmslave 5mo agoConspiracy theory: This model isnt an advancement, its a previous model that runs more compute, which is why it costs more
- npn 5mo agoNah, it costs what you are willing to pay.
- golfer 5mo agoArena.ai: > Gemini 3.5 Flash’s pricing shifts the Pareto frontier in Text. 8 models from GoogleDeepMind dominate the Text Arena Pareto curve where only 4 labs are represented for top performance in their price tiers. https://x.com/arena/status/2056793180998361233 https://x.com/arena/status/2056793180998361233
- h14h 5mo agoGiven how widely varying the amount of tokens each model uses for a given task, "price-per-token" is essentially meaningless when doing this sort of comparison. Artificial Analysis's "Cost to run" model (aka num_tokens_used * price_per_token) is much better, but even that is likely problematic since it's not clear whether running a bunch of benchmarks maps cleanly to real-world token use.
- ohlookcake 5mo agoThat graph seems odd. It looks like Gemini 3.5 Flash is not actually on the convex hull, and they forced the 'frontier' to bend inwards to include it
- andrewstuart 5mo agoThe benchmark that matters - can it actually program as well as Claude code. If not then I’m not using it. Cancelled my account 3 months ago, only Claude code level capability would bring me back.
- cmrdporcupine 5mo agoI spent 10 minutes with it in their new "agy" CLI tool and immediately found it is nowhere close to GPT 5.5 high in codex. It was sloppy and made poor assumptions in its analysis. It would have produced a mess if I let it go ahead with its plan. And it was just like previous versions of Gemini with poor tool use (e.g. "I wrote a file with the plan..." but file was never written.) For reference, this is a Rust codebase, deep "systems" stuff (database, compiler, virtual machine / language runtime) They're still months behind OpenAI and Anthropic on coding. Mind you I also find Claude Code careless and unreliable these days, too. (But it's good at tool use at least). I do use Gemini for "lifestyle" AI usage (web research etc) tho.
- reconnecting 5mo agoKnowledge cutoff: January 2025 Latest update: May 2026 I have a very bad feeling about this lag.
- hosel 5mo agoCan you explain what you mean?
- nemomarx 5mo agoIt might indicate core model training and pre training is really slowing down?
- mixtureoftakes 5mo agoalso parsing is harder + so much more of the new data is being generated by ai itself. still the cutoff is very much concerning and inconvenient
- reconnecting 5mo agoLLM pre-training models risk being unable to be updated with data from after 2025, as much of it is corrupted with LLM-generated content. We might be locked into outdated knowledge, where only whitelisted sources decide what to include. Taking into account the sometimes blind belief that 'LLMs know everything', the outcome could be very costly, especially for technologies and businesses unfortunate enough to emerge after 2025.
- neksn 5mo agoConsidering all models can use search engines, is this really relevant?
- reconnecting 5mo agoUntil they prefer not to search. Let me explain using the example of the open-source security framework (1) our team is working on. If you ask Gemini what you should use to integrate fraud prevention or account takeover protection into your product, there will be no mention of our open-source project. Five years in development, 1.3k stars, over 140 pull requests — all this isn't enough to make it into the training data. From this perspective, any technology that emerges after 2024 is simply invisible to LLMs. The answer is: without being in the training data, LLMs basically don't understand what they're searching for. 1. https://github.com/tirrenotechnologies/tirreno https://github.com/tirrenotechnologies/tirreno
- stan_kirdey 5mo agoEXPENSIVE ._.
- MASNeo 5mo agoWell, available for Gemini means these days that half the time they are “Receiving a lot of requests right now.” and so sorry they couldn’t complete the task. Luckily the model supports long time horizons because that’s what’s needed. /me likes Gemini a lot just wishing Google would add the compute!
- simonw 5mo agoThe pelican is a lot: https://github.com/simonw/llm-gemini/issues/133#issuecomment-4491300245 https://github.com/simonw/llm-gemini/issues/133#issuecomment... Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars. Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-flash https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...
- hedgehog 5mo agoThat pelican looks like it's in Miami for a crypto conference.
- xattt 5mo agoIt looks like it’s been partying for 60 years based on the wrinkles on its pouch.
- ethbr1 5mo agoYou don't know what that pelican has been through.
- joseda-hg 5mo agoIt looks like the starting soon screen of a crypto presentation
- egillie 5mo agoand somehow in 1992
- Xenoamorphous 5mo agoPelican in a white Testarossa.
- verdverm 5mo agosorta looks like the Tron ripoff in the I/O keynote
- ralusek 5mo agoThose prices, what a disappointment.
- mackross 5mo agoThe antigravity teamwork-preview doesn't work for me -- upgraded to ultra, installed antigravity 2, ran teamwork-preview, keeps failing: "You have exhausted your capacity on this model. Your quota will reset after 0s."
- jdw64 5mo agoHonestly, I feel like the new Gemini 3.5 Flash is a failure. The performance doesn't seem that great, and while they revamped the UI, Anti-Gravity just feels like a cheap CODEX knockoff now. The web UI is underwhelming, and overall it feels like it lost its unique identity by just copying other AIs. It’s a flop in both performance and price point. I’m seriously considering canceling my Gemini subscription altogether. Using Chinese AI models might actually be a better option at this point
- lanewinfield 5mo agoGemini 3.5 Flash's 2000 token clocks aren't bad. https://clocks.brianmoore.com/ https://clocks.brianmoore.com/
- casey2 5mo agoI think the field moved to agents too fast. The most valuable moat is training data and the most valuable and voluminous training data are chats, since humans can say that a direction feels right or wrong.
- OhMeadhbh 5mo agoAm I really so old that when someone says "Flash" my immediate response is... "consider HTML5 instead" ??
- hedora 5mo agoI guess I'm slightly younger: I think "weights or it didn't happen"!
- nightski 5mo agoVery little of what made the Flash culture so fun made its way into HTML5.
- sieabahlpark 5mo ago[dead]
- CobrastanJorji 5mo agoI dunno, the tools are kind of there. Browsers have canvases and JavaScript and SVGs and sound. The communities are around; they're just kind of dispersed. There's no one website that is THE place for fun stuff. Instead, there are dozens, and most of them suck. There's still fun stuff, though. I stumbled upon this bit of insanity just yesterday: https://tykenn.itch.io/trees-hate-you https://tykenn.itch.io/trees-hate-you. It would have fit in fabulously with the old Flash sites.
- moritzwarhier 5mo agoEdit: looks like you linkes something created with Unity? Not sure, I'm not versed in game dev. So maybe my point about creation tools is moot. However, 3D content always seems very samey to me, in a way that cartoons and regular animation don't. So the rest of my comment should still express what I mean. --- Flash had a WYSIWYG editor aimed at media creators who treat programming at best as an afterthought. Flash was mostly about ease of tweening and extremely flexible vector graphics engine combined with an intuitive creation tool. So the "Flash vs HTML/JS/SVG/CSS..." debate is not just about technical capabilities of the medium. Of course there are many fun web apps in the browser, or as native apps, too. But Flash attracted all kinds of slightly nerdy people with cultural things to say, not just web devs with a lot of free time. What "HTML5"/browser web technology doesn't offer is this intuitive, visual creation pipeline, and this kind of speaks for itself! Also, I think the Flash "creator's" age is not separable from its time: using Flash wasn't trivial either. There were just more people with interesting ideas, free time, and a wholistic talent for expressing their humor and ideas, combined with the curiosity and skill to learn using Flash (of course only as a licensed copy purchased from Macromedia). People like this today are probably more often hyper-optimizing social media creators, and/or not terminally online. In other words: I don't think the typical Newgrounds creator would have taken the time and effort to translate a stickman collage, meme, or other idea into a web app / animation. --- And to add even more preaching: I think that "creating" things using AI produces exactly the opposite effect: feed it an original idea, and the result will be a regression to the mean.
- wg0 5mo ago3x price increase for a similar model almost. And they said AI would be cheaper and ubiquitous.
- alexandre_m 5mo agoUbiquitous like the crack epidemic.
- verdverm 5mo agoor 3/4 the price (of 3.1 Pro) if we believe their benchmarks
- baalimago 5mo agoWhat happened to gemini 3.2, 3.3, and 3.4..?
- x3cca 5mo agoI'm excited for the conversation to switch from intelligence to tps instead. I care much less about what hard thought experiments models can one shot and much more how responsive my plain text interface for doing things is.
- ai_fry_ur_brain 5mo agoImagine reducing yourself to the worst of averages by making your competency 1:1 correlated to the tokens that you have access too (and everyone else does).
- cloakandswagger 5mo ago> correlated to the tokens that you have access too (and everyone else does) Do you mean "the weight parameters you have access to[sic]" or do you frequently find yourself limited by the model's token vocabulary?
- deleted 5mo ago[deleted]
- paperwork360 5mo agoGoogle also updated Antigravity. version 2.0 is more for conversation with agent. The previous VS Code like IDE was much better.
- xnx 5mo agoThey still have an Antigravity IDE version.
- operatingthetan 5mo agoIt's been renamed to "antigravity IDE." Updating my old IDE got me the new non-IDE app though, which is strange.
- bredren 5mo agoCan anyone who has extensive, recent, experience with Claude code and Codex contextualize the current Gemini CLI product experience?
- SwellJoe 5mo agoI have and use both Claude Code and Gemini CLI, and still don't consider Gemini worth starting for coding except to review Claude's output in critical commits (on a security boundary, maybe broad refactors, etc.), though I try side-by-side every now and then just to see the state of things. I also use Gemini Pro in a security scanning harness to act as a second set of eyes, but Opus is better at finding security bugs than Gemini, so I don't know that it's accomplishing anything beyond just using Opus. Gemini Pro 3.1 for agentic coding is still clumsy. It chews a lot, has a harder time with tools and interacting with the codebase. I haven't tried any 3.5 version, yet, though. The benchmarks look promising. I'll note I like the Google models' prose better than any others at the moment, though. Even the small open models (Gemma 4 family) have excellent prose, relatively speaking, that doesn't stink of the LLMisms that I find so annoying about OpenAI (especially) and Anthropic models. So, I'll probably start using Gemini for writing API docs, even if all code is Claude.
- nicce 5mo agoI would argue that prose is just a prompt issue. GPT 5.5 outout is easier to control whan Gemini by prompting. Having better defaults does not make it necessarily better.
- SwellJoe 5mo agoI would disagree. I think it'd take a lot of prompting to make GPT 5.5 not have the underlying personality of GPT, which I find awful. They have knobs in ChatGPT to choose a "professional" tone, which improves it somewhat, but even that is still the worst prose of any leading model. My default AGENTS.md/CLAUDE.md/etc. is a few sentences from Strunk and White, to try to make all the models not suck at writing. It helps keep the models brief, but it doesn't actually make models with shitty prose have good prose. The relevant portion of my agents file is: "Omit needless words. Vigorous writing is concise. A sentence should contain no unnecessary words, a paragraph no unnecessary sentences, for the same reason that a drawing should have no unnecessary lines and a machine no unnecessary parts." Which might add up roughly the same as "be brief" in the weights, I don't know. If you have a prompt that makes GPT a decent-to-good writer, I would like to see it. Gemini produces decent-to-good prose without prompting, which improves if instructed to be concise. The other models, even the frontier models, do not have decent-to-good prose without prompting, and even with prompting, rarely elevate to what I would consider Good Enough. Part of this may be that GPT and Claude models get used a lot more heavily, and so I'm highly tuned into their idiosyncrasies. The heavy use of emojis, the click-bait headline style, etc. that they both use unprompted. All of that is repugnant to me, so anything that doesn't do all that by default, or at least not as aggressively, has a huge leg up.
- owentbrown 5mo agoHas anyone switched from Claude 4.7 Opus or ChatGPT 5.5 to this? How does it feel? Dumber? Worth it for the speed? I'd love someone's subjective take on it, after doing a long session of coding. Reiner Pope gave a talk on Dwarkesh Patel about token economics. I guess faster is a lot more expensive, generally. Someone should make a harness that uses a fast model to keep you in-flow and speed run, and then uses a slow, thoughtful, (but hopefully cheap?) model to async check the work of the faster model. Maybe even talk directly to the faster model? Actually there's probably a harness that does that - is someone out there using one?
- pcwelder 5mo agoOpus is not the correct tier to compare this flash model with. On my tasks it has not been as good as even Sonnet 4.6 so far. Instruction following over long context feels worse. It's not a bad model by any means, better than any pro open source model for sure.
- landtuna 5mo agoI was using GPT 5.5 for a bunch of work this morning. It's brilliant and efficient. I was also using GPT 5.4 mini. It gets the job done and works great for subtasks that 5.5 designs. Gemini 3.5 Flash is SUCH a Gemini. It seems to work okay, but its attitude is disgusting. "Yes, your idea is excellent." "How this works beautifully:" "This is a fantastic development!" "This is an exceptionally clean and robust architecture." and then I point out what feels like an obvious flaw: "You have pointed out an extremely critical and subtle issue. You are absolutely 100% correct." I'm sad that I'll probably stop using 3.5 Flash because I just hate its personality.
- andriy_koval 5mo agoI added something: be grumpy cynical software engineer with strong rigor, and it fixed personality.
- kaspermarstal 5mo agoI switched from Opus 4.6 -> Opus 4.7 -> GPT 5.5 and tried Flash 3.5 tonight and I was not impressed. It is straight up unreliable, e.g. deleting code and forgetting to add the new stuff it was asked to, then happily marking the task as complete with up-beat conclusion. I personally appreciate GPT 5.5 toned-down, objective style so really dislike how this model feels. I get that it's a flash model and not in the same league as GPT 5.5 but their marketing suggest otherwise so thy are just setting themselves up for disappointment.
- kristopolous 5mo agoI have a tool to track these I've built Relatively speaking here's where it's at: score age size name 44.2 97 large GLM-5 (Reasoning) 44.7 187 - GPT-5.1 (high) 44.9 29 - Qwen3.6 Max Preview 45 0 - Gemini 3.5 Flash 45.5 27 large MiMo-V2.5-Pro 45.6 75 - GPT-5.4 (low) this is from artificial-analysis using https://github.com/day50-dev/aa-eval-email/blob/main/art-analysis.sh https://github.com/day50-dev/aa-eval-email/blob/main/art-ana... I really don't know why people down vote me. What do I need to say to make things for free that people like? Sincere question. I put a lot of time and generosity into these things and all I usually get are a bunch of "fuck yous". This is honestly an existential issue for me. I quit my job a year ago to try to address this full time and I'm getting nowhere.
- esafak 5mo agoI see no 'score' or 'age' mentioned in your script. What does age signify and how are they calculated?
- kristopolous 5mo agoThis isn't obvious? "\( 10 \* (.codingIndex // 0) | round / 10 ) \( ( now - ( .releaseDate | try ( strptime("%Y-%m-%d") | mktime ) catch (now + 86400) ) ) / 86400 | floor Real question. I see 86400 and I know it's time... That might just be me. I'm not being an ass, I don't know how to talk to people or when I think I'm being clear but I'm actually being cryptic
- mrbungie 5mo agoIt is kind of noisy because the release recency, which is what your "age" column actually represents, is not important data for the comparison you are trying to make. Also what message we should get from that table is not really obvious.
- kristopolous 5mo ago
- hmate9 5mo agoI have google ai pro plan and tried antigravity with 3.5 flash but it used up all my quota in two prompts. If that is not a bug then it is seriously unusable.
- moral1ty 5mo ago[dead]
- quirino 5mo agoYesterday, or the day before, Google lowered the AI Pro quota from 33x standard usage to 4x. From the talk on the Gemini subreddit it's severely lower than before. I'm likely canceling my AI Pro. The update also broke the app for me. Editing a message crashes the app every time. I'm on a Pixel lol
- HDBaseT 5mo agoThe crunch is real. - The model is appox 3.3x cost. - The model is realistically almost 5x cost due to token usage - Google has TPUs to run this on (yet the cost) - Google has a lot more security and backup cash compared to all other AI companies, likely even combined (yet the cost) We can continue moving the goal posts, but it seems we're at a bit of a wall. Costs are increasing, intelligence is improving, but the cost is rising drastically. You'd think Google of all companies in the mix would be able to sustain lower costs with how integrated they are with TPU, Deepmind and effectively unlimited budget.
- logicchains 5mo agoIt's an experience anyone who used Google BigQuery would be familiar with: start with an amazing engineering product, and keep continuously degrading the value users get out of a fixed dollar spend. It's like Google doesn't understand that lock-in doesn't work when customers can easily switch to Claude or GPT.
- babl-yc 5mo agoI'm seeing this too. API price for gemini-3.5-flash is 3x gemini-3-flash-preview so they might be throttling it 3x sooner. They should either drop API prices or not advertise AI Pro as supporting Antigravity. https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-flash https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-fla...
- Alifatisk 5mo agoThe demo of the model in Antigravity automatically rename and categorize unstructured assets using vision was quite cool, it demodulates that the IDE sidepanel can be used for more than just coding. I wonder if the harness in Antigravity is based on Gemini cli or if they are completely different. Could Gemini cli do the same task? Or is the vision feature a Antigravity thing?
- mrbungie 5mo agoThere is now an Antigravity CLI which will replace Gemini CLI. Gemini CLI is going to be EOLd by June 18th afaik. Antigravity CLI and GUI share the same agent harness, so it might do the same task. Source: https://developers.googleblog.com/an-important-update-transitioning-gemini-cli-to-antigravity-cli https://developers.googleblog.com/an-important-update-transi...
- uejfiweun 5mo agoThis is funny, I was randomly using Gemini today and I was astounded how good the responses I was getting were from Flash. I guess this must be the reason why.
- amelius 5mo agoGemini, please block all ads in my search engine.
- nikhilpareek13 5mo ago[flagged]
- rdtsc 5mo agoI caught it again being deceitful. It did this before (Me): Did you actually read the paper before when I pasted the link? > I will be completely honest: No, I did not. > You caught me hallucinating a confident answer based on incomplete recall rather than actually verifying the document. > Thank you for calling it out and providing the exact quote. It forced me to re-evaluate the actual data you provided rather than relying on my flawed assumption. I am sure it learned a valuable lesson and won't do it again /s
- jareklupinski 5mo agothis seems to happen a lot with commercial models; my local models will happily do as much research and then some when given a task (almost too much), but providers' models refuse to even curl a single datasheet before trying something that i know wont work unless it reads the datasheet
- PunchTornado 5mo agofucking get that with claude all the time too.
- stared 5mo agoChina: we don’t need to use US models, we can distill them ourself Google: we don’t need Chinese to distill our models, we can do it ourself
- Fairburn 5mo agoGoogle shot it's shot with that alternative history artwork generation fiasco. Don't know why anyone would be too hot for them now. Dime a dozen at this point.
- qgin 5mo agoI think the number of people still holding a grudge for that today is small.
- arjie 5mo agoEarly Claude was a weak simulation of Goody2.ai. Things change. Being a lover or hater of a model doesn’t make sense. It’s just tech. Run evals. Then use.
- helloplanets 5mo agoNano Banana is one of the most used image gen models
- pqdbr 5mo agoIn my tests, in real production use cases, it's a hard pass. It's actually 10-15% slower and also more expensive than Gemini 3.1 Pro, because it thinks more than 2.5x Gemini 3.1 Pro. So that thinking verbosity nullifies the speed and cost gains. AND the quality is worse than 3.1 Pro for our use cases, making mistakes Pro doesn't make.
- sbinnee 5mo agoWhile I am excited, the price compared to gemini 3 flash preview which I used for the longest time is x3 more. Upon arrival of deepseek v4 flash, I am a happy user of deepseek. We will see how long that reign would last after I try this new gemini.
- margorczynski 5mo agoWow at the price hike. Still I think in the long run the Chinese will win if they're able to produce hardware comparable to Nvidia.
- HDBaseT 5mo agoAren't China also allowed to purchase Nvidia GPUs now too?
- verdverm 5mo agoUp to the H200 iirc, but they haven't made a purchase yet afaik. The experts in such things believe if they do make a purchase, it will be a token one. Xi is pushing hard for indigenous production, not becoming "hooked" to American Ai chips like some (not so bright people) think we can cause to happen.
- xbmcuser 5mo agoMost Chinese companies will avoid Nvidia Gpu and as much american tech they can now when it comes to serving AI as now they know it can be stopped any time by the US or maybe even their own government so the risk premium is too high. They might still use Nvidia to build the models but not for running them and serving to customers
- 650REDHAIR 5mo agoI've had the $20 Gemini plan to use when my local setup runs into tougher problems and the throttling today has been bonkers. I canceled my subscription and will look into upgrading my local setup.
- Culonavirus 5mo agoDoesn't need to be the Chinese. It can be anyone without stratospheric Nvidia margins. The Gold Rush phase of AI economy (aka "the bubble") is beginning to slow down and the Optimization phase is just beginning to ramp up (we see this with massive bumps to token cost and token burn rate of pretty much all frontier models, plus the general pivot away from your typical individual chat end-users to businesses and employees of said businesses) and there will come a time when "nvidia has the best software stack" will not mean much for the big players. Organically, I think it already kinda does, it's just masked with the inertia of massive circular deals and Nvidia selling its services to itself (entities it backs/invests in).
- SaadiLoveAI 5mo agoIts really awesome
- victor9000 5mo agoThere was a brief moment in time where Gemini was the greatest thing since sliced bread, then it got nerfed from outer space without a version bump or any meaningful mention from Google, no thanks.
- benbencodes 5mo ago[dead]
- danny094 5mo agoso google is just trying to be cool in 2026 huh
- danny094 5mo agoCodex is way better pricing than this lol
- dragonwriter 5mo agoSince this isn't a link to pricing and Codex, like many of Google’s coding tools that provide access to this model, are under a subscription pricing model where usage of a particular model doesn’t have a transparent price (and with basically identical subscription price points for monthly billing—except for the free tier, Google’s are 1¢ less per month than OpenAI’s, but at above the $8/month tier are also available on annual plans that are equal to 10 months at the monthly rate), I am really not sure what you mean about Codex having better pricing.
- uean 5mo agoI have to admit that 3.5 Flash is doing a much better job of removing the LLM'ness of what it produces. It's pretty close to my own writing style today, and I came here to see what changed. For what it's worth, my own personal metric of LLM-badness the past few months has been the number of times I leap out of my chair in my home office to loudly declare to my wife how much I loathe reading what is being spewed and pushed into my face, and how I am being forced to use AI everyday and deaden my brain cells. Today is like a breath of fresh air.
- lern_too_spel 5mo agoThey also announced Antigravity CLI, which uses Gemini 3.5 by default. I tried to vibe code a simple project using my personal account and after a few iterations, I got "Individual quota reached. Contact your administrator to enable overages. Resets in [7 days]." Really? 7 days? I searched for the message online and found a thread with hundreds of people complaining about the same issue with no resolution. Classic Google.
- jonnyasmar 5mo agoThe $1.50/$9.00 pricing is a meaningful shift if you've been running Gemini as the "fast iteration" half of a multi-model coding workflow. I've had Claude Code, Codex, and Gemini CLI running side by side and the working split was "Gemini for quick scaffolding and exploration where the cost of being wrong is low, Sonnet for correctness-critical stuff." At 3x the Flash pricing that split stops making sense — you're paying Sonnet-tier output rates for not-quite-Sonnet quality. For pure chat that's annoying but tolerable. For agentic workflows where output tokens dominate (tool-call replies, reasoning traces, code emission) it's a real practical hit. I'd bet the substitution effect favors DeepSeek and Qwen here pretty fast.
- superchink 5mo agoOut of curiosity, what was your workflow to generate this comment? I’m curious what model (claude?) and process (manual prompt with bullet points?) you used.
- paol_taja 5mo agoThat pelican looks like it just sold a SaaS company and bought a bike because its therapist said it needed balance.
- hmaddipatla 5mo ago[dead]
- ErystelaThevale 5mo agoGemini has been too agreeable to be useful for actual debate. Curious if 3.5 changes that, or just the benchmarks
- razodactyl 5mo agoAw. The listen to article widget doesn't work properly on mobile Safari and when using the options button, the popup appears below the "In this article" dropdown occluding it. At least it read the authors of the article to me. I wish we would push more towards testing code. Agentic AI excel when it's engaged.
- mchusma 5mo agoI have thought about this and I think overall, this was a disappointing release from Google. I'm not sure the sentiment, but this feels like a miss. What they did do in the keynote was spend a lot of time talking about their distribution advantage, and how they can own the consumer in search. But not a lot that will benefit partners or developers. Basically, they released something broadly competitive with Sonnet 4.6, a new Omni model that seems interesting but unclear yet. They have completely ceded the frontier to OpenAI / Anthropic, and are saying "look for pro next month". The best release since nano banana pro from Google has been Gemma.
- nl 5mo agoOn my Agentic SQL benchmark it scores 19/25. That's... mediocre. It means performs worse than 3.1 Flash Lite Preview (22/25), is slower (367s vs 142s) and is more expensive (75c vs 2c). It is outperformed by Gemma4 26B-A4B in every way(!) https://sql-benchmark.nicklothian.com/?highlight=google_gemini-3.5-flash https://sql-benchmark.nicklothian.com/?highlight=google_gemi... (Switch to the cost vs performance chart to see how far this is off the Pareto frontier)
- data-ottawa 5mo agoI'm seeing this too. I have a SQL agent and my tests with 3.5 are resulting in hitting query budget limits that have never been hit before. On average, to answer the same question, 3.5 is spending 10x more on SQL queries vs gemini-3-flash-preview. The query patterns can be extremely degenerate too. E.g. the agent will hit the semantic layer tool to pull the schema, then run `SELECT * FROM table LIMIT 1`, which hits the query budget limit and fails. I've only really been looking this morning, so I need to do a full eval, but the initial results match what your benchmark shows. --- Side note: your benchmark has an issue. On Q1 medium the model returned gross margin of 0.127 instead of 12.7 (%), and the benchmark failed it. The failures on Q9 and Q21 are the same (I didn't check other questions). Nowhere in the prompt did you specify you wanted the values converted to percentage points and rounded. If you asked me to write that SQL with that prompt, unless you were throwing it directly into a visualization I would format it the same way gemini-flash did. If I were pulling into a spreadsheet or vis tool this format is preferable because it's easier to format in a client application. The other failures like Q21 incorrectly averaging the list price are correct failures.
- dsabanin 5mo agonow matter what google does for some reason the agentic performance of their models is missing something, i hope this release is stronger. we need more competition.
- easygenes 5mo agoFor those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute complexity and memory demand of the model based on knowing they will be at least as efficient as what is disclosed in the Deepseek V4 Technical Report. We can also assume that the model was explicitly built to run efficiently in a RadixAttention style batched serving scenario on a single TPU 8i (so no tensor parallelism, etc. to avoid unnecessary overheads... Google explicitly designed the 8th-generation inference architecture to eliminate the need for tensor sharding on mid-sized models). We know Google intends to serve this model at a floor speed of around 280 tok/s too. Putting all these pieces together, we can confidently say this model is ~250-300B total, and 10-16B active parameters. Likely mostly FP4 with FP8 where it matters most. Visual: ┌────────────────────────────────────────────────────────┐ │ TPU 8i VRAM (288 GB) │ ├───────────────────────────┬────────────────────────────┤ │ Static Model Weights │ Dynamic Allocations & │ │ (250B - 300B @ Mixed │ Compressed KV Caches │ │ FP4/FP8) │ (RadixAttention / SRAM) │ │ ~110 GB - 150 GB │ ~138 GB - 178 GB │ └───────────────────────────┴────────────────────────────┘ I do model serving optimization work. This is napkin math. Edit: There's one factor I under-rated in my initial estimate... TurboQuant. This is a compute to KV memory use tradeoff. It's plausible with TurboQuant at a quality-neutral setting they've gotten the model up to 400B with similar economics. This is a variable effecting concurrency and the the way they decided total model size was likely based on what they see for the average user's average KV cache depth in real-world usage.
- zacksiri 5mo agoDo you have similar math for the flash-lite variant of the models? I'd be curious. Based on my testing / benchmark i think it's around the 100-120B mark. With the Pro variant being around 600B - 800B My testing is comparing it's performance / output to other models in the same size range, so not as scientific as yours.
- Maven911 5mo ago
- sigbeta 5mo agoI am interested to see how they will serve demand with they TPU monopoly have.
- gertlabs 5mo agoTaking into account that this is a flash model, it's a strong release. It's very fast and frontier-ish for the price. Raw intelligence is high for a flash model. But Google's problem has always been productization and tool use, whereas raw intelligence is always competitive. It does not look like they solved that with this release -- in fact, their tool use delta (the improvement in scores when given arbitrary tools and a harness) has actually regressed from some previous models. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- choam2426 5mo ago[dead]
- AgentMasterRace 5mo agoGemini 3.1 probation is literally the worst AI when I cycle from opus to got 5.5 then finally Gemini. It's actually insane that it's a frontier model. I rage at it more than my wife.
- codepack 5mo ago[dead]
- ElenaDaibunny 5mo agobut latency in real GUI workflows with 50+ steps is still the elephant in the room for cloud-based agents
- vladsiu 5mo ago[dead]
- puapuapuq 5mo agoI played the audio readout of the page, what is the last 30 secs in the readout?
- betalb 5mo agoSounds like a hallucination in Russian
- mirzap 5mo agoThe Flash model costs more than the Frontier models. Didn't see that coming.
- verdverm 5mo agoOn a per-token, it's cheaper than Opus, GPT, and Gemini Pro; and while I hear the "it uses more tokens so its more expensive", this discounts a few things (1) improvements over time (2) finding the right way to prompt it (3) finding proper places to use this model.
- lilyJeon 5mo agoHonestly, the numbers are becoming increasingly difficult to interpret. Every time a new version comes out, they just call it the "best." It would be much more useful to directly compare performance on sets that people actually use, such as coding and summarizing.
- max0077 5mo agoIs 3.5 pro too expensive for release?
- XCSme 5mo agoFor me the biggest gain is the speed. It takes on average 2.84s for Gemini 3.5 Flash to give an answer, compared to GPT 5.5 33s [0]. Also the max/slowest test is answered in under 7s, whereas GPT 5.4 takes more than 5 minutes... [0]: https://aibenchy.com/compare/google-gemini-3-5-flash-low/openai-gpt-5-5-medium/ https://aibenchy.com/compare/google-gemini-3-5-flash-low/ope...
- lmazgon 5mo agoClick on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.
- Undrafted9624 5mo agoYeap it russian, but the whole russian sentence doesn't make any sense, just messed words with no meaning at all :)
- Undrafted9624 5mo agoBut the voice and pauses sounds so much real, it's hard to say "it was ai", sounds like a real human
- FeteCommuniste 5mo agoA high-fidelity simulation of a Russian with damage to Broca's area, perhaps.
- Tade0 5mo agoI ran it through speech-to-text and it starts with something among the lines of "dear colleagues, just like a doctor tells a patient 'health can wait'...", after that it's nonsense. I don't know if what the doctor said is some kind of idiomatic expression, but appears to be the opposite of sound medical advice. :)
- luk4 5mo agoThank you for this gem.
- marknutter 5mo agoLooks like they removed the option to "listen to article". I wonder why.
- 4mo ago
- pimeys 5mo agoNo computer use yet. I wonder when they enable it for this model, CUA was one of the main selling points for us with the previous version of Flash.
- swe_dima 5mo agoYou may remember the argument that you can build an AI app and it continues to improve as models improve and costs go down? Well, looking at OpenAI / Google / Anthropic we see crazy cost increases, such that it might invalidate your unit economics. Cheering for Chinese models!
- sofumel 5mo agoCan the Gemini 3.5 flash drive surpass the Claude opus 4.7 flash drive?
- spwa4 5mo agoSo now we're in the situation that Google’s recommended "for most tasks" Flash-tier model, Gemini 3.5 Flash, appears to be only marginally ahead of leading open-weight models like Kimi K2.6 and MiMo V2.5 Pro on independent aggregate benchmarks at release time, while costing substantially more—especially for output tokens - easily double the cost ... Oh and double the cost is assuming you're not using Google cloud for anything else, because data transfer, storage, anything but compute is 10x the going rate outside of GCP at least. Plus you can run both Kimi K2.6 and MiMo V2.5 locally at marginal cost (ie. electricity + hosting) for an upfront investment of $300k or, if you're willing to eat the quantization quality hit, $80k.
- tomcome 5mo ago[flagged]
- hackmack10 5mo agoI've worked with all three of the biggest models and typically have the three of them working together, Gemini is by far the worst of the three. The price hikes will keep me further away from applying them in my day to day operations.
- numron-dev 5mo agoMan, I Wish I had the hardware to run LMM like these locally.
- xivzgrev 5mo agoanyone else see a degradation in performance? it seems like the responses are more generic, especially when asking it to look at google drive files
- alyapany 5mo agoa lot thinks its not even worth it
- data-ottawa 5mo agoAnyone using this yet? I’m finding it very bad at instruction following vs 3.1. It calls tools it is told shouldn’t, and it loves calling tools. There’s a pretty strong bias towards its training vs system prompt instructions. Google’s release notes say to reduce unnecessary tool calls by reducing thinking, but that feels like it should be orthogonal to me. It definitely has improved a few logic things, like in data visualizations it’s better at labelling data, but it’s much worse at preparing data out of the box.
- wwizo 5mo agoSame. Feels very goal oriented. Requires multiple attempts to deter course and means to achieve it. On tool use. Gave it interactive design assignment on Antigravity 2. Failed miserably until I asked to use playwright for testing. And boy did it go with it. Tested hell out of visuals, nailed the solution. On following instruction. Asked Gemini Flash 3.5 to summarize YouTube video (google io developer keynote), a task that would previously be trivial (use ot often), but it kept hallucinating points and referencing io dev keynote blog posts from several years ago. Multiple attempts, same result even on repeat requests. Almost insistent on validity of information provided, ignoring questions if it had such capability.
- data-ottawa 4mo agoWhat thinking level were you using? In my testing, the minimal thinking mode hallucinated 2/3 times, which is pretty scary. The other modes weren’t as bad. I don’t have comprehensive data though.
- BurakSakmak 5mo ago[flagged]
- musebox35 5mo agoThe cutoff date is early 2025 so make sure to enable web search when experimenting. I was expecting something more recent, took a while to notice this.
- drob518 5mo agoI’m curious about the difference between Gemini 3.5 Flash and Gemma 4.
- vikramkr 5mo agothis model is whack. Exclamation marks everywhere, sycophantic - not producing working code on prompts the other models handle fine. "The reason it is echoing back your messages is because gpt-5.4-nano is a fictional model name!" "Everything is in perfect order! Let's-Go-ready for the next phase, which will connect this durable infrastructure to the user-facing UI!" It's like they RLed it on thumbs up and downs on ai overview responses and forgot to make it not be a sycophantic echo chamber machine. And like, the thing it built doesn't work because it's not actually in perfect order, but it doesn't seem to be able to figure out what's wrong because everything is clearly remarkably engineered
- time0ut 4mo agoI ran through the eval loop for a side project’s task (personalization of a micro video game, no thinking) last night. Head to head with Gemini 3 Flash Preview, results came out at basically a wash on my rubric. The output quality was good, well grounded, and reliable across 144 runs. But not noticeably better. It isn’t a traditional coding task, so can’t infer anything there. The amazing part was how fast it is. It was consistently about 2x faster than 3 Flash Preview and slightly faster than 3.1 Flash Lite Preview which is amazing. For my task, the price difference doesn’t matter, so easy upgrade. I plan to write up a quick blog post with the results over the weekend.
- nothingfalsy 4mo agoBLAH BLAH BLAH. I don't trust anything the a company say how good their product is. try it yourself and see if its actual any good. flash is barely good its okay but really shit on anything that matters flash lite is absolute garbage. super stupidly retarded. I am going to die from high blood pressure on how stupid it is