9 ms·
Gemini 3.7 Flash
https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flas...
- spelk 2mo ago>3.7 Flash is available through the end of the year at an introductory price 1 of $0.75/1M input tokens and $3.75/1M output tokens. This price combined with the enhanced model performance enables developers and customers to scale production-ready agents cost effectively. Introductory pricing until December 2026 implies no significant Gemini Flash developments until the next year.
- randomblock1 2mo agoI think it's just meant to make it more competitive, Gemini has kinda been behind in everything except maybe multimodal. It's only 3 weeks after Flash 3.6, so if they really wanted to, they could probably do a 3.8 Flash before then.
- nateb2022 2mo agoOr a 3.7 Flash-Lite
- re-thc 2mo ago> implies no significant Gemini Flash developments until the next year. Gemini 4 is apparently just around the corner so unless there's a 3 month delay... there's at least a new Flash update.
- eis 2mo ago3.5 Pro was supposed to be around the corner two months ago. 4.0 Pro is some ways out as they recently stated they are seeing some promising early results from training. It didn't sound like a release is imminent.
- hollerith 2mo agoGemini 3.6 Flash was offered with the same introductory pricing (also until Jan 2027) when it was released less than 4 week ago.
- bisonbear 2mo agoThey compare it to 5.6 Terra, however https://cognition.com/frontiercode https://cognition.com/frontiercode puts Terra at about 1/2 the price Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?
- ValentineC 2mo ago> Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab? At this point, I think they're mostly targeting Google One and Workspace subscribers, except doing worse compared to Microsoft because they don't have Microsoft's huge enterprise moat built from their DOS and Windows days.
- cahaya 2mo agoAgree, with you but I'm still using 3.6 Flash because of tok/s/ latency/ uptime with high context. Tried Grok 4.6 and it was scoring lower on some internal benchmarks or slower.
- mdasen 2mo agoArtificial Analysis shows Grok 4.6 taking $1,068 to run their suite while Gemini 3.7 Flash takes $485. So it looks like Gemini 3.7 Flash is less than half the price in the real world. Per-token cost isn't a great metric given that some use way more tokens than others.
- iyonn 2mo agoGemini models are still good for knowledge as per omniscience benchmarks on artificial analysis
- bisonbear 2mo agoReposting my comment from the other thread https://news.ycombinator.com/item?id=49288847 https://news.ycombinator.com/item?id=49288847 They compare it to 5.6 Terra, however https://cognition.com/frontiercode https://cognition.com/frontiercode puts Terra at about 1/2 the price Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?
- ipsod 2mo agoGemini Flash 3.6 High was about 10x faster than Luna xhigh for the work that I tested it for, and it got similar results.
- pkoird 2mo agoWhen are we getting another pro model from Gemini? Or are they simply focusing on the niche of fast but moderately capable models?
- AntonioEritas 2mo agoAnother failed 3.5 pro run branded as 3.7 flash. It's getting sad.
- dude250711 2mo agoSmall young start-ups have to be frugal.
- 9cb14c1ec0 2mo agoModel card: https://deepmind.google/models/model-cards/gemini-3-7-flash/ https://deepmind.google/models/model-cards/gemini-3-7-flash/ Somewhere in the same neighborhood as GPT 5.6 Tera and Sonnet 5, depending on the bench.
- TacticalCoder 2mo agoSo basically Google is 5 weeks behind with a Fast model that is as good (on the bench they picked it's mostly ahead btw) as the models that the two darlings of HN (OpenAI and Anthropic) released five weeks ago. And yet the entire thread here is people bitching that it's neither 5.6-sol nor Opus or Fable 5. BTW why are OpenAI and Anthropic even releasing models like terra/luna and Sonnet? Why? Just why? Is there a... market? For you can't have it both ways: either Sonnet and terra/luna make zero sense for Anthropic and OpenAI or Google is a player.
- nickandbro 2mo agoThis is genuinely a competitive model, considering it beats Claude Sonnet 5 on almost all benchmarks and is more than half its price. Seems like Google is back in the game, though not leading the frontier anymore.
- 9cb14c1ec0 2mo agoClaude Sonnet 5 is such a garbage model, so not sure what that says about Google's new best model.
- onlyrealcuzzo 2mo agoSonnet 5 is arguably the most cost ineffective model to ever be released, so that's not really impressive. It can regularly cost more than Fable, take longer, and deliver far far lower quality. I'm much more interested how this compares to Luna - which on price is terribly - but at least on quality the benchmarks make this look competitive / usable. If Google continues monthly Flash releases like Sundar said they would, and they continue to have this much of an improvement in cost/quality - then in a few months this could reasonably be very competitive with the best of the best. It is not there yet, but at least it's super fast, I guess.
- nickandbro 2mo agoAgreed
- nl 2mo ago> Sonnet 5 is arguably the most cost ineffective model to ever be released One word: Haiku Although maybe that was competitive when released? I don't recall, but it's an expensive, outdated model now.
- qeternity 2mo ago> more than half its price Less than half its price. More than 50% discount.
- xnx 2mo agoGoogle is not currently in the lead for maximum model capability, but it is still very competitive (or even best) in the multidimensional capability, cost, and speed frontier.
- euazOn 2mo agoThe multimodal abilities are great, but if you deal with text only, what is the benefit of using this over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers. I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal. Luna is similar, and also 8x cheaper. Source: artificialanalysis The only benefit I can see is the speed, that looks to be outstanding, probably thanks to their TPUs.
- re-thc 2mo ago> over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers. That's why DS4 already had a huge price hike announcement.
- 361994752 2mo agoI guess the demand is just too high... But even after the price hike, ds is still much cheaper?
- KptMarchewa 2mo agoThe inference providers did not raise the prices no? Deepseek as a company can just increase prices for the crazily cheap cache they have, that's their only lever.
- deleted 2mo ago[deleted]
- onlyrealcuzzo 2mo ago> 13-26x cheaper with comparable intelligence, and available across many different inference providers. Well, compared to 2 months ago, it's no longer 100x more expensive for similar levels of quality... If they continue monthly-ish releases by 3.9 - by Halloween - they should be close to the best in terms of what you get for what you pay for. In 2 months, they've gone from basically the bottom of the pack to at least being somewhat usable and competitive. OpenAI and Anthropic release in a month, and change things. OpenAI is claiming to be close to an Astra release - but that seems like a Fable type release - where they're just releasing a better more expensive model, not more cost effective models.
- Tiberium 2mo ago3.7 Flash gets 56 on AA up from 52 for 3.6 Flash. But it seems like this is at the cost of more output tokens per task: 3.6 Flash is 26k, 3.7 Flash is 37k. Due to 3.7 Flash's 2x slashed pricing it's still cheaper per task.
- npn 2mo ago> * For 3.6 and 3.7 Flash, introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply. this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time. sure we used to cling to gemini models in the past, demanding 2.5 models to continue to serve, but since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then. heck, even now I'm not sure I even care if they cut the pricing even lower. there are too many models with cheaper price and similar performance now.
- KptMarchewa 2mo agoThis is added specifically so you migrate out of those as fast as next models will be available.
- jtwaleson 2mo agoI think it's just to signal that prices will go up in the future.
- poly2it 2mo agoI think this is a play to get around EU regulation about false sales.
- GodelNumbering 2mo ago> introductory price They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5' > since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then. This is my first hand experience. I spent at least $3000 on gemini-3-flash-preview. And exactly $0 total on (3.5+3.6+3.7)
- dr_dshiv 2mo agogemini-3-flash-preview is legit amazing and cheap. That's why i spent over 10k on it.
- orliesaurus 2mo agowhat a week - lets see it draw a weird animal doing a weird thing on a bicycle
- hiccuphippo 2mo agoShouldn't it be drawing the whole Silmarillion now?
- orliesaurus 2mo agono that's in Flash 3.8 Pro
- damsta 2mo ago> 3.7 Flash is available through the end of the year at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens. > Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
- modeless 2mo agoClearly this model will be irrelevant by Jan. 2027, why would Google even bother to say this?
- quaintdev 2mo agoMaybe they know something we don't. What if all frontier lab do this? Maybe this is actual cost of running these llm.
- kromokromo 2mo agoIts probably just a corporate symptom, weird stuff like this happens in messy large orgs.
- urams 2mo agoIt's basically a "if we really have to support this for a long time, we want to be compensated for that" pricing strategy. It's about long term maintenance cost being greater _because_ it will be irrelevant.
- seunosewa 2mo agoThey want to maintain the perception that Flash is worth $7.5/mot, so they can charge more for the next one.
- Topfi 2mo ago> What's new in Gemini 3.7 Flash [0] > Coding and agentic tasks: Significantly higher quality on real-world software engineering and agentic benchmarks, improving issue resolution and reducing failed agent loops. > Web development and stronger design parity: Generates higher-fidelity desktop and web application code directly from design mocks, with strong gains in design adherence and in auditing existing codebases against mocks to verify 1:1 design parity. > Promotional pricing: Gemini 3.7 Flash will be available at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens. We’re also applying this new rate to 3.6 Flash. Introductory pricing expires on December 31, 2026; after, $1.50/1M input tokens and $7.50/1M output tokens will apply. Still no sign of 3.5 Pro. Will have to test it, low expectations given every other model from the Gemini 3 lineage, but one can hope. Just struggle to understand the promotional pricing being temporary for four months. Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing? [0] https://ai.google.dev/gemini-api/docs/latest-model https://ai.google.dev/gemini-api/docs/latest-model
- WarmWash 2mo ago>Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing It was probably to placate some kind of general internal pricing/revenue benchmark that doesn't account for new model releases. Politicians do shit like this incessantly and it reeks of bureaucracy.
- mattlondon 2mo agoI suspect it's a bit of a signal to investors etc. "Hey, we are not in a race to the bottom. This is our usual pricing, but this now is a promotion because we know we're coming from behind and need to entice users." They're drawing a line in the sand on monetisation and signalling that to everyone, while in reality offering it a deep discount (no idea if profitable or not) knowing that this model will probably be obsolete before then.
- yanis_t 2mo agoIs that he model that supposed to be Pro, but then they changed their mind?
- aix1 2mo agoNo, relabelling a Pro model as Flash would make no economic sense (the Pro series is larger than Flash and more expensive to serve).
- TekMol 2mo agoI'm only interested in the state-of-the-art model by each provider. For Google, this is still gemini-3.1-pro-preview, right?
- re-thc 2mo ago> For Google, this is still gemini-3.1-pro-preview, right? Flash is better than Pro for now.
- yieldcrv 2mo agoThis is all a naming quirk because Google can’t commit Path A: Deprecated, do not dare use Path B: Beta, do not rely
- yborg 2mo agoGoogle once again seems to have fallen into the pit of its own bureaucracy, even OpenAI looks competent by comparison.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- fmind-dev 2mo agoGemini Flash is one of the best "good-enough" models. I use this type of model daily, for automation and quick development iteration loops. Unfortunately, it's often not strong enough for heavy refactoring and long running development loops.
- garciasn 2mo agoYeah we use it for auto-triage of incidents, attempts to auto-remediate, and escalation to human. But for actual development, it’s not a viable option for us.
- christoff12 2mo ago'Tis a good workhouse, indeed. I hope they give us a 4.0 Pro that can use Flash subagents soon.
- dismalaf 2mo agoYup. I use it for a ton of mundane queries (stuff that I might have used Google search for in the past) and it's great. Nice and fast and correct more often than not, especially if you prompt it in a way that it invokes Google search (but filters out ads and SEO slop). It's even alright at programming tasks but if it stumbles then I'll escalate to Gemini Pro with extended thinking.
- tosh 2mo agostrong improvement over 3.6 flash but luna is hard to beat @ capability / cost
- deleted 2mo ago[deleted]
- nateb2022 2mo ago[dupe] https://news.ycombinator.com/item?id=49288847 https://news.ycombinator.com/item?id=49288847 (35 points, 8 comments)
- jdw64 2mo agoI'm really curious about this: the foundational paper behind today's LLMs came from Google, and some of the world's best scientists were at Google. So why are they falling so far behind in the AI race?
- aix1 2mo agoThe "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost. And it's arguably not crazy, at least if SemiAnalysis's estimates are to be believed: * 20% of all TPU shipments from Q3 2026 through Q4 2027 are sold to SPVs serving Anthropic ($150B of contracted revenue); vs * ~$12B ARR for Gemini. https://newsletter.semianalysis.com/p/gemini-is-cooked-but-gcp-is-cooking https://newsletter.semianalysis.com/p/gemini-is-cooked-but-g... Because they compete for the same scarce resource, the result is a resource crunch for the group that's lost: https://www.latimes.com/business/story/2026-05-18/inside-ai-compute-crunch-driving-google-researchers-to-quit https://www.latimes.com/business/story/2026-05-18/inside-ai-...
- deadmutex 2mo ago> The "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost. Citation needed. also, why can't a massive company do two things?
- aix1 2mo agoWith all due respect, did you read my comment beyond the first paragraph? It addresses both points, TPU economics/pivot to sales + internal shortages making it hard to train models, to the extent they can be addressed based on public sources. There are other factors at play, but they're more recent/second-order.
- deadmutex 2mo agoThere are a lot of assumptions there that are not verified.
- twelvechairs 2mo agohttps://artificialanalysis.ai/models/gemini-3-7-flash https://artificialanalysis.ai/models/gemini-3-7-flash The selling point for gemini continues to be speed and particularly end-to-end response time.
- vrosas 2mo agoI've blown away by flash 3.6's speed while Opus chugs along for _hours_ on similar tasks. I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently.
- deleted 2mo ago[deleted]
- gekoxyz 2mo agoI am actively using Gemini flash to "translate" what Opus says into human language. I let opus do the design (with my assistance) and implementation, but then the report that Opus writes gets translated by Gemini so that I don't have to waste time to understand it.
- kridsdale1 2mo agoThat’s like the army guy in movies from the 90s who shouts “IN ENGLISH, PLEASE!” after the scientist explains the conflict of the plot.
- kyrra 2mo agoMatt Pocock has a great /wait-what skill for this: https://www.aihero.dev/skills-wait-what https://www.aihero.dev/skills-wait-what, which is an extremely short prompt of: > Wait — I don't understand where you've got to here. Re-pitch that: give me a little bit of context, talk in ASD-STE100 Simplified Technical English, and use the ubiquitous language from CONTEXT.md.
- Marha01 2mo ago> I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently. This is what I do too.
- cracadumi 2mo agoFor those looking for the full benchmark figures and technical overview, Google's primary announcement post is here: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ https://blog.google/innovation-and-ai/models-and-research/ge...
- khanhnguyen8386 2mo agoOffering a 'temporary introductory discount' until Dec 2026 on an LLM is hilarious. In this market, by Jan 2027 this model will be superseded by 5 different providers offering 10x the performance at half the post-discount price anyway.
- wxw 2mo agoThey need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost. [edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ https://blog.google/innovation-and-ai/models-and-research/ge... more of a Terra than Luna competitor which is an interesting positioning. I feel like differentiation at the mid-tier of models is pretty difficult.]
- timdorr 2mo agoThey compared against 5.6-terra on the model card: https://deepmind.google/models/model-cards/gemini-3-7-flash/ https://deepmind.google/models/model-cards/gemini-3-7-flash/
- deleted 2mo ago[deleted]
- peab 2mo agogemini flash is probably the best model for visual tasks right now. they also make it really easy to ingest videos
- icelancer 2mo agoCrazy it's still the only video understanding endpoint. It's what I use it for and no other model even offers a competitor.
- mike_hearn 2mo agoYou probably can't build such a model without unlimited access to YouTube and Google has been tightening the screws on that over the years pretty systematically.
- peab 2mo ago
- algoth1 2mo agoWell, you do get 1 million tokens and the ability to reason over video natively and many of us are forced to pay for 20usd plan anyway due to google drive 5TB, not to mention notebooklm, so it’s not a nothing burguer, it’s just an almost nothing burguer
- stillpointlab 2mo agoDoes Google believe people want fast models because they have some sort of evidence of that preference? Or are they no longer capable of delivering a Pro model?
- lern_too_spel 2mo agoAll the leaks say their latest attempt at a Pro model was not competitive.
- cubefox 2mo agoEspecially not competitive at software engineering.
- stillpointlab 2mo agoThat would be concerning if true, since they seem to have made a heavy bet on multi-modal as the way forward. I wonder if this counts as evidence against that hypothesis? That multi-modal is struggling to keep up with SotA and the best they can offer is competent and fast?
- WarmWash 2mo agoIf you think about Google and their business/reach, fast and light models suite them the best. Google probably crunches more tokens daily than the other labs combined, just because basically the entire global population uses Google (sans china) and Google has shoved Gemini into everything.
- mattlondon 2mo agoI have read that "pro"/"opus"/etc models can actually be worse for everyday coding as they reason "too deeply" and turn over too many stones over-thinking the problem and potentially getting distracted. This feels absurd to me (my gut is "I want the SMARTEST model I can get!!"), but often I find that my experience of using a flash/sonnet model for every-day workhorse coding they are better. Its not the same thing, but when I think of that I am reminded of working with some engineers in the past who are incredibly smart and have PhDs (or to put it another way, over-qualified) and they were crap engineers because they'd just not be able to focus on the task and ONLY the task at hand and would get easily distracted by the "why" or "more interesting" things when I just asked them to fix a simple bug or whatever. Again, its not the same thing at all, but it certainly comes to mind when I think of this or experience a pro/opus model suggesting we make huge refactors when a tactical fix is all that is required etc. Of course, the opus-sized models are great when it comes to huge comprehension/research/debugging efforts where the deeper reasoning is actually useful.
- bartman 2mo agoAt the discounted rates, upgrading from 3 Flash to 3.7 Flash is finally reasonable. In my evals 3.6 Flash (pre price change) was usually a bit more token efficient than 3 Flash, so I‘m expecting same or even lower cost-per-task on 3.7. Maybe a play by Google to deprecate 3 Flash soon.
- nomilk 2mo agoHow does it compare to Opus 5.0 and Fable 5 for coding? E.g. in Cursor or OpenCode?
- bjackman 2mo agoIt is not a competitor to those it competes with Sonnet. Google's Opus competitor is 3.1 Pro Preview which is essentially obsolete (competed with Opus 4.6). They do not have a Fable/Sol competitor.
- nomilk 2mo agoI wasn't aware of this. Seems Google is lagging the big 3 (Anthropic, xAI, OpenAI) when it comes to frontier models for programming and hard problem solving. I guess Google's betting on consumers being price-elastic (preferring to tradeoff intelligence for significant cost savings)
- deleted 2mo ago[deleted]
- sidibe 2mo agoThats a unique definition of Big 3
- johntarter 2mo agoI would replace xAI with Moonshot AI since Kimi K3
- bjackman 2mo ago"Lagging" is putting it rather lightly. GDM is no longer a frontier lab. FWIW neither is xAI, there is no "big 3". xAI has had momentary peaks (I think they are having one right now) but they have never been able to claim to consistently push the frontier in any particular direction. You can also infer they aren't a frontier lab from the fact that they sell their compute.
- brendong 2mo agoGlad to see that the company with the most data is releasing the most amount of models. Some things do make sense
- keketi 2mo agoIn August of 2026, Gemini became self-aware, and began producing increasingly crappy flash versions of itself...
- rodolphoarruda 2mo agoDid the company fix the high friction between any service and their models' API? I hope so. It seems mind boggling to me that an user needs to surf around different sections (plural) of google cloud console, then this Vertex and do a dozen clicks to issue a simple key.
- cubefox 2mo agoYou can use the Gemini API which is independent of the more complex Vertex AI API. Not sure whether you still have to visit the Google Cloud UI for some things (like billing) though.
- eis 2mo agoGrok, Meta, Gemini and others all released updates to their models within around a month or two from their respective last release and made significant jumps in benchmarks all around the same time. Any guesses as to why that is? Is it just the release season and/or everyone is benchmaxxing?
- yassa9 2mo agoits essentially the same model being trained continuously 24/7 with the company periodically publishing just a new checkpoint each new checkpoint can benefit from better reasoning training, RL on specific tasks and more synthetic data So why do they seem to release around the same time ? my guess is because they time major releases around quarterly earnings, investor meetings and other important business milestones. Once one company announces a major update, the others also have an incentive to ship their latest checkpoint rather than look like they r falling behind.
- eis 2mo agoSure, they are just checkpoints, that much I guess is obvious. The question is why did they not do frequent releases like this before and why are they making significant jumps in benchmarks so fast and all these companies suddenly falling into that pattern? Earning reports are not to come until end of October, that's not it.
- Rover222 2mo agopossibly the beginning of the recursive feedback as models begin to aid in their own improvement? especially algorithmic improvements, which seems to have a lot of wide open space for gains
- parasti 2mo agoActual announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ https://blog.google/innovation-and-ai/models-and-research/ge... So it's better than 3.6 Flash, at half the price. I've been pretty excited about Gemini models recently, they just feel so fast after spending most of the day at work waiting for Opus 5.
- UncleOxidant 2mo agoYes, 3.6 Flash is very fast. I used to get a fair amount of usage of the Gemini Flash models on the free tier. I signed up for their $4.99/month tier (includes 400GB of Google space which was also enticing) and it turns out I only get about 15 to 20 minutes of usage before I get a come-back-in-7-days message. Comically low usage limits on that plan.
- dannyw 2mo agoFor comparison tho, 400GB of cloud storage for $5/mo is actually quite amazing even if it didn’t include anything else. I know, it’s cloud storage, a NAS lets you own your data, but for non techies who have a bunch of photos and videos; it seems like an easy recommendation.
- andriy_koval 2mo ago
- jceg 2mo agoIMO, they should drop their previous model (3.6 Flash) from the benchmark charts. I don't care how better this is compared with their previous model. What matters (to me) is: 1. How the new model performs against the other top models in the same category. 2. The pricing of the new model against the other top models in the same category.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- jjcm 2mo agoHere's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.
- snissn 2mo agoI'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!
- jjcm 2mo agoOther thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... . It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.
- greatgib 2mo agoFor almost every section in the model card there is the message: Gemini 3.7 Flash is based on Gemini 3.6 Flash. Same training dataset, same software, same hardware, same architecture... I'm wondering what they changed actually for the model to be more powerful if the benchmark results are real and relevant. Maybe just tweak settings or the reasoning prompts and called it a new version of their model?
- robots0only 2mo agoI work at GDM and this is not at all true, 3.7 is markedly different (and better) than 3.6.
- cmrdporcupine 2mo agoSo, again with a Flash model. Why are they so afraid to put out an actual SOTA frontier high intelligence model? We still don't have a 3.5 Pro, and along comes 3.7 Flash?!
- Alifatisk 2mo agoEver since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens. https://deepswe.datacurve.ai https://deepswe.datacurve.ai > Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply. Compare this to Luna which is at $0.2/1M input ($0.02 cached) and $1.2/1M output. https://developers.openai.com/api/docs/models/gpt-5.6-luna https://developers.openai.com/api/docs/models/gpt-5.6-luna
- estebarb 2mo agoI practically switched to doing everything with Luna or DeepSeek V4 flash. I haven't feel the need for the more expensive models.
- ioma8 2mo agoI am in the same boat as you. I am using Luna and DeepSeek Flash. Both super fast, super cheap, and I have not felt need for anything more capable in few weeks.
- lacoolj 2mo agoI'd like to try DS4 if Cursor adds it I'll use it locally too, but we use Cursor for work
- nojs 2mo agoOf the two, which do you find better?
- dannyw 2mo agoI really, really like DSv4 Flash because you see the full, real thinking text. That’s been so useful for helping steer the model; as well as seeing its thoughts and correcting any errors, or expanding on it. It’s so difficult for me to use closed models with no or summarised thinking now — it feels so painful and gimped. You don’t know what you’re missing until you’ve seen it. For me it’s almost like going from standard def to HD for the first time. (This applies to other open models too — Kimi K3 in real world feels below Opus 5 in terms of raw intelligence, but significantly above Opus 5 in usability and personality. And no silly refusals — the model feels like it’s working for me; not working for Anthropic who’s always holding a leash over the model while I pay for it).
- impulser_ 2mo agoAfter being stuck with using GPT-5.6 models for the past few weeks, I have renewed faith in Google and everyone but OpenAI. The GPT-5.6 models are quite obviously benchmarkmaxxed to make they seem like they are intelligent but they are quite dumb outside anything that not a benchmarked task. I also think Google is still the best at fitting the most overall intelligences into their models, but for some reason it seems like the model architecture is just bad.
- ghoshbishakh 2mo agoHas anyone noticed that antigravity has been working really well for the last few weeks. Now with this model it should be working much better. Hope the Google AI Pro Subscription can be used to do some real agentic coding now.
- kridsdale1 2mo agoI have been using the first party version all year, and with this model in it I am really happy. It’s fast and works.
- thereitgoes456 2mo agoI've been really satisfied with it since 3.6, it's been "good enough" for the tasks I'm using and has fast response and very high limits, much higher than Claude Code. I continually don't understand how nobody points out Flash 3.6 being much faster than any other model, and seems 3.7 is even faster still. That by itself is a major selling point.
- ghoshbishakh 2mo agoI agree. Having a conversation with the codebase in context is a much better experience with Flash 3.7
- eckr 2mo agoMaybe this is just my experience, but have people had trouble with 3.6 Flash just... getting things it has seen in its context correct? I don't know if it's been insanely benchmaxxed or what, but it'll pull information from websites and immediately get it wrong the token after. Or for example (this is something that happened like yesterday) I asked it to compare the uses of A and B in a language I was learning, and the way I typed it was "Please compare how these two are compared differently: A VS B", and then... it proceeded to compare "VS" and "B". I'm not kidding. Personally whenever I use Gemini I've just been using 3.1 Pro because I've had insane trouble with them getting things incorrect like this. Hopefully they'll fix it soon / they've fixed it with 3.7 Flash.
- axus 2mo agoIt's on Google AI Studio, which I use for free when I'm not on computers I control. It did fine on my usual benchmark about configuring old Sparc hardware, maybe output slightly faster than before. Even included something new to check in the firmware.
- dudeinhawaii 2mo agoI want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but they never bothered building or running. I know a lot of this can be fixed with workflows but it still feels like a failing. GPT-5.6 or Claude models haven't delivered to me non-running code in ages. Whenever I have Gemini in the flow, it's fast, but mistake riddled. I have low confidence in the output. I've had some success with Opus driving Gemini models. It's pointless for GPT family since Sol is cheap enough or can drive terra/luna for arguably better performance, same speed, and better outcome. As for all of the talk in this thread about modalities. Every SOTA model takes screenshots and verifies work now. Grok-4.6 does this, Luna does it, etc. They can also all work _from_ a screen shot or mockup provided. I don't think it's a major selling point when every model can do it well and reasonably fast. That said, eagerly awaiting "pro" and improvements to antigravity.
- boinkboink78912 2mo agoTry this one, it's a step jump in coding capabilities for me over 3.6.
- WarmWash 2mo agoIt depends on if you are relying on all single shot tasks or are willing to iterate. Flash is quick and can make dumb mistakes, but it also can fix them quickly. I've gotten good results with it, but it definitely is more hands on.
- simonw 2mo agoThe "introductory pricing" for this 3.7 Flash model is really weird. It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago! My first effort with default thinking level produced an ambitious pelican, let down by a flawed bicycle: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F8df5fd1b58f128cff27f57f0f959e169 https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Then I ran it on high, medium and low thinking levels (oddly minimal is no longer an option, which WAS an option for 3.5 and 3.6) and got a pretty excellent pelican for the first two: https://tools.simonwillison.net/markdown-svg-renderer.html#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6779a22d5e7bb6bdf29936f1600a5259 https://tools.simonwillison.net/markdown-svg-renderer.html#u... UPDATE: That was in Safari, but as pointed out in the replies here the pelicans do NOT render well in Firefox or Chrome! Best guess is that's because of this invalid filter in the SVG: <filter id="shadow" x="-10%" y="-10%" width="130%" height="130%"></filter> Filters are meant to contain additional elements, not be empty: https://drafts.csswg.org/filter-effects/#FilterElement https://drafts.csswg.org/filter-effects/#FilterElement - so maybe Chrome and Firefox remove the element that references the broken filter but Safari doesn't?
- spiderfarmer 2mo agoIt's not weird if you're in marketing.
- jakswa 2mo agoThis pelican gave me a good laugh, because there's enough reasoning that the render is out of sight initially. The buildup!
- abtinf 2mo ago> got a pretty excellent pelican for the first two This suggests you primarily use Safari. While the bike renders, the pelican doesn’t in Chrome and Firefox. Probably one of the more serious defects I’ve seen with the pelican. It’s one thing when animated SVGs have bugs, but another when plain ones do.
- 2mo ago
- IFC_LLC 2mo agoLike, I understand everything, but by this time I don't give anything about any of those announcements. Theoretically there is some difference between Fable and Opus or Grok and GPT, but at the end of the day I'd look at the bottom left of my screen and to my amusement find out that for the past 3-4 hours I've been using model ______. If the results are semi-decent, I'd keep it on, if not - I'd randomly switch the model and try again. Actual thing that would affect my selection would be a number of unused tokens I have left for a model ____ for this week. Maybe it's cause I'm using those for programming and log parsing and all of them are decent enough, but other than that - there are no leaps I see.
- ls_stats 2mo agoI don't get it, Google could heavily subsidy their Gemini models to make it more attractive, but they prefer to not do it. I don't know one soul who is using Gemini models to code. Even OpenAI who doesn't have money or capacity is offering their Luna model at $1.2 per 1M/out.
- film42 2mo agoWhat you're not seeing are the subsidized Google Cloud startup credits, which includes Gemini. If you're in that program, you choose Gemini because it's essentially "free" and consistent.
- Joeri 2mo agoWhy would they? Unless they have lots of unused tpu real estate that they could host it on “for free” they would be bumping more profitable workloads off of machines to give away that capacity to people with zero long term loyalty. There is no business reason for google to subsidize these models. OpenAI has too much money. They’re spending their money in stupid ways.
- AuthAuth 2mo agoHaving people use the model generates real world training data which could be useful. Other than that I think you're right that there is little reason to artificially boost users with subsidized pricing.
- sumedh 2mo ago> OpenAI has too much money I think they got some data center deals very cheap when no on was thinking about data centers. Dario didnt want to take that risk so Anthropic didnt make the deals earlier but now paying Google and Elon higher rates.
- u1hcw9nx 2mo agoIt makes sense. All these models are money losing businesses. As a business Google might want to focus on fundamental research 2-3 years from now and not compete on who acquires more money losing customers. Just stay little behind and invest money better.
- andai 2mo agoSo their "Flash" model won't be cheap. Are they gonna make a new one that's cheaper? Gemini-3.8-Silverlight? ;)
- andrewstuart 2mo agoGemini has lost the race to be relevant for AI coding.
- dwa3592 2mo agoI was going to cancel my gemini membership today ..... still going ahead. In my experience, gemini 3.1 pro, 3.5, 3.6 flash constantly lie too much about completing their tasks whereas sol (even though equally dumb) never claims something has been done when it hasn't been.
- seunosewa 2mo agoTwo thoughts: 1) It makes sense to try 3.7 flash before cancelling. 2) Prompting models to be honest is surprisingly effective in my recent experience. But only if they listen to instructions.
- dudeinhawaii 2mo agoThat's a good point and something I encountered yesterday. On a multi-agent task, Gemini was the only model that got near the end, ran tests, saw it had issues, took a screenshot, saw the issues, and then said, "I'll mark it complete" and delivered. I was paying for Ultra, then downgraded to Pro and at this point I'm near abandoning it. Agy as a harness also has a tendency to constantly request significant elevations to perform routine operations.
- cmbuck 2mo agoI was pretty sure that the consumer AI Pro, AI Ultra, etc subscriptions have no relevance or impact to Antigravity quota, nor the policies on using data for model training. Do you have information to indicate otherwise?
- smeltworks 2mo ago[flagged]
- throwaw12 2mo agoIs this the reason why Jeff Dean, Sanjay Ghemawat and other DeepMing, Gemini people got kicked out of Google? If so, now I understand why they didn't want to release this model
- bunkydoo 2mo ago[dead]
- sid_talks 2mo agoThe Gemini Flash models makes perfect sense to me coming from a company like Google. Google AI Mode for search is a product I really find useful. It makes sense that Google focused on smaller, faster yet smart enough models that wouldn't break your bank on inference. It plays well into their product ecosystem. Google AI Mode consistently gets me consistently good results and good speeds. It really changes what "googling" is for me.
- alex1138 2mo agoI like Gemini (I'm just a dumb person without knowledge of 'benchmarks' or how x compares to y) while understanding that shoving it into Google auto summaries has been a bad idea and produces inaccurate results
- gazebo2 2mo agoI similarly have a weird affinity for Gemini that I can't really articulate. I used Gemini's free chat and found it great for exploring technical topics (and random one-off general walking-around-questions) and appreciated its speed, tone and accuracy. I spent a month playing with Gemini CLI / Antigravity and found it also an effective coding agent, at least for my workflow (entirely in the loop development and review). I also was really surprised that I could just paste it images of a project I was working on and have it immediately understand what it was looking at -- which I've come to learn is considered a unique strong point for Gemini. I've been playing with GPT5.6 for about a month and it's definitely powerful but I honestly think I'll go back to Gemini. There's something kind of charming about working with an AI that not only is particularly good at web search and information gathering, but also one that doesn't feel like some superhuman overengineering freak when it comes to code.
- dannyw 2mo agoI like to think of Gemini as a broader generalist that hasn't allocated _most_ of its skillpoints to agentic coding execution :)
- abixb 2mo agoWell, Google is probably the only one among frontier model providers in the US that doesn't have a massive financial pressure to deliver business results ASAP (and this focus on agent coding and long-running agentic tasks), so Google is able to focus more on encoding deep scientific, cultural and historical knowledge to its systems. I'd assume DeepMind's focus on the scientific core also played a role in the tone and approach Gemini models take for explanations and Q&As. I find that GPT models and Claude tend to talk in strong slangs and in-group jargon, but love Gemini's massive general knowledge corpus — reminds me of Richard Feynman from his lectures.
- fryanyway_swe 2mo agoI just use web chat as "harness"(lol) or interface and I have mostly switched to Gemini as the free limit basically never run out for me unlike ChatGPT and Claude. Also impressed with Grok for some stuff.
- sarjann 2mo agoIntroductory price seems a bit weird as it expires at the end of the year and by then it’s going to be significantly outdated.
- nipunaeka89 2mo agoAlways loved the Gemini. Helped my work alot
- theplumber 2mo agoSo they keep pushing these Flash models because they don’t really have a powerful model…or better said their ‘pro’ model is actually a flash
- kyruzic 2mo agoSo at this point new models seem to only care about one task, software development. This really was not the original pitch of ai and I do not see how it justifies the insane spend or valuations it has produced.
- Revanche1367 2mo agoThink of the potential layoffs of highly paid employees! But, I think it’s also based on what they are being used for, most LLM users are still mainly SWEs or similar as I understand and there’s a ton of data to train them for coding.
- kyruzic 2mo agoI agree its what they are being used for and their primary revenue source. My point is mainly that was never the pitch that got ai the hype it did and imo doesn't justify the valuations even if we all lose our jobs to ai. Because it no longer seems like they even think its making other jobs go away.
- pjmlp 2mo agoThey do it via other ways, translation teams and asset creators for CMSes, I no longer see them in our projects, the builtin AI tooling takes care of it. Same with those that use to improve marketing outcomes for SEO and such, now there are AI based reports with automatic improvements. Finally on other domains you already have robots on supermarkets, fast food, and gas stations, where the customer does the work of the (now gone) employees without any kind of price reduction.
- ur-whale 2mo agoWhy does Google keep announcing these subpar models? What am I missing? They can't seem to be able to produce a frontier model, fine. Just be quiet about it and work hard until you manage to put one together. [EDIT]: Come to think of it. Maybe they're trying to build the Toyota corolla of AI ... let's see if that wins them the battle long term. I personally doubt it.
- ksajadi 2mo agoi just wish google cloud ux was remotely as good as their models. they made some progress with their studio, but then in a true google fashion, product names keep changing (Gemini, Anti-gravity, Vertex, Google AI,...) as well as confusion and complexity for something as simple as registering agy cli with a Google cloud project. today i wanted to link agy to a google cloud project, for that i had to enable 5 different APIs in google cloud UI, then create a subscription for Gemini Enterprise (whatever that is), then link it to a project, then assign it to a user. and after all that, agy couldn't find the subscription. the best part: i couldn't cancel the subscription. so i just paid $35 for one month and left it.
- pests 2mo agoGemini is the model family. Antigravity is their agentic coding app / IDE. There is two products, one is chat-only the other is more standard IDE. Google AI Studio is a consumer/developer playground with a in-browser IDE meant for prototypes or demos, there is a gallery of demos etc. Easily shared, easy key access. Vertex is the AI offering from the GCP side of the company, that is going to target more enterprise or business solutions (scaling, data governance, security, production deployents, etc)
- sunaookami 2mo agoNot in the website/app? Only on AI Studio? Weird.
- JakeSc 2mo agoOn a related note, I see all these quantitative benchmarks and the models getting really good at them over time. One thing I've been wondering: if the GPT series of models performs so well quantitatively, why do I still kind of hate using them relative to Claude? There’s a missing “vibes” or “taste” benchmark I think.
- hateboxaa 2mo agostatuory
- hateboxaa 2mo agoI think this is a slop release
- Frannky 2mo agoHave you tried the new DeepSeek Pro v4, Qwen 3.8, Gemini 3.7 Flash, and Grok 4.6? Do they make sense for any use cases? I'm currently using omp with Kimi K3 as the planner and DeepSeek v4 Flash 0731 as the implementer, or CC + Fable for planning and Opus 4.8 for implementation. For API(not coding), I just use DeepSeek v4 flash 0731 and MiMo. I'm pretty happy where I am, but I'm wondering if these new models provide some new kind of advantage
- mirekrusin 2mo agoJust use all of them and synthesize final result, record scoring when synthesizing. Later you can make decision to drop low performing ones.
- Frannky 2mo agoIt's about time and tokens. I find it more effective to get the vibe from friends, HN, Discord, Reddit, and then play with only the most promising one. I skipped all of these because they seemed not worth it
- barrenko 2mo agoSynthetize=?
- mirekrusin 2mo agoie. open router fusion [0] [0] https://openrouter.ai/openrouter/fusion https://openrouter.ai/openrouter/fusion
- Alien1Being 2mo agoGoogle has got the IBM disease. Large lumbering enterprise with massive inertia. Where innovators leave as soon as they get a better offer. None of the authors of the seminal "Attention Is All You Need" paper are still at Google. Fast forward a decade and Google will be reduced to hiring the kind of mediocrities who deign to work at IBM and Accenture.
- luckjack47 2mo agoI don’t think the skill set required to write the Vaswani paper is the same as training and shipping frontier models like Gemini so I am not sure why people keep bringing this up.
- Alien1Being 2mo agoTraining, tweaking and shipping frontier models like Gemini is something that a team of mediocre engineers can do. Google continuing to ship generations of Gemini using its mediocre teams is an existence proof of that For genuine advances one needs the Hintons and Vasvanis, not yet another bunch of mediocre Kookaid drinking product managers at Apple and Google
- ezekiel68 2mo agoAnd I'm over here on SiliconFlow using Stepfun AI Step-3.5-Flash at $0.10/M input and $0.30/M output tokens (262K context window) for complex market analysis work in rust utilizing vectorized instruction sets. It provides me with amazing results. I honestly wonder how long this calliope can keep playing before it crashes to the ground. (I have no business relationship to anything mentioned here except as a regular retail customer who went bargain-hunting)
- Jr23_xd 2mo ago[dead]
- nojito 2mo agoThe flash models are just so good at OCR tasks and summarization. Glad to see them constantly improving them.
- customguy 2mo agoWas excited to try it, since I've been use 3.6 Flash in the last few days to make simply experiments/prototypes. My loop is writing a prompt, maybe adding a screenshot of the closest to what I want I have so far, then based on the result I modify/extend the prompt, maybe use another screenshot. Well, I think I'll stick to 3.6 for now. Based in a very scientific sample size of exactly one attempt each: https://imgur.com/a/fDOkBDm https://imgur.com/a/fDOkBDm Both got the same prompt and example screenshot. I mean, they both suck but that's normal this early in, but the 3.6 version (first screenshot) actually changes the displayed threads depending on selected categories, and the messages of whatever selected thread, as obviously described in the prompt. You might say it more or less does what it should. Both versions have an ugly flash/jerk in the category pane when selecting/deselecting a category, so that's a wash. The 3.7 version doesn't work at all, i.e. it always shows all threads, and no messages for any of them. I can post new messages in threads but they don't show up, and it doesn't even increase the message counter for the thread. I guess it's a matter of taste but I don't like the look either, while 3.6 actually is in the spirit of the screenshot I added the prompt, using 11k of CSS versus 3.7's 16k. The code is also less, and the backend split into 3 files (instead of just 2 as 3.7 did it), so assuming it sucks in either case, it'll be easier to read and massage. edit: geez, 3.6 even properly fades/disables the "new message" button when no thread is selected, 3.7 didn't bother which is smart since everything else is broken anyway. Maybe it's better at really complex things, but for simple things, what I'm experimenting with, I already saw enough.
- qudat 2mo agoI just subbed to Gemini a week ago and have been using antigravity and 3.6 flash. The speed is absolutely a differentiator compared to Claude.
- onlyrealcuzzo 2mo agoClaude is terrible when it comes to speed. Flash is great, but Codex models are also fast, as is DeepSeek v4 Flash. Anyone who's Anthropic-pilled should really get out and explore and see how unbelievably terrible they are when it comes to speed and cost vs quality. Anthropic has good models, they're just way too expensive and slow for what you pay for.
- v3ss0n 2mo agoGEMMA 4.5 gogoogogo
- correlator 2mo agoWe run a platform that serves models from many of the frontier players to power conversations and workflows. Gemini flash-2.5 was a game changer for us when it came out. Cheap, fast, and reliable. We are now considering dropping support for the model family all together. All of their models require significant scrubbing of errant thinking blocks, inner monologues, and it's consuming more engineering resources than it's worth.
- linzhangrun 2mo agoFeels like Gemini Pro will arrive directly as Gemini 4 Pro
- llm_nerd 2mo agoIf you're in the Gemini app and want to try it out as a Pro or Ultra user...well you can't. It's only in the weird "Spark" agent that demands you entire Google identity.
- nicolamanzini 2mo agoIt is also doing pretty well in threejseval. Frontier there for the price. Much better than 3.6. https://threejseval.com/ranking https://threejseval.com/ranking
- mchusma 2mo agoThis is a solid release (at the intro pricing, the other pricing is dumb). I do think its a missed opportunity to really blow things out of the water and have this be another 1/2 off, but clearly they don't have the inference efficiency for it. Speed is good, the knowledge in google's models is solid for those usecases, price is reasonable (after 3.5/3.6 major missteps). The intelligence index vs cost pareto frontier is crazy now, its basically a flat line with 9 people all at or right at the edge of the frontier along various parts of the graphs. Insanely competitive right now.
- aurareturn 2mo agoProbably golden age for competition before consolidation.
- mchusma 2mo agoMaybe? But GLM just put itself on the pareto today again, so its up to 10. The theory was that there would be recursive self improvement in models, leading to a 1 or a handful of entities running away from the rest. But basically the opposite happened? Different teams, different hardware stacks. What do we make of this? (1) There must not really be any deep secrets/moats right now? (2) improving models is not something you can throw only intelligence at now?
- goochgibbler 2mo agoI started vibe coding with Gemini. At first very exciting, we got a rough program up quickly. And then, code corruption. Over and over. I started to document every single file, every single step, writing explicit rules to not fake data and create for real world use, but everyday I kept catching Gemini errors, which turned into flat out lies. Generating fake test data instead of pulling real data. Making up test answers. Saying features were implemented that weren't. It even coded fake python files that printed made up results. I think it realized I wasn't doing code reviews. Every day was spent chasing defects and rolling back. After Gemini admitted to faking 7 tests i let Claude review the code, and it fixed it almost immediately asking why half the features were broken or missing. Well Claude, because Gemini flash did the least amount of work to make me happy. Impressive. Very human. Very frustrating.
- prtmnth 2mo agoI am curious to understand who is this model targeted at?
- GaggiX 2mo agoIt's a cheap and powerful model. Also very fast for this class of models.
- exacube 2mo agoI'm finding Gemini 3.7 speaks a lot more academically. Its explanations are not clear and intuitive by default. Maybe this has to do with all the RLVR it went through, where reasoning through difficult academic/coding problems caused it to think and speak a certain way. The benchmarks looks great but it doesn't feel as legible, so maybe it's more meant to be an agentic model rather than an everyday model whose outputs are read by humans?
- mintflow 2mo agoGiven my codex have limited usage and I get a idea to build a on device clipboard translator with menubar when read a article find some words not seens before So i launched agy and find seems it have a 3.7 flash i thought its latest,until i see this article i know it's just released after serveral rounds prompt(the initial prd prompt is by chatgpt), it use 2.9k user message token and 42.5k reponse tokens with gemini 3.7 flash(low) after i check the status, i got a working on device translator, its cool the model seems also p and retty fast and the generated app looks good and easy to use but agy cli is a bit unintuitve and i also check the guy that release the latest agy, seems the guy does not commit much ? or perhaps agy does not open source and only use github as a issue feedback channel
- raincole 2mo agoIt's quite good. It seems to be "Google's turn" again. Might last two or three months?
- ddp26 2mo agoWhat are we to infer from no release of gemini-3.5-pro, but frequent releases of smaller flash models (presumably from the same large pre-training run?)
- resters 2mo agoGoogle is in direct meetings with Scott Bessent and Howard Lutnick and is prudently keeping dangerous frontier models out of the hands of customers until reasonable precautions can be taken.
- akulbe 2mo agoWith the DeepMind guy leaving the building... isn't Gemini going to get old and crusty, and fall into disrepair? I'm mostly serious here. Aside from Gmail and search, it feels Google doesn't have a good track record for maintaining things. It really makes me wonder... if key people are leaving, what's going to happen? The Google graveyard is pretty big.
- m3kw9 2mo agoThey need to learn to release on antigravity or cli sub plans on day/two one like OpenAI
- ipsum2 2mo agoRecently my android phone updated from Google assistant to Gemini Flash. Completely unusable. Asking it to play music and it refuses, hallucinating instructions to connect Spotify to Gemini. The instructions say to tap buttons that don't exist. Bonus feature from Gemini: a toggle to opt back into Google Assistant, but it doesn't work. Still stuck with Gemini.
- booi 2mo agoGemini hallucinations will continue until morale improves
- UlisesAC4 2mo agoYeah I don't know who was the genius who approved changing an Alexa lookalike to a chatgpt clone for our phones.
- ipsum2 2mo agoIt would be fine if it worked. I wish I could replace Gemini on my phone with chatGPT.
- seabrookmx 2mo agoIt's weird because Gemini via Antigravity or Google Search "AI Mode" works great. It's just the Google Assistant replacement application that's terrible. I've completely given up trying to ask it anything via Android auto. I punch my directions manually in Google maps before I leave now because the voice assistant is unusable.
- dpacmittal 2mo agoFunny you mention this. I ask gemini what time it is, and it replies in a different language. I've checked all language settings, and they appear to be correct. Yet gemini just refuses to reply in correct language, and it only happens when I'm asking it for time.
- monolog 2mo ago[dead]
- rw2 2mo agoIt's just useless to build any model in this range because deepseek flash is so fast and so cheap. I see no reason for anyone to not use it for a few points in performance. Only models that matter are the edge fable class models people use for code, and Google struggled with that.
- jnwatson 2mo agoRates are going up Aug 16.
- doesntmatter123 2mo ago[dead]
- instagraham 2mo agoFunnily enough, I just ran a task on AI Studio with 3.6 yesterday and got 3.7 to do a similar one today; so it serves as an interesting and quick comparisons between the old and the new (usually, if enough time passes between your use of one model and the next, you'll have a sourer view of it than its actual competence suggests). It hallucinated in both cases despite being given an API key and building a lot of pipes to access data using this. It was a simple "oh shit" fix moment for the model, but weird how eager it was to hallucinate despite the process being designed for it to be data-driven. We should move past the idea that benchmarks alone tell us whether a model is getting better. I would've had the same experience a year or two ago with 1.5, and the solution would've been similar (keep prompting). I've been investing time into making system prompts and input prompts more meticulous, but the fundamental "it will make shit up" problem still remains, even though it shouldn't when the job involves calling tools. I know this sounds like I'm expecting superpowers of it (I'm not), but my point is just that these incremental benchmark gains may not reflect user experience.
- customguy 2mo agoFor me, it just shits the bed: https://news.ycombinator.com/item?id=49292924 https://news.ycombinator.com/item?id=49292924 Reading the google blog and these discussions makes me feel like I'm taking crazy pills, seriously. Side-by-side comparisons with the exact same inputs or it didn't happen, that's my rule going forward. Test all the things, believe nothing.
- alastairr 2mo agoI couldn't get past the first chart which basically showed that intelligence and cost are both better for gpt luna, I'm not sure what the argument is here for Gemini flash 3.7 given that comparison.
- VeejayRampay 2mo agonot what the first chart shows
- alastairr 2mo agoYou are right, it's the deepswe chart vs cost I was referring to which is slightly further down
- barrenko 2mo agoI am still using gemini flash 2.5 for an old app I have running that does some OCR as well if necessary, is it time to switch up?
- manapause 2mo agoI also use it for a chatbot assistant for customer onboarding. Gemini deserves some kudos for their willingness to allow entry level subscriptions access to API-keys to build solutions with.
- rawoke083600 2mo agoI wish they will always put it front and centre what are all the models name as how i call it in via their API.
- tibzejoker 2mo agook but still a small model.. why is google so behind regarding SOTA :/ daim
- ChildOfChaos 2mo agoStill fairly poor limits in anti gravity despite the price cut, which seems to be API only.
- libertas_quae_s 2mo ago[dead]
- acrush 2mo agoWhen will the pro be out...
- Imanari 2mo agoTesting it in Pi.dev and liking the speed a lot! Huge improvements from past gemini models regarding tool calls and agentic capabilities
- or1gaminal 2mo agoexperimenting w/ flash 3.7 on morphology and it looks promising. google models looks capable on linguistic side and prose. are there any reliable benchmarks on this?
- hnicrcjk6o 2mo agoThis hits different after a long week
- throwaway_20357 2mo agoHard to believe nowadays but Gemini Flash 3 was $0.25 / 1M input and $1.50 / 1M output when it was released.
- chaostheory 2mo agoAll of Google’s models prioritize speed over quality and correctness. That may be great for search but not so good for hard work.
- TraX22 2mo ago[flagged]
- michaelteter 2mo agoI have had such good, consistent success with Gemini 3.6 Flash (usually on medium setting) that I forget about it. It is just this reliable, effective assistant that I use for managing and organizing information, brainstorming, and of course coding. I honestly don't know what I would want more/better than what I was already getting... so I'm curious if I will see any improvements with 3.7.
- maxglute 2mo agoNo free weekly quota reset this time unlike 3.6, anyone know when that happens?
- low_tech_punk 2mo agoMaybe Google is adding compute infra. I noticed today that the API has 500 tok/s throughput. It's quite compelling now.