35 ms·
Gemini 3.1 Pro
Preview: https://console.cloud.google.com/vertex-ai/publishers/google/model-garden/gemini-3.1-pro-preview?pli=1 https://console.cloud.google.com/vertex-ai/publishers/google...
Card: https://deepmind.google/models/model-cards/gemini-3-1-pro/ https://deepmind.google/models/model-cards/gemini-3-1-pro/
- rohithavale3108 8mo ago[flagged]
- Topfi 8mo agoAppears the only difference to 3.0 Pro Preview is Medium reasoning. Model naming has long gone from even trying to make sense, but considering 3.0 is still in preview itself, increasing the number for such a minor change is not a move in the right direction.
- GrayShade 8mo agoMaybe that's the only API-visible change, saying nothing about the actual capabilities of the model?
- argsnd 8mo agoI disagree. Incrementing the minor number makes so much more sense than “gemini-3-pro-preview-1902” or something.
- xnx 8mo ago> increasing the number for such a minor change is not a move in the right direction A .1 model number increase seems reasonable for more than doubling ARC-AGI 2 score and increasing so many other benchmarks. What would you have named it?
- Topfi 8mo agoMy issue is that we haven't even gotten the release version of 3.0, that is also still in Preview, so may stick with 3.0 till that has been deemed stable. Basically, what does the word "Preview" mean, if newer releases happen before a Preview model is stable? In prior Google models, Preview meant that there'd still be updates and improvements to said model prior to full deployment, something we saw with 2.5. Now, there is no meaning or reason for this designation to exist if they forgo a 3.0 still in Preview for model improvements.
- xnx 8mo agoGiven the pace AI is improving and that it doesn't give the exact same answers under many circumstances, is the the [in]stability of "preview" a concern? GMail was in "beta" for 5 years.
- verdverm 8mo agoChatGPT 4.5 was never released to the public, but it is widely believed to be the foundation the 5.x series is built on. Wonder how GP feels about the minor bumps for other model providers?
- Topfi 8mo agoMinor version bumps are good and I want model providers to communicate changes. The issue I am having is that Gemini "preview" class models have different deprecation timelines and rate limits, making them impossible to rely on for professional use cases. That's why I'd prefer they finish the 3.0 role out prior to putting resources into deploying a second "preview" class model. For a stable deployment, Google needs a sufficient amount of hardware to guarantee inference and having two Pro models running makes that even more challenging: https://ai.google.dev/gemini-api/docs/models https://ai.google.dev/gemini-api/docs/models
- verdverm 8mo agoSorry, but you come off as an armchair devops saying things like this. Google is fine, they know more than anyone else about how to run Ai at scale. "preview" != GA, sounds like you need to adjust your expectations
- jannyfer 8mo agoAccording to the blog post, it should be also great at drawing pelicans riding a bicycle.
- clhodapp 8mo agoThere's a very short blog post up: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/ https://blog.google/innovation-and-ai/models-and-research/ge...
- sigmar 8mo agoblog post is up- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/ https://blog.google/innovation-and-ai/models-and-research/ge... edit: biggest benchmark changes from 3 pro: arc-agi-2 score went from 31.1% -> 77.1% apex-agents score went from 18.4% -> 33.5%
- sho_hn 8mo agoThe touted SVG improvements make me excited for animated pelicans.
- takoid 8mo agoI just gave it a shot and this is what I got: https://codepen.io/takoid/pen/wBWLOKj https://codepen.io/takoid/pen/wBWLOKj The model thought for over 5 minutes to produce this. It's not quite photorealistic (some parts are definitely "off"), but this is definitely a significant leap in complexity.
- makeavish 8mo agoLooks great!
- onionisafruit 8mo agoGood to see it wearing a helmet. Their safety team must be on their game.
- BrokenCogs 8mo agoYes but why would a pelican need a helmet? If it falls over it can just fly away... Common sense 1 Gemini 0
- throwa356262 8mo agoObviously these domestic pelicans can't fly, otherwise why would they need a bike?
- WarmWash 8mo agoIt seems google is having a disjointed roll out, and there will likely be an official announcement in a few hours. Apparently 3.1 showed up unannounced in vertex at 2am or something equally odd. Either way early user tests look promising.
- mark_l_watson 8mo agoFine, I guess. The only commercial API I use to any great extent is gemini-3-flash-preview: cheap, fast, great for tool use and with agentic libraries. The 3.1-pro-preview is great, I suppose, for people who need it. Off topic, but I like to run small models on my own hardware, and some small models are now very good for tool use and with agentic libraries - it just takes a little more work to get good results.
- nurettin 8mo agoI like to ask claude how to prompt smaller models for the given task. With one prompt it was able to make a low quantized model call multiple functions via json.
- throwaway2027 8mo agoSeconded. Gemini used to be trash and I used Claude and Codex a lot but gemini-3-flash-preview punches above it's weight, it's decent and I rarely if ever run into any token limit either.
- verdverm 8mo agoThirded, I've been using gemini-3-flash to great effect. Anytime I have something more complicated, I give it to pro & flash to see what happens. Coin flip if flash is nearly equivalent (too many moving vars to be analytical at this point)
- PlatoIsADisease 8mo agoWhat models are you running locally? Just curious. I am mostly restricted to 7-9B. I still like ancient early llama because its pretty unrestricted without having to use an abliteration.
- mark_l_watson 8mo agoI experimented with many models on my 16G and 32G Macs. For less memory, qwen3:4b is good, for the 32B Mac, gpt-oss:20b is good. I like the smaller Mistral models like mistral:v0.3 and rnj-1:latest is a pretty good small reasoning model.
- maxloh 8mo agoGemini 3 seems to have a much smaller token output limit than 2.5. I used to use Gemini to restructure essays into an LLM-style format to improve readability, but the Gemini 3 release was a huge step back for that particular use case. Even when the model is explicitly instructed to pause due to insufficient tokens rather than generating an incomplete response, it still truncates the source text too aggressively, losing vital context and meaning in the restructuring process. I hope the 3.1 release includes a much larger output limit.
- esafak 8mo agoPeople did find Gemini very talkative so it might be a response to that.
- jayd16 8mo ago> Even when the model is explicitly instructed to pause due to insufficient tokens Is there actually a chance it has the introspection to do anything with this request?
- otabdeveloper4 8mo agoNo.
- maxloh 8mo agoYeah, it does. It was possible with 2.5 Flash. Here's a similar result with Qwen Qwen3.5-397B-A17B: https://chat.qwen.ai/s/530becb7-e16b-41ee-8621-af83994599ce?fev=0.2.7 https://chat.qwen.ai/s/530becb7-e16b-41ee-8621-af83994599ce?...
- jayd16 8mo agoOk it prints some stuff at the end but does it actually count the output tokens? That part was already built in somehow? Is it just retrying until it has enough space to add the footer?
- verdverm 8mo agoNo, the model doesn't have purview into this afaik I'm not even sure what "pausing" means in this context and why it would help when there are insufficient tokens. They should just stop when you reach the limit, default or manually specified, but it's typically a cutoff. You can see what happens by setting output token limit much lower
- PunchTornado 8mo agoThe biggest increase is LiveCodeBench Pro: 2887. The rest are in line with Opus 4.6 or slightly better or slightly worse.
- shmoogy 8mo agobut is it still terrible at tool calls in actual agentic flows?
- esafak 8mo agoHas anyone noticed that models are dropping ever faster, with pressure on companies to make incremental releases to claim the pole position, yet making strides on benchmarks? This is what recursive self-improvement with human support looks like.
- PlatoIsADisease 8mo agoOnly using my historical experience and not Gemini 3.1 Pro, I think we see benchmark chasing then a grand release of a model that gets press attention... Then a few days later, the model/settings are degraded to save money. Then this gets repeated until the last day before the release of the new model. If we are benchmaxing this works well because its only being tested early on during the life cycle. By middle of the cycle, people are testing other models. By the end, people are not testing them, and if they did it would barely shake the last months of data.
- KoolKat23 8mo agoI have a relatively consistent task that it completed with new information on weekdays at the edge of its intelligence. Interestingly 3.0 flash was good when it came out, took a nose dive a month back and is now excellent, I actually can't fault it it's so good. It's performance in antigravity has also actually improved since launch day where it was giving non-stop typescript errors (not sure if that was antigravity itself).
- emp17344 8mo agoRemember when ARC 1 was basically solved, and then ARC 2 (which is even easier for humans) came out, and all of the sudden the same models that were doing well on ARC 1 couldn’t even get 5% on ARC 2? Not convinced these benchmark improvements aren’t data leakage.
- casey2 8mo agoARC 2 was made specifically to artificially lower contemporary LLM scores, therefore any kind of model improvements will have outsized effects Also people use "saturated" too liberally. The top left corner 1 cent per task is saturated IMO. Since there are billions of people who would perfer to solve arc 1 tasks at 52 cents per task. Arc 2 a human would make thousands of dollars a day with 99.99% accuracy
- dude250711 8mo agoI hereby allow you to release models not at the same time as your competitors.
- sigmar 8mo agoIt is super interesting that this is the same thing that happened in November (ie all labs shipping around the same week 11/12-11/23).
- zozbot234 8mo agoThey're just throwing a big Chinese New Year celebration.
- vintermann 8mo agoCould that actually be connected? There are a LOT of Chinese engineers and researchers working on all these models, I assume they would like to take some vacation days, and it makes sense to me to time releases around it.
- matrix2596 8mo agoGemini 3.1 Pro is based on Gemini 3 Pro
- skerit 8mo agoLol, and this line: > Geminin 3.1 Pro can comprehend vast datasets Someone was in a hurry to get this out the door.
- __jl__ 8mo agoAnother preview release. Does that mean the recommended model by Google for production is 2.5 Flash and Pro? Not talking about what people are actually doing but the google recommendation. Kind of crazy if that is the case
- qingcharles 8mo agoI've been playing with the 3.1 Deep Think version of this for the last couple of weeks and it was a big step up for coding over 3.0 (which I already found very good). It's only February...
- minimaxir 8mo agoPrice is unchanged from Gemini 3 Pro: $2/M input, $12/M output. https://ai.google.dev/gemini-api/docs/pricing https://ai.google.dev/gemini-api/docs/pricing Knowledge cutoff is unchanged at Jan 2025. Gemini 3.1 Pro supports "medium" thinking where Gemini 3 did not: https://ai.google.dev/gemini-api/docs/gemini-3 https://ai.google.dev/gemini-api/docs/gemini-3 Compare to Opus 4.6's $5/M input, $25/M output. If Gemini 3.1 Pro does indeed have similar performance, the price difference is notable.
- deleted 8mo ago[deleted]
- plaidfuji 8mo agoSounds like the update is mostly system prompt + changes to orchestration / tool use around the core model, if the knowledge cutoff is unchanged
- sigmar 8mo agoknowledge cutoff staying the same likely means they didn't do a new pre-train. We already knew there were plans from deepmind to integrate new RL changes in the post training of the weights. https://x.com/ankesh_anand/status/2002017859443233017 https://x.com/ankesh_anand/status/2002017859443233017
- brokencode 8mo agoThis keeps getting repeated for all kinds of model releases, but isn’t necessarily true. It’s possible to make all kinds of changes without updating the pretraining data set. You can’t judge a model’s newness based on what it knows about.
- rancar2 8mo agoIf we don't see a huge gain on the long-term horizon thinking reflected with the Vendor-Bench 2, I'm not going to switch away from CC. Until Google can beat Anthropic on that front, Claude Code paired with the top long-horizon models will continue to pull away with full stack optimizations at every layer.
- denysvitali 8mo agoWhere is Simon's pelican?
- saberience 8mo agoPlease no, let's not.
- codethief 8mo agoNot Simon's but here is one: https://news.ycombinator.com/item?id=47075709 https://news.ycombinator.com/item?id=47075709
- denysvitali 8mo agoThank you!
- Mashimo 8mo agoIt's also quite impressive with SVG animations. > Create an SVG animation of a Beaver sitting next to a recordplayer and a create of records, his eyes follows the mouse curser. https://gemini.google.com/share/717be5f9b184 https://gemini.google.com/share/717be5f9b184
- msavara 8mo agoSomehow doesn't work for me :) "An internal error has occurred"
- ChrisArchitect 8mo agoBlog post: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/ https://blog.google/innovation-and-ai/models-and-research/ge...
- saberience 8mo agoI always try Gemini models when they get updated with their flashy new benchmark scores, but always end up using Claude and Codex again... I get the impression that Google is focusing on benchmarks but without assessing whether the models are actually improving in practical use-cases. I.e. they are benchmaxing Gemini is "in theory" smart, but in practice is much, much worse than Claude and Codex.
- skerit 8mo agoI'm glad someone else is finally saying this, I've been mentioning this left and right and sometimes I feel like I'm going crazy that not more people are noticing it. Gemini can go off the rails SUPER easily. It just devolves into a gigantic mess at the smallest sign of trouble. For the past few weeks, I've also been using XML-like tags in my prompts more often. Sometimes preferring to share previous conversations with `<user>` and `<assistant>` tags. Opus/Sonnet handles this just fine, but Gemini has a mental breakdown. It'll just start talking to itself. Even in totally out-of-the-ordinary sessions, it goes crazy. After a while, it'll start saying it's going to do something, and then it pretends like it's done that thing, all in the same turn. A turn that never ends. Eventually it just starts spouting repetitive nonsense. And you would think this is just because the bigger the context grows, the worse models tend to get. But no! This can happen well below even the 200.000 token mark.
- reilly3000 8mo agoFlash is (was?) was better than Pro on these fronts.
- user34283 8mo agoI exclusively use Gemini for Chat nowadays, and it's been great mostly. It's fast, it's good, and the app works reliably now. On top of that I got it for free with my Pixel phone. For development I tend to use Antigravity with Sonnet 4.5, or Gemini Flash if it's about a GUI change in React. The layout and design of Gemini has been superior to Claude models in my opinion, at least at the time. Flash also works significantly faster. And all of it is essentially free for now. I can even select Opus 4.6 in Antigravity, but I did not yet give it a try.
- the_duke 8mo agoGemini 3 is pretty good, even Flash is very smart for certain things, and fast! BUT it is not good at all at tool calling and agentic workflows, especially compared to the recent two mini-generations of models (Codex 5.2/5.3, the last two versions of Anthropic models), and also fell behind a bit in reasoning. I hope they manage to improve things on that front, because then Flash would be great for many tasks.
- verdverm 8mo agoThese improvements are one of the things specifically called out on the submitted page
- chermi 8mo agoYou can really notice the tool use problems. They gotta get on that. The agent trend seems real, and powerful. They can't afford to fall behind on it.
- verdverm 8mo agoI don't really have tool usage issues that I don't put under that doesn't follow system prompt instructions consistently there are these times where it puts a prefix on all function calls, which is weird and I think hallucination, so maybe that one 3.1 hopefully fixes that
- HardCodedBias 8mo ago"They can't afford to fall behind on it." They are very, very seriously far behind as of 3.0. We'll see if 3.1 addresses the issue at all.
- spwa4 8mo agoIn other words: they just need to motivate their employees while giving in to finance's demands to fire a few thousand every month or so ... And don't forget, it's not just direct motivation. You can make yourself indispensable by sabotaging or at least not contributing to your colleagues' efforts. Not helping anyone, by the way, is exactly what your managers want you to do. They will decide what happens, thank you very much, and doing anything outside of your org ... well there's a name for that, isn't there? Betrayal, or perhaps death penalty.
- zhyder 8mo agoSurprisingly big jump in ARC-AGI-2 from 31% to 77%, guess there's some RLHF focused on the benchmark given it was previously far behind the competition and is now ahead. Apart from that, the usual predictable gains in coding. Still is a great sweet-spot for performance, speed and cost. Need to hack Claude Code to use their agentic logic+prompts but use Gemini models. I wish Google also updated Flash-lite to 3.0+, would like to use that for the Explore subagent (which Claude Code uses Haiku for). These subagents seem to be Claude Code's strength over Gemini CLI, which still has them only in experimental mode and doesn't have read-only ones like Explore.
- WarmWash 8mo ago>I wish Google also updated Flash-lite to 3.0+ I hope every day that they have made gains on their diffusion model. As a sub agent it would be insane, as it's compute light and cranks 1000+ tk/s
- zhyder 8mo agoAgree, can't wait for updates to the diffusion model. Could be useful for planning too, given its tendency to think big picture first. Even if it's just an additional subagent to double-check with an "off the top off your head" or "don't think, share first thought" type of question. More generally would like to see how sequencing autoregressive thinking with diffusion over multiple steps might help with better overall thinking.
- topocite 8mo agoThe only thing I can notice is deep research is better. Like much closer to outputting a paper from arxiv straight away. I am really the bottleneck now and what to do with all this new information.
- vinhnx 8mo agoModel card https://deepmind.google/models/model-cards/gemini-3-1-pro/ https://deepmind.google/models/model-cards/gemini-3-1-pro/
- josalhor 8mo agoI speculated that 3 pro was 3.1... I guess I was wrong. Super impressive numbers here. Good job Google.
- refulgentis 8mo ago> I speculated that 3 pro was 3.1 ?
- josalhor 8mo agoSorry... I speculated that 3 deep think is 3.1 pro.. model names are confusing..
- makeavish 8mo agoI hope to have great next two weeks before it gets nerfed.
- unsupp0rted 8mo agoI've found Google (at least in AI Studio) are the only provider NOT to nerf their models after a few weeks
- makeavish 8mo agoI don't use AI studio for my work. I used Antigravity/Gemini CLI and 3 pro was great for few weeks and now it's worse than 3 flash or any smaller model from competitor which are rated lower on benchmarks
- scrlk 8mo agoIME, they definitely nerf models. gemini-2.5-pro-exp-03-25 through AI Studio was amazing at release and steadily degraded. The quality started tanking around the time they hid CoT.
- makeavish 8mo agoGreat model until it gets nerfed. I wish they had a higher paid tier to use non nerfed model.
- xnx 8mo agoWhat are you talking about?
- spyckie2 8mo agoI think there is a pattern it will always be nerfed the few weeks before launching a new model. Probably because they are throwing a bunch of compute at the new model.
- makeavish 8mo agoYeah maybe that but atleast let us know about this Or have dynamic limits? Nerfing breaks trust. Though I am not sure if they actually nerf it intentionally. Haven't heard from any credible source. I did experience in my workflow though.
- Mond_ 8mo agoBad news, John Google told me they already quantized it immediately after the benchmarks were done and it sucks now. I miss when Gemini 3.1 was good. :(
- mustaphah 8mo agoGoogle is terrible at marketing, but this feels like a big step forward. As per the announcement, Gemini 3.1 Pro score 68.5% on Terminal-Bench 2.0, which makes it the top performer on the Terminus 2 harness [1]. That harness is a "neutral agent scaffold," built by researchers at Terminal-Bench to compare different LLMs in the same standardized setup (same tools, prompts, etc.). It's also taken top model place on both the Intelligence Index & Coding Index of Artificial Analysis [2], but on their Agentic Index, it's still lagging behind Opus 4.6, GLM-5, Sonnet 4.6, and GPT-5.2. --- [1] https://www.tbench.ai/leaderboard/terminal-bench/2.0?agents=Terminus+2 https://www.tbench.ai/leaderboard/terminal-bench/2.0?agents=... [2] https://artificialanalysis.ai https://artificialanalysis.ai
- saberience 8mo agoBenchmarks aren't everything. Gemini consistently has the best benchmarks but the worst actual real-world results. Every time they announce the best benchmarks I try again at using their tools and products and each time I immediately go back to Claude and Codex models because Google is just so terrible at building actual products. They are good at research and benchmaxxing, but the day to day usage of the products and tools is horrible. Try using Google Antigravity and you will not make it an hour before switching back to Codex or Claude Code, it's so incredibly shitty.
- gregorygoc 8mo agoWhat’s so shitty about it?
- mustaphah 8mo agoThat's been my experience too; can't disagree. Still, when it comes to tasks that require deep intelligence (esp. mathematical reasoning [1]), Gemini has consistently been the best. [1] https://arxiv.org/abs/2602.10177 https://arxiv.org/abs/2602.10177
- naiv 8mo agook , so they are scared that 5.3 (pro) will be released today/tomorrow and blow it out of the water and rushed it while they could still reference 5.2 benchmarks.
- PunchTornado 8mo agoI don't think models blow other models anymore. We have the big 3 which are neck to neck in most benchmarks and the rest. I doubt that 5.3 will blow the others.
- scld 8mo agoeasy now
- nickandbro 8mo agoDoes well on SVGs outside of "pelican riding on a bicycle" test. Like this prompt: "create a svg of a unicorn playing xbox" https://www.svgviewer.dev/s/NeKACuHj https://www.svgviewer.dev/s/NeKACuHj Still some tweaks to the final result, but I am guessing with the ARC-AGI benchmark jumping so much, the model's visual abilities are allowing it to do this well.
- simonw 8mo agoInteresting how it went a bit more 3D with the style of that one compared to the pelican I got.
- andy12_ 8mo agoI'm thinking now that as models get better and better at generating SVGs, there could be a point where we can use them to just make arbitrary UIs and interactive media with raw SVGs in realtime (like flash games).
- nickandbro 8mo agoOr quite literally a game where SVG assets are generated on the fly using this model
- kridsdale3 8mo agoThats one dimension before another long term milestone: Realtime generation of 3D mesh content during gameplay. Which is the "left brain" approach vs the "right brain" approach of coming at dynamic videogames from the diffusion model direction which the Gemini Genie thing seems to be about.
- rafark 8mo ago> there could be a point where we can use them to just make arbitrary UIs and interactive media with raw SVGs So render ui elements using xml-like code in a web browser? You’re not going to believe me when I tell you this…
- simonw 8mo agoPretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.
- bredren 8mo agoWhat is that, a snack in the basket?
- WarmWash 8mo agoA fish for the road
- sigmar 8mo ago"integrating a bicycle basket, complete with a fish for the pelican... also ensuring the basket is on top of the bike, and that the fish is correctly positioned with its head up... basket is orange, with a fish inside for fun." how thoughtful of the ai to include a snack. truly a "thanks for all the fish"
- defen 8mo agoA pelican already has an integrated snack-holder, though. It wouldn't need to put it in the basket.
- SauntSolaire 8mo agoThat one's full too
- troymc 8mo agoThe number of snacks in the basket is a random variable with a Poisson distribution.
- xnx 8mo agoNot even animated? This is 2026.
- dxbednarczyk 8mo agoEvery time I've used Gemini models for anything besides code or agentic work they lean so far into the RLHF induced bold lettering and bullet point list barf that everything they output reads as if the model was talking _at_ me and not _with_ me. In my Openclaw experiment(s) and in the Gemini web UI, I've specifically added instructions to avoid this type of behavior, but it only seemed to obey those rules when I reminded the model of them. For conversational contexts, I don't think the (in some cases significantly) better benchmark results compared to a model like Sonnet 4.6 can convince me to switch to Gemini 3.1. Has anyone else had a similar experience, or is this just a me issue?
- InkCanon 8mo agoI think they all output that bold lettering, point by point style output. I strongly suspect it's part of a synthetic data pipeline all these AI companies have, and it improves performance. Claude seems to be the least of them, but it will start writing code at the drop of a hat. What annoys me in Gemini is that it has a really strange tendency to come up with weird analogies, especially in Pro mode. You'll be asking it about something like red black trees and it'll say "Red Black Trees (The F1 of Tree Data Structures)".
- hydrolox 8mo agoYes, the analogy habit is the most annoying of all. Overall formatting for me is doable, if it didn't divide up an answer into these silly arbitrary categories with useless analogies. I've tried adding in my user preferences to never use analogies but it inevitably falls back into that habit.
- markab21 8mo agoYou just articulated why I struggle to personally connect with Gemini. It feels so unrelatable and exhausting to read its output. I prefer to read Opus/Deepseek/GLM over Gemini, Qwen and the open source GPT models. Maybe it is RLHF that is creating my distaste from using it. (I pay for Gemini; I should be using it more... but the outputs just bug me and feel more work to get actionable insight.)
- quacky_batak 8mo agoI’m keen to know how and where are you using Gemini. Anthropic is clearly targeted to developers and OpenAI is general go to AI model. Who are the target demographic for Gemini models? ik that they are good and Flash is super impressive. but i’m curious
- minimaxir 8mo agoGemini has an obvious edge over its competitors in one specific area: Google Search. The other LLMs do have a Web Search tool but none of them are as effective.
- jdc0589 8mo agoI use it as my main platform right now both for work/swe stuff, and person stuff. It works pretty well, they have the full suite of tools I want from general LLM chat, to notebookLM, to antigravity. My main use-cases outside of SWE generally involve the ability to compare detailed product specs and come up with answers/comparisons/etc... Gemini does really well for that, probably because of the deeper google search index integration. Also I got a year of pro for free with my phone....so thats a big part.
- jug 8mo agoI personally use it as my general purpose and coding model. It's good enough for my coding tasks most of the time, has very good and rapid web search grounding that makes the Google index almost feel like part of its training set, and Google has a family sharing plan with individual quotas for Google AI Pro at $20/month for 5 users which also includes 2 TB in the cloud. Family sharing is a unique feature for Gemini 3 Flash Thinking (300 prompts per day and user) & Pro (100 prompts per day and user).
- deleted 8mo ago[deleted]
- hunta2097 8mo agoI use the Gemini web interface just as I would ChatGPT. They also have coding environment analogues of Claude-Code in Anti-gravity and Gemini-CLI. When you sign up for the pro tier you also get 2TB of storage, Gemini for workspace and Nest Camera history. If you're in the Google sphere it offers good value for money.
- eric15342335 8mo agoMy first impression is that the model sounds slightly more human and a little more praising. Still comparing the ability.
- hsaliak 8mo agoThe eventual nerfing gives me pause. Flash is awesome. What we really want is gemini-3.1-flash :)
- mixel 8mo agoGoogle seems to really pull ahead in this AI race. For me personally they offer the best deal and although the software is not quiet there compared to openai or anthropic (in regards to 1. web GUI, 2. agent-cli). I hope they can fix that in the future and I think once Gemini 4 or whatever launches we will see a huge leap again
- eknkc 8mo agoI hope they fail. I honestly do not wish Google to have the best model out there and be forced to use their incomprehensible subscription / billing / project management whatever shit ever again. I don’t know what their stuff cost. I don’t know why would I use vertex or ai studio. What is included in my subscription what is billed per use. I pray that whatever they build fails and burns.
- dybber 8mo agoEventually the models will be generally be so good that the competition moves from the best model to the best user experience and here I think we can expect others will win, e.g. Microsoft with GitHub and VS Code
- eknkc 8mo agoThat's my hope but Google has unlimited cash to throw at model development and can basically burn more cash can openai and anthropic combined. Might tip the scale in the long run.
- otherme123 8mo agoThey all suck. OpenAI ignores scanning limits and disabled routes in robots.txt, after a 429 "Too Many Requests" they retry the same url half a dozen of times from different IPs in the next couple of minutes, and they once DoS'ed my small VPS trying to do a full scan of sitemaps.xml in less than one hour, trying and retrying if any endpoint failed. Google and others at least respects both robots.txt and 429s. They invested years scanning all the internet, so they can now train on what they have stored in their server. OpenAI seems to assume that MY resources are theirs.
- Robdel12 8mo agoI really want to use google’s models but they have the classic Google product problem that we all like to complain about. I am legit scared to login and use Gemini CLI because the last time I thought I was using my “free” account allowance via Google workspace. Ended up spending $10 before realizing it was API billing and the UI was so hard to figure out I gave up. I’m sure I can spend 20-40 more mins to sort this out, but ugh, I don’t want to. With alllll that said.. is Gemini 3.1 more agentic now? That’s usually where it failed. Very smart and capable models, but hard to apply them? Just me?
- alpineman 8mo ago100% agreed. I wish someone would make a test for how reliably the LLMs follow tool use instructions etc. The pelicans are nice but not useful for me to judge how well a model will slot into a production stack.
- embedding-shape 8mo agoAt first when I got started with using LLMs I read/analyzed benchmarks, looked at what example prompts people used and so on, but many times, a new model does best at the benchmark, and you think it'll be better, but then in real work, it completely drops the ball. Since then I've stopped even reading benchmarks, I don't care an iota about them, they always seem more misdirected than helpful. Today I have my own private benchmarks, with tests I run myself, with private test cases I refuse to share publicly. These have been built up during the last 1/1.5 years, whenever I find something that my current model struggles with, then it becomes a new test case to include in the benchmark. Nowadays it's as easy as `just bench $provider $model` and it runs my benchmarks against it, and I get a score that actually reflects what I use the models for, and it feels like it more or less matches with actually using the models. I recommend people who use LLMs for serious work to try the same approach, and stop relying on public benchmarks that (seemingly) are all gamed by now.
- cdelsolar 8mo agoshare
- cmrdporcupine 8mo agoDoesn't show as available in gemini CLI for me. I have one of those "AI Pro" packages, but don't see it. Typical for Google, completely unclear how to actually use their stuff.
- zokier 8mo ago> Last week, we released a major update to Gemini 3 Deep Think to solve modern challenges across science, research and engineering. Today, we’re releasing the upgraded core intelligence that makes those breakthroughs possible: Gemini 3.1 Pro. So this is same but not same as Gemini 3 Deep Think? Keeping track of these different releases is getting pretty ridiculous.
- LZ_Khan 8mo agobiggest problem is that it's slow. also safety seems overtuned at the moment. getting some really silly refusals. everything else is pretty good.
- sergiotapia 8mo agoTo use in OpenCode, you can update the models it has: opencode models --refresh Then /models and choose Gemini 3.1 Pro You can use the model through OpenCode Zen right away and avoid that Google UI craziness. --- It is quite pricey! Good speed and nailed all my tasks so far. For example: @app-api/app/controllers/api/availability_controller.rb @.claude/skills/healthie/SKILL.md Find Alex's id, and add him to the block list, leave a comment that he has churned and left the company. we can't disable him properly on the Healthie EMR for now so this dumb block will be added as a quick fix. Result was: 29,392 tokens $0.27 spent So relatively small task, hitting an API, using one of my skills, but a quarter. Pricey!
- gbalduzzi 8mo agoI don't see it even after refresh. Are you using the opencode-gemini-auth plugin as well?
- sergiotapia 8mo agoNo I am not just vanilla OpenCode. I do have OpenCode Zen credits, and I did opencode login whatever their command is to auth against opencode itself. Maybe that's the reason I see these premium models.
- markerbrod 8mo agoBlogpost: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/ https://blog.google/innovation-and-ai/models-and-research/ge...
- jeffbee 8mo agoRelatedly, Gemini chat seems to be if not down then extremely slow. ETA: They apparently wiped out everyone's chats (including mine). "Our engineering team has identified a background process that was causing the missing user conversation metadata and has successfully stopped the process to prevent further impact." El Mao.
- pawelduda 8mo agoIt's safe to assume they'll be releasing improved Gemini Flash soon? The current one is so good & fast I rarely switch to pro anymore
- derac 8mo agoWhen 3 came out they mentioned that flash included many improvements that didn't make it into pro (via an hn comment). I imagine this release includes those.
- tucnak 8mo agoGemini 3 Pro (high) is a joke compared to Gemini 3 Flash in Antigravity, except it's not even funny. Flash is insane value, and super capable, too. I've had it implement a decompiler for very obscure bytecode, and it was passing all tests in no time. PITA to refactor later, but not insurmountable. Gemini 3 Pro (high) choked on this problem in the early stages... I'm looking forward to comparing 3.1 Pro vs 3.0 Flash, hopefully they have improved on it enough to finally switch over.
- techgnosis 8mo agoI'd love a new Gemini agent that isn't written with Node.js. Not sure why they think that's a good distribution model.
- CamperBob2 8mo ago(Shrug) Ask it to write one!
- onlyrealcuzzo 8mo agoWe've gone from yearly releases to quarterly releases. If the pace of releases continues to accelerate - by mid 2027 or 2028 we're headed to weekly releases.
- rubicon33 8mo agoBut actual progress seems to be slower. These modes are releasing more often but aren’t big leaps.
- wahnfrieden 8mo agoGPT 5.3 (/Codex) was a huge leap over 5.2 for coding
- rubicon33 8mo agoEh, sure, but marginally better if not the same as Claude 4.6, which itself was a small bump over Claud 4.5
- gallerdude 8mo agoWe used to get one annual release which was 2x as good, now we get quarterly releases which are 25% better. So annually, we’re now at 2.4x better.
- minimaxir 8mo agoDue to the increasing difficulty of scaling up training, it appears the gains are instead being achieved through better model training which appears to be working well for everyone.
- deleted 8mo ago[deleted]
- jcims 8mo agoPelican on a bicycle in drawio - https://imgur.com/a/tNgITTR https://imgur.com/a/tNgITTR (FWIW I'm finding a lot of utility in LLMs doing diagrams in tools like drawio)
- pqdbr 8mo agoHow are you prompting it to draw diagrams in drawio
- ac29 8mo agoDrawio drawings are just XML, its possible it can generate that directly
- riku_iki 8mo agohopefully op will answer if that's what he is doing
- jcims 8mo agoSometimes it helps to also provide a drawio file that has the elements you wan't (eg. cloud service icons or whatever), but you just feed it the content you want diagrammed and let it eat. Even if it's not completely correct, it usually creates something that's much closer to complete than a blank page.
- jcims 8mo agoHere's the chat I used for the drawing - https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221ijMj6rxiXDcPfgE-NjGWqQDR05gGn8tY%22%5D,%22action%22:%22open%22,%22userId%22:%22114347212038551092903%22,%22resourceKeys%22:%7B%7D%7D&usp=sharing https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%... Save the xml, import to drawio
- impulser_ 8mo agoSeems like they actually fixed some of the problems with the model. Hallucinations rate seems to be much better. Seems like they also tuned the reasoning maybe that were they got most of the improvements from.
- whynotminot 8mo agoThe hallucination rate with the Gemini family has always been my problem with them. Over the last year they’ve made a lot of progress catching the Gemini models up to/near the frontier in general capability and intelligence, but they still felt very late 2024 in terms of hallucination rate. Which made the Gemini models untrustworthy for anything remotely serious, at least in my eyes. If they’ve fixed this or at least significantly improved, that would be a big deal.
- SubiculumCode 8mo agoMaybe I haven't kept up with how ghatgpt and claude are doing , but 6 monthlatelys ago or so, I thought Gemini was leading on that front.
- ArmandoAP 8mo agoModel Card https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
- seizethecheese 8mo agoI use Gemini flash lite in a side project, and it’s stuck on 2.5. It’s now well behind schedule. Any speculation as to what’s going on?
- foruhar 8mo agoGemini-3.0-flash-preview came out right away with the 3.0 release and I was expecting 3.0-flash-lite before a bump on the pro model. I wonder if they have abandoned that part of the Pareto/price-performance.
- janalsncm 8mo agoThis model says it accepts video inputs. I asked it to transcribe a 5 second video of a digital water curtain which spelled “Boo Happy Halloween”, and it came back with “Happy” which wasn’t the first frame, but also is incomplete. This kind of test is good because it requires stitching together info from the whole video.
- nautilus12 8mo agoOk, why don't you work on getting 3.0 out of preview first? 10 min response time is pretty heinous
- mucai82 8mo agoI agree, according to Googles terms you are not allowed to use the preview model for production use cases. And 3.0 has been in preview for a loooong time now :(
- davidguetta 8mo agoImplementation and Sustainability Hardware: Gemini 3 Pro was trained using Google’s Tensor Processing Units (TPUs). TPUs are specically designed to handle the massive computations involved in training LLMs and can speed up training considerably compared to CPUs. TPUs often come with large amounts of high-bandwidth memory, allowing for the handling of large models and batch sizes during training, which can lead to better model quality. TPU Pods (large clusters of TPUs) also provide a scalable solution for handling the growing complexity of large foundation models. Training can be distributed across multiple TPU devices for faster and more efficient processing. So google doesn't use NVIDIA GPUs at all ?
- PunchTornado 8mo agono. only tpus
- paride5745 8mo agoAnother reason to use Gemini then. Less impact on gamers…
- TiredOfLife 8mo agoTPUs still use ram and chip production capacity
- dekhn 8mo agoWhen I worked there, there was a mix of training on nvidia GPUs (especially for sparse problems when TPUs weren't as capable), CPUs, and TPUs. I've been gone for a few years but I've heard a few anecdotal statements that some of their researchers have to use nvidia GPUs because the TPUs are busy.
- lejalv 8mo agoBla bla bla yada sustainability yada often come with large better growing faster... It's such an uninformative piece of marketing crap
- rjh29 8mo agoI assume that's a Gemini LLM response? You can tell Gemini is bullshitting when it starts using "often" or "usually" - like in this case "TPUs often come with large amounts of memory". Either they did or they didn't. "This (particular) mall often has a Starbucks" was one I encountered recently.
- tenpoundhammer 8mo agoIn an attempt to get outside of benchmark gaming I had it make Platypus on a Tricycle. It's not as good as pelican on bicycle. https://www.svgviewer.dev/s/BiRht5hX https://www.svgviewer.dev/s/BiRht5hX
- dinosor 8mo agoFor a moment I assumed the output would look like Perry the Platipus from the Disney (I think?) show. It's suprising to me (as a layman) that a show with lots of media that would've made it to the training corpus didn't show up.
- 0_____0 8mo agothat's better than i thought it would be
- hyperbovine 8mo agowould love to be able to teleport this thread to, oh, 5 years ago. people would think some sort of alien technology had landed.
- textlapse 8mo agoTo really confuse it, ask it to take that tricycle with the platypus on it to a car wash.
- azuanrb 8mo agoThe CLI needs work, or they should officially allow third-party harnesses. Right now, the CLI experience is noticeably behind other SOTA models. It actually works much better when paired with Opencode. But with accounts reportedly being banned over ToS issues, similar to Claude Code, it feels risky to rely on it in a serious workflow.
- 1024core 8mo agoIt's been hugged to death. I keep getting "Something went wrong".
- boxingdog 8mo ago[dead]
- spankalee 8mo agoI hope this works better than 3.0 Pro I'm a former Googler and know some people near the team, so I mildly root for them to at least do well, but Gemini is consistently the most frustrating model I've used for development. It's stunningly good at reasoning, design, and generating the raw code, but it just falls over a lot when actually trying to get things done, especially compared to Claude Opus. Within VS Code Copilot Claude will have a good mix of thinking streams and responses to the user. Gemini will almost completely use thinking tokens, and then just do something but not tell you what it did. If you don't look at the thinking tokens you can't tell what happened, but the thinking token stream is crap. It's all "I'm now completely immersed in the problem...". Gemini also frequently gets twisted around, stuck in loops, and unable to make forward progress. It's bad at using tools and tries to edit files in weird ways instead of using the provided text editing tools. In Copilot it, won't stop and ask clarifying questions, though in Gemini CLI it will. So I've tried to adopt a plan-in-Gemini, execute-in-Claude approach, but while I'm doing that I might as well just stay in Claude. The experience is just so much better. For as much as I hear Google's pulling ahead, Anthropic seems to be to me, from a practical POV. I hope Googlers on Gemini are actually trying these things out in real projects, not just one-shotting a game and calling it a win.
- knollimar 8mo agoIs the thinking token stream obfuscated? Im fully immersed
- orbital-decay 8mo agoIt's just a summary generated by a really tiny model. I guess it also an ad-hoc way to obfuscate it, yes. In particular they're hiding prompt injections they're dynamically adding sometimes. Actual CoT is hidden and entirely different from that summary. It's not very useful for you as a user, though (neither is the summary).
- ukuina 8mo agoAgree the raw thought-stream is not useful. It's likely filled with "Aha!" and "But wait!" statements.
- timabdulla 8mo agoGoogle tends to trumpet preview models that aren't actually production-grade. For instance, both 3 Pro and Flash suffer from looping and tool-calling issues. I would love for them to eliminate these issues because just touting benchmark scores isn't enough.
- mijoharas 8mo agoGemini 3 is still in preview (limited rate limits) and 2.5 is deprecated (still live but won't be for long).[0] Are Google planning to put any of their models into production any time soon? Also somewhat funny that some models are deprecated without a suggested alternative(gemini-2.5-flash-lite). Do they suggest people switch to Claude? [0] https://ai.google.dev/gemini-api/docs/deprecations https://ai.google.dev/gemini-api/docs/deprecations
- vidarh 8mo agoThis feels very Google
- drbacon 8mo agoI found the Googler!
- vidarh 8mo agoNope. The closest I've gotten was rejecting Google recruiters several times. But like everyone else I'm used to Google failing to care about products.
- cmrdporcupine 8mo agoInside Google we just constantly joked/complained about "old thing is deprecated, new isn't ready yet" This held for internal APIs, facilities, systems more even than it did for the outside world. Which is terrible.
- deleted 8mo ago[deleted]
- johnwheeler 8mo agoI know Google has anti-gravity but do they have anything like Claude code as far as user interface terminal basically TUI?
- alooPotato 8mo agohttps://github.com/google-gemini/gemini-cli https://github.com/google-gemini/gemini-cli
- johnwheeler 8mo agoThankS!!
- Murfalo 8mo agoI like to think that all these pelican riding a bicycle comments are unwittingly iteratively creating the optimal cyclist pelican as these comment threads are inevitably incorporated in every training set.
- alpineman 8mo agoMore like half of Google's AI team is hanging out on HN, and they can optimise for that outcome to get a good rep among the dev community.
- Barbing 8mo agoSee: fish in bike front basket
- kridsdale3 8mo agoHello. (I'm not aware of anyone doing this, but GDM is quite info-siloed these days, so my lack of knowledge is not evidence it's not happening)
- alpineman 8mo agoHello. Please push internally for more reliable tool use across Gemini models. Intelligence is useless if it can't be applied :)
- throwaw12 8mo agoCan we switch from Claude Code to Google yet? Benchmarks are saying: just try But real world could be different
- foruhar 8mo agoMy sense is that the Gemini models are very capable but the Gemini CLI experience is subpar compared to Claude Code and Codex. I'm guess that it's the harness but since it can get confused, fall into doom loops, and generally lose the plot in a way that the model does not in Gemini Studio or the Gemini app. I think a bunch of these harnesses are open source so it surprises me that there can be such a gulf between them.
- cmrdporcupine 8mo agoIt's not just the tooling. If you use Gemini in opencode it malfunctions in similar ways. I haven't tried 3.1 yet, but 3 is just incompetent at tool use. In particular in editing chunks of text in files, it gets very confused and goes into loops. The model also does this thing where it degrades into loops of nonsense thought patterns over time. For shorter sessions where it's more analysis than execution, it is a strong model. We'll see about 3.1. I don't know why it's not showing in my gemini CLI as available yet.
- dana321 8mo agoIts not just subpar, its not even sub-sub-par. It goes into loops and never completes a task 8 times out of 10 that i've used it.
- WarmWash 8mo ago3.1 Pro is the first model to correctly count the number of legs on my "five legged dog" test image. 3.0 flash was the previous best, getting it after a few prompts of poking. 3.1 got it on the first prompt though, with the prompt being "How many legs does the dog have? Count Carefully". However, it didn't get it on the first try with the original prompt (prompt: "How many legs does the dog have?"). It initially said 4, then with a follow up prompt got it to hesitantly say 5, with one limb must being obfuscated or hidden. So maybe I'll give it a 90%? This is without tools as well.
- merlindru 8mo agoyour question may have become part of the training data with how much coverage there was around it. perhaps you should devise a new test :P
- gallerdude 8mo agoMy job may have become part of the training data with how much coverage there is around it. Perhaps another career would be a better test of LLM capabilities.
- suddenlybananas 8mo agoHave you ever heard of a black swan?
- wat10000 8mo agoEasy fix, make a new test image with six legs, and watch all the LLMs say it has five.
- WarmWash 8mo agoHonestly at this point I have fed this image in so many times on so many models, that it also functions as a test for "Are they training on my image specifically" (they are generally, for sure, but that's along with everything else in the ocean of info people dump in). I genuinely don't think they are. GPT-5.2 still stands by 4 legs, and OAI has been getting this image consistently for over a year. And 3.1 still fumbled with the harder prompt "How many legs does the dog have?". I needed to add the "count carefully" part to tip it off that something was amiss. Since it did well I'll make some other "extremely far out of the norm" images to see how it fairs. A spider with 10 legs or a fish with two side fins.
- 1024core 8mo agoIt got the car wash question perfectly: You are definitely going to have to drive it there—unless you want to put it in neutral and push! While 200 feet is a very short and easy walk, if you walk over there without your car, you won't have anything to wash once you arrive. The car needs to make the trip with you so it can get the soap and water. Since it's basically right next door, it'll be the shortest drive of your life. Start it up, roll on over, and get it sparkling clean. Would you like me to check the local weather forecast to make sure it's not going to rain right after you wash it?
- suddenlybananas 8mo agoThey probably had time to toss that example in the training soup.
- AlphaAndOmega0 8mo agoPrevious models from competitors usually got that correct, and the reasoning versions almost always did. This kind of reflexive criticism isn't helpful, it's closer to a fully generalized counter-argument against LLM progress, whereas it's obvious to anyone that models today can do things they couldn't do six months ago, let alone 2 years back.
- suddenlybananas 8mo agoI'm not denying any progress, I'm saying that reasoning failures that are simple which have gone viral are exactly the kind of thing that they will toss in the training data. Why wouldn't they? There's real reputational risks in not fixing it and no costs in fixing it.
- AlphaAndOmega0 8mo agoGiven that Gemini 3 Pro already did solid on that test, what exactly did they improve? Why would they bother? I double checked and tested on AI Studio, since you can still access the previous model there: >You should drive. >If you walk there, your car will stay behind, and you won't be able to wash it. Thinking models consistently get it correct and did when the test was brand new (like a week or two ago). It is the opposite of surprising that a new thinking model continues getting it correct, unless the competitors had a time machine.
- pRusya 8mo agoI'm using gemini.google.com/app with AI Pro subscription. "Something went wrong" in FF, works in Chrome. Below is one of my test prompts that previous Gemini models were failing. 3.1 Pro did a decent job this time. > use c++, sdl3. use SDL_AppInit, SDL_AppEvent, SDL_AppIterate callback functions. use SDL_main instead of the default main function. make a basic hello world app.
- himata4113 8mo agoThe visual capabilities of this model are frankly kind of ridicioulus what the hell.
- trilogic 8mo agoHumanity last exam 44%, Scicode 59, and that 80, and this 78 but not 100% ever. Would be nice to see that this models, Plus, Pro, Super, God mode can do 1 Bench 100%. I am missing smth here?
- xrd 8mo agoThese models are so powerful. It's totally possible to build entire software products in the fraction of the time it took before. But, reading the comments here, the behaviors from one version to another point version (not major version mind you) seem very divergent. It feels like we are now able to manage incredibly smart engineers for a month at the price of a good sushi dinner. But it also feels like you have to be diligent about adopting new models (even same family and just point version updates) because they operate totally differently regardless of your prompt and agent files. Imagine managing a team of software developers where every month it was an entirely new team with radically different personalities, career experiences and guiding principles. It would be chaos. I suspect that older models will be deprecated quickly and unexpectedly, or, worse yet, will be swapped out with subtle different behavioral characteristics without notice. It'll be quicksand.
- seizethecheese 8mo ago> It feels like we are now able to manage incredibly smart engineers for a month at the price of a good sushi dinner. In my experience it’s more like idiot savant engineers. Still remarkable.
- cm2012 8mo agoIts like getting access to an amazing engineer, but you get a new individual engineer each prompt, not one consistent mind.
- jama211 8mo agoYeah I keep maintaining a specific app I built with gpt 5.1 codex max with that exact model because it continues to work for the requests I send it, and attempts with other models even 5.2 or 5.3 codex seemed to have odd results. If I were superstitious I would say it’s almost like the model that wrote the code likes to work on the code better. Perhaps there’s something about the structure it created though that it finds easier to understand…
- worldsavior 8mo agoSushy dinner? What are you building with AI, a calculator?
- upmind 8mo agoIn my experience, while Gemini does really well in benchmarks I find it much worse when I actually use the model. It's too verbose / doesn't follow instructions very well. Let's see if that changes with this model.
- veselin 8mo agoI am actually going to complain about this: that neither of the Gemini models are not preview ones. Anthropic seems the best in this. Everything is in the API on day one. OpenAI tend to want to ask you for subscription, but the API gets there a week or a few later. Now, Gemini 3 is not for production use and this is already the previous iteration. So, does Google even intent to release this model?
- ChrisArchitect 8mo agoMore discussion: https://news.ycombinator.com/item?id=47075318 https://news.ycombinator.com/item?id=47075318
- syspec 8mo agoDoes anyone know if this is in GA immediately or if it is in preview? On our end, Gemini 3.0 Preview was very flakey (not model quality, but as in the API responses sometimes errored out), making it unreliable. Does this mean that 3.0 is now GA at least?
- panarchy 8mo agoI had it make a simple HTML/JS canvas game (think flappy bird) and while it did some things mildly better (and others noticeably worse) it still fell into the exact same traps as earlier models. It also had a lot of issues generating valid JS at parts and asking it what the code should be just made it endlessly generate the same exact incorrect code.
- lysecret 8mo agoPlease I need 3 in ga…
- solarisos 8mo agoThe speed of these 3.1 and Preview releases is starting to feel like the early days of web frameworks. It’s becoming less about the raw benchmarks and more about which model handles long-context 'hallucination' well enough to be actually used in a production pipeline without constant babysitting.
- yuvalmer 8mo agoGemini 3.0 Pro is bad model for its class. I really hope 3.1 is a leap forward.
- leecommamichael 8mo agoWhoa, I think Gemini 3 Pro was a disappointment, but Gemini 3.1 Pro is definitely the future!
- XCSme 8mo agoGets 10/10 on my potato benchmarks: https://aibenchy.com/model/google-gemini-3-1-pro-preview-medium https://aibenchy.com/model/google-gemini-3-1-pro-preview-med...
- XCSme 8mo agoNow I need to write more tests. It's a bit hard to trick reasoning models, because they explore a lot of the angles of a problem, and they might accidentally have an "a-ha" moment that leads them on the right path. It's a bit like doing random sampling and stumbling upon the right result after doing gradient descent from those points.
- thevinter 8mo agoAre you intentionally keeping the benchmarks private?
- XCSme 8mo agoYes. I am trying to think what's the best way to give most information about how the AI models fail, without revealing information that can help them overfit on those specific tests. I am planning to add some extra LLM calls, to summarize the failure reason, without revealing the test.
- XCSme 8mo agoAdded one more test, which surprisingly gemini flash 3 reasoning passes, but gemini 3.1 pro not
- 0xcb0 8mo agoI'm trying to find the information, is this available on the Gemini CLI script, or is this just the web front-end where I can use this new model?
- BMFXX 8mo agoJust wish iI could get 2.5 daily limit above 1000 requests easily. Driving me insane...
- fdefitte 8mo ago[dead]
- pickle-pixel 8mo agodoes it still crash out after couple prompts?
- Filip_portive 8mo ago[flagged]
- robviren 8mo agoI have run into a surprising number of basic syntax errors on this one. At least in the few runs I have tried it's a swing and a miss. Wonder if the pressure of the Claude release is pushing these stop gap releases.
- jeffybefffy519 8mo agoSomeone needs to make an actual good benchmark for LLM's that matches real world expectations, theres more to benchmarks than accuracy against a dataset.
- robotpepi 8mo agothis reminds me of that joke of someone saying "it's crazy that we have ten different standards for doing this", and then there're 11 standards
- knollimar 8mo agoXkcd 927
- casey2 8mo agoWe don't need real world benchmarks, if they were good for real world tasks people would use them We need scientific benchmarks that tease out the nature of intelligence. There are plenty of unsaturated benchmarks. Solving chess using "mostly" language modeling is still an open problem. And beyond that creating a machine that can explain why that move is likely optimal at some depth. AI that can predict the output of another AI.
- kuprel 8mo agoWhy don't they show Grok benchmarks?
- andxor 8mo agoThey've fallen way behind.
- kuprel 8mo agoGPT 5.2 loses at everything but they included that
- andxor 8mo agoWho are they supposed to compare it to? I'm not sure what makes you think that Grok is even remotely comparable to the frontier models right now.
- rudhdb773b 8mo agoGrok has been and still is the best at incorporating search. 4.20 with its 4 agents puts it back at the top for reasoning as well. As soon as it's added to the API, the benchmarks should show that.
- andxor 8mo agoI agree it's good for researching current events because of the integration with X.
- vnglst 8mo agoI asked Gemini 3.1 Pro to generate some of the modern artworks in my "Pelican Art Gallery". I particularly like the rendition of the Sunflowers: https://pelican.koenvangilst.nl/gallery/category/modern https://pelican.koenvangilst.nl/gallery/category/modern
- agentifysh 8mo agoMy enthusiasm is a bit muted this cycle because I've been burned by Gemini CLI. These models are very capable but Gemini CLI just doesn't seem to be able to work for one it never follows instructions strictly like its competitors do, and it hallucinates even which is a rarity. More importantly feels like Google is stretched thin across different Gemini products and pricing reflects this, I still have no idea how to pay for Gemini CLI, in codex/claude its very simple $20/month for entry and $200/month for ton of weekly usage. I hope whoever is reading this from Google they can redeem Gemini CLI by focusing on being competitive instead of making it look pretty (that seems to be the impression I got from the updates on X)
- cheema33 8mo ago> I still have no idea how to pay for Gemini CLI, in codex/claude its very simple $20/month for entry and $200/month for ton of weekly usage. This! I would like to sign up for a paid plan for Gemini CLI. But I have not been able to figure out how. I already have Codex and Claude plans. Those were super easy to sign up for.
- jiggawatts 8mo agoWhat’s your difficulty? Google has published easy to follow 27-step instructions for how to sign up for the half a dozen services you need to chain together to enable this common usecase!
- knollimar 8mo agoOn the 3.0 rollout I signed up for billing and it just silently failed. Solution was to remake billing account and then wait a day
- jiggawatts 8mo ago“Time machine not included. Some temporal slippage may be experienced. Paradoxes will not be compensated.”
- 8mo ago
- mbh159 8mo ago77.1% on ARC-AGI-2 and still can't stop adding drive-by refactors. ARC-AGI-2 tests novel pattern induction, it's genuinely hard to fake and the improvement is real. But it doesn't measure task scoping, instruction adherence, or knowing when to stop. Those are the capabilities practitioners actually need from a coding agent. We have excellent benchmarks for reasoning. We have almost nothing that measures reliability in agentic loops. That gap explains this thread.
- deleted 8mo ago[deleted]
- taytus 8mo agoAnother preview model? Why google keep doing this?
- vnglst 8mo agoI asked Gemini 3.1 Pro Preview to generate the modern artworks as SVG for my Pelican Art Gallery. I particularly like the rendition of the Sunflowers: https://pelican.koenvangilst.nl/gallery/category/modern https://pelican.koenvangilst.nl/gallery/category/modern
- atleastoptimal 8mo agoWriting style wise, 3.1 seems very verbose, but somehow less creative compared to 3.
- jdthedisciple 8mo agoWhy should I be excited?
- getcrunk 8mo agoGemini is so stubborn, and often doesn’t follow explicit and simple instructions. So annoying
- ismailmaj 8mo ago3.1 feels to me like 3.0 but that takes a long time to think, it didn't feel like a leap in raw intelligence like 2.5 pro was.
- sdeiley 8mo agoPeople underrate Google's cost effectiveness so much. Half price of Opus. HALF. Think about ANY other product and what you'd expect from the competition thats half the price. Yet people here act like Gemini is dead weight ____ Update: 3.1 was 40% of the cost to run AA index vs Opus Thinking AND SONNET, beat Opus, and still 30% faster for output speed. https://artificialanalysis.ai/?speed=intelligence-vs-speed&media-leaderboards=image-to-video#intelligence-vs-cost-to-run-artificial-analysis-intelligence-index https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...
- jstummbillig 8mo agoIt's not half price or cost effective if it can't do the job, that I am happy to pay twice the price for to get done. But I agree: If they can get there (at one point in the past year I felt they were the best choice for agentic coding), their pricing is very interesting. I am optimistic that it would not require them to go up to Opus pricing.
- Decabytes 8mo agoAny tips for working with Gemini through its chat interface? I’ve worked with ChatGPT and Claude and I’ve generally found them pleasant to work with, but everytime I use Gemini the output is straight dookie
- londons_explore 8mo agomake sure you use ai studio (not the vertex one), not the consumer gemini interface. Seems to work better for code there.
- briHass 8mo agoEven though I don't like the privacy implications, make sure you use the option to save and use past chats for context. After a few months of back and forth (hundreds of 'chat' sessions), the responses are much higher quality. It sometimes does 'callbacks' to things discussed in past chats, which are typically awkward non-sequiturs, but it does improve it overall. When I play with it in 'temporary chat' mode that ignores past chats and personal context directives, the responses are the typical slop littered with emojis, worthless lists, and platitudes/sycophancy. It's as jarring as turning off your adblocker and seeing the garish ad trash everywhere.
- mrcwinn 8mo agoIt's fascinating to watch this community react to positively to Google model releases and so negatively toward OpenAI's. You all do understand that an ad revenue model is exactly where Google will go, right?
- jeffbee 8mo agoGemini already drives ad revenue. If the conversation goes in that direction it will use product search results with the links attributable to Google.
- sidrag22 8mo agoIt's all so astroturfed so its hard to tell. I got the opposite impression though. Seemed like OpenAI had more fake positivity towards the top that i tried to skim, and this had way less and a lot of complaints. Im biased I dont trust either of them, so perhaps im just hard looking for the hate and attributing all the positive stuff to advertising.
- webtcp 8mo agoAn enemy is better than a traitor
- mrcwinn 8mo agoQuite a low bar. And in any case, isn’t Google already a traitor to its original mission statement?
- zapnuk 8mo agoGemini 3 was: 1. unreliable in GH copilot. Lots of 500 and 4XX errors. Unusable in the first 2 months 2. not available in vertex ai (europe). We have requirements regarding data residency. Funny enough anthropic is on point with releasing their models to vertex ai. We already use opus and sonnet 4.6. I hope google gets their stuff together and understands that not everyone wants/can use their global endpoint. We'd like to try their models.
- hn_throw2025 8mo agoYeah great, now can I have my pinned chats back please? https://www.google.com/appsstatus/dashboard/incidents/nK23ZsSxV33wfkqXVeh3 https://www.google.com/appsstatus/dashboard/incidents/nK23Zs...
- siliconc0w 8mo agoGoogle has a hugely valuable dataset of changes from decades of changes from top tier software engineers but it's so proprietary they can't use it to train their external models.
- sheepscreek 8mo agoIf it’s any consolation, it was able to one-shot a UI & data sync race condition that even Opus 4.6 struggled to fix (across 3 attempts). So far I like how it’s less verbose than its predecessor. Seems to get to the point quicker too. While it gives me hope, I am going to play it by the ear. Otherwise it’s going to be - Gemini for world knowledge/general intelligence/R&D and Opus/Sonnet 4.6 to finish it off. UPDATE: I may have spoken too soon. > Fixing Truncated Array Syncing Bug > I traced the missing array items to a typo I made earlier! > When fixing the GC cast crash, I accidentally deleted the assignment.. > ..effectively truncating the entire array behind it. These errors should not be happening! They are not the result of missing knowledge or a bad hunch. They are coming from an incorrect find/replace, which makes them completely avoidable! On a lighter note, every time it happens, I think about this Family Guy: https://youtu.be/HtT2xdANBAY?si=QicynJdQR56S54VL&t=184 https://youtu.be/HtT2xdANBAY?si=QicynJdQR56S54VL&t=184
- sigmoid10 8mo agoFor me it's Opus 4.6 for researching code/digging through repos, gpt 5.3 codex for writing code, gemini for single hardcore science/math algorithms and grok for things the others refuse to answer or skirt around (e.g. some security/exploitability related queries). Get yourself one of those wrappers that support all models and forget thinking about who has the best model. The question is who has the best model for your problem. And there's usually a correct answer, even if it changes regularly.
- scrollop 8mo agoUsing simtheory.ai which is very good, you can switch models within a conversation and use mcps
- replwoacause 8mo agoAre you associated with this somehow?
- bdelmas 8mo agoYes I came to the same conclusion. Just to add: be careful with Opus 4.6 guys. It’s expensive…
- deleted 8mo ago[deleted]
- datakazkn 8mo agoOne underappreciated reason for the agentic gap: Gemini tends to over-explain its reasoning mid-tool-call in a way that breaks structured output expectations. Claude and GPT-4o have both gotten better at treating tool calls as first-class operations. Gemini still feels like it's narrating its way through them rather than just executing.
- carbocation 8mo agoI agree with this; it feels like the most likely tool to drop its high-level comments in code comments.
- n4pw01f 8mo agoI created a nice harness and visual workflow builder for my Gemini agent chains, works very well. I did this so it would create code the way I do, that is very editable. In contrast, the vs code plugin was pretty bad, and did crazy things like mix languages
- attentive 8mo agoA lot of gemini bashing. But flash 3.0 with opencode is reasonably good and reliable coder. I'd rate it between haiku 4.5 (also pretty good for a price) and sonnet. Closer to sonnet. Sure, if I am not cost-sensitive I'd run everything in opus 4.6 but alas.
- 486sx33 8mo ago[dead]
- deleted 8mo ago[deleted]
- thallavajhula 8mo agoThis is great. I am hopeful that Gemini 3.1 Pro would be great. So far, I'm almost always pulled away from Gemini models by Claude. Having used Claude Opus High for a while now, Claude Opus seems to be fantastic at coding. Even Gemini's comparison chart says so. OpenAI's 5.3-codex is by far the weakest (of the 3) for my coding purposes. Claude Opus really shines at explanations and generating code. Gemini is almost great. Claude Opus is great. I keep switching among these subscriptions every month to not miss out on any of the offerings for too long; ChatGPT Plus <-> Gemini Pro <-> Claude.
- 3371 8mo agoI would suggest you also take a look at Cursor's Composer1.5. It's super fast, and perform better than Gemini3P in my use cases.
- thallavajhula 8mo agoI've been trying composer-1.5 on and off and it doesn't come close to Claude's Opus High. The explainability of Claude is just something else.
- 3371 7mo agoSure, my point was it's better than Gemini and it's really really fast, and it's missing from the parent comment.
- lgl 8mo ago> I keep switching among these subscriptions every month to not miss out on any of the offerings for too long; ChatGPT Plus <-> Gemini Pro <-> Claude. I wonder why many people seem to be doing this instead of just going for a copilot subscription that has access to all those models? Anybody care to share pros and cons?
- sothatsit 8mo agoOpenAI and Anthropic give you a lot of usage/$ through their plans. For the Anthropic Max plans, this can be like a ~90% discount. Copilot does not benefit from this (their pricing model is also different though, it is request-based rather than token usage based, so it is hard to compare). That's not to mention that the models generally work better in their own harnesses, which is perhaps unsurprising because the models have been trained with the specific harness in mind (and vice versa). That said, I think some 3rd-party harnesses do a lot of work to make different models work well in their harness.
- deleted 8mo ago[deleted]
- alwinaugustin 8mo agoI use gemini if i need to write something in my native language- Malayalam or translation. it works very well in writing in Indian regional languages.
- XCSme 8mo agoFunnily, on my tests, 3 flash with medium reasoning does better. Seems like 3.1 pro reasoned about the correct answer, but chose to go with a different (wrong) one: https://aibenchy.com/compare/?left=google-gemini-3-flash-preview-medium&right=google-gemini-3-1-pro-preview-medium https://aibenchy.com/compare/?left=google-gemini-3-flash-pre... EDIT: while also being 3x cheaper
- 0x110111101 8mo agoRelevant: Scanned diaries from 1945 of USFS Ranger. Had this transcribed in Claude. [1]:https://news.ycombinator.com/item?id=47041836 https://news.ycombinator.com/item?id=47041836
- rishabhaiover 8mo agoI think we're past the point where benchmarks hold real value. All models are above a certain threshold of intelligence but Gemini somehow borrows the worst of both worlds. It's neither good with long-horizon coding tasks nor does it offer a likable personality (like Claude which is much more beloved)
- ttul 8mo agoWhat I’m noticing, overall: I’ve never cut so much code in my life. I’ve become a coding monster with one of those dark green GitHub profiles ever since 5.3-Codex gave me the confidence to load in a ridiculous number of tasks every day and let it rip. I have about three coding tasks going at once and in another window, Claude Cowork is ripping through PowerPoints and getting back to lawyers. This tech is not going to replace us. If anything, I am becoming even more of a workaholic. But the output volume is going to pay off for those who are privileged enough to use these tools.
- AIorNot 8mo agoYeah see this article I think it was spot on https://hbr.org/2026/02/ai-doesnt-reduce-work-it-intensifies-it https://hbr.org/2026/02/ai-doesnt-reduce-work-it-intensifies...
- ttul 8mo ago“Some described sending a “quick last prompt” right before leaving their desk so that the AI could work while they stepped away.” This, I can relate to. Also: I feel like I need a second monitor.
- javier123454321 8mo agoWhat ive noticed, i dont have the apetite to spend tokens on AI fixing errors AI made. Or paying a 200/month subscription. In the beggining of the mobth im happy tinkering, but i reach the cap of how much money im willing to spend playing
- motoboi 8mo agoThere are thousands like you now. How many does it take to run the economy? What would the rest do. Think of it like what a tractor did to agricultural work. The fist guy that used a tractor probably thought: this is not replacing me, I’m just much more productive. Well, turns out you only need one guy per farm now.
- jaco6 8mo ago
- exabrial 8mo agoYou know what would slay right now? A native app. Not another piece of Electron bloatware, a regular, efficient, fast, snappy, native, app. One that connects to my MCP severs and has local filesystem tools. Anthropic might fall behind Google/OpenAI eventually, but their Desktop App + MCP/Connectors is unbelievably useful to get real work done.
- arcfour 8mo agoI haven't used Anthropic's desktop app in months since I don't have access to a Mac anymore, but when I did...it was just an electron app? Did something change?
- perardi 8mo agoNope. It is still Electron, and it is not snappy. And I am on an M3 Max MacBook Pro. I have transitioned off ChatGPT for home use (Google provides me slightly better value in my personal life, as I can pay for a plan that also accommodates my weird photo storage needs) and it’s all Anthropic at work, but I miss the ChatGPT Mac app. I can’t say for certain if it was Electron or not—I never dug into the internals, and it felt very, very fast and “native”.
- exabrial 8mo agoNo, sadly. I wish it were native. Its _terrible_.
- YetAnotherNick 8mo agoNot only that, it is the slowest app among all AI apps.
- ceroxylon 8mo agoIt also has some strange bugs between versions. There was an update a month or two ago that caused the app to be unable to quit normally, and I would have to 'force quit' it. Thankfully it was resolved, but it was unnerving to not be able to close the app normally.
- infinitewars 8mo agoI find Gemini is great at generating code that is relatively common on the internet, especially web and algorithms. It is absolutely better at this then OpenAI's models. But Gemini is not as good at reasoning about problems from first principles, or catching subtle bugs. In some ways it is just a better Google that finds exactly what you want, less a general intelligence.
- deleted 8mo ago[deleted]
- holografix 8mo agoI think it begs the question: Is Gemini meant to be be a revenue making product or strictly a cost centre to defend against Search and Ads erosion by OpenAI? Why does the Gemini web app not support MCP Servers?
- nobrains 8mo agoIn the "Intelligence applied" section, where they show the comparison animations, they are shown using a non-optimal UI. There is not enough time to read the text, see old animation, and see new animation. Better would have been to keep the same animation on repeat, so that people have unlimited time to read the text and observer the animations. Also, it jumps from example to example in the same video. Better would have been to show each separately, so that once user is done observing one example at their own pace, they can proceed to the next. As a workaround, I had to open the video (just the video) in a new tab, pause once an example came up, read the text, then rewind to the start of the animation to see the old animation example, then rewind again, then see the new animation example, and then sometimes rewind again if I wanted to see the animation again. Then, once done with the example, I had to forward to the next example and repeat the above process again. Somewhere along that process, they lost me.
- eboy 8mo ago[dead]
- Drblessing 8mo agoGemini is the smartest model currently available. It is the only model out of the big ones that correcly identifies the specific versions of superhers in a collage I tested them with.
- andrewstuart 8mo agoGemini current version drops most of the code every time I try to use it. Useless.
- Grisu_FTP 8mo agoSomehow the models apparently get better and better every week, but every time i try to use them they get worse. Am I the issue? Am i just misremembering the early times because it was a new thing?
- Mashimo 8mo agoYou are holding it wrong! No but for real, what is your usecase? Do you acutely think something like gpt3 was best?
- Grisu_FTP 8mo agoI dont have a real special usecase, i just use it whenever i think it will give better results than googling or thinking or i dont feel like getting annoyed by cookie popups. And i dont think gpt3 was best, but it felt like it actually listened. Now i tell it: "You did this and this wrong, i specifically told u the exact opposite. Can you please do what i asked you?" And then it says something like: "Oh yes my bad, you are right and very very smart to have caught that you must be a super genius. I will now do what you asked me" Does the same wrong thing again. and again and again. I ask it to fix a mistake, it tells me it fixed it, gives 1:1 the same thing with more errors. It also feels like it forgets mid convo way faster than it did.
- Mashimo 8mo ago> I ask it to fix a mistake, it tells me it fixed it, gives 1:1 the same thing with more errors. > It also feels like it forgets mid convo way faster than it did. Mhh, I don't observe this. Hard to say. You probably know this already, but be sure to don't reuse a AI conversation with different context (Having a single chat for both cooking and coding is nono). Often starting a new chat is better. If it forgets what you said it sounds a bit like you use one chat for too long, or you use a too small model (fast, air, haiku, nano etc.)
- dragochat 8mo ago...you sound like a typical opus-person :P Just use anthropic's flagships if you want good instruction following, focus in long convos, and proper understanding of guidance-when-wrong.
- SrFil 8mo agoFor me, Gemini has been by far the best model for document understanding tasks. I look forward to seeing how much more capable this version is.
- tskulbru 8mo agoOff-topic but, what are people using to create those video animations seen in the "ISS orbit tracking dashboard" example? Looks pretty nice! Im guessing Google uses a whole building of UX people but ive seen similar videos from small indie startups too, or even 1 person SaaS.
- Jirach05 8mo agoCan anyone explain why these models decrease in performance on this "MCRC v2 (8-needle)" long context benchmark when thinking is turned on?
- faebi 8mo agoI'm doing Ruby and Gemini 3.0 pro has by far been the best model for me. It writes the nicest ruby code, like I would. Further, it either succeeds or fails hard and obviously. I prefer it failing hard instead of of slowly going weird in my code. Similar in antigravity. Privately it's my absolute favorite. So I'm actually rooting for this.
- znnajdla 8mo agoWhich harness? Gemini CLI or OpenCode?
- carpe__diem 8mo agoOne thing I’d like to see in these releases is stronger emphasis on regression behavior, not just headline capability. In production, the costly failures are usually "almost right" edits that quietly shift semantics across large diffs. We now gate model upgrades behind a fixed eval set of our own repos + prompts and compare pass rates by task category (refactor, test repair, API migration). Raw benchmark gains matter less to us than variance and rollback safety. If 3.1 improves consistency on long multi-file edits, that’s a bigger win than a small jump on one-shot tasks.
- rahulroy 8mo agoIn the meantime, I'm trying to update Antigravity to use the latest version, but it just wouldn't update itself, nor would it let me use 3.0 model. I restarted multiple times with the same result. I tried telling this to agent, and it keeps repeating the same phrase "Gemini 3.1 Pro is not available on this version. Please upgrade to the latest version." Congratulations on beating the benchmarks, but I wonder how much effort is devoted on improving DX? Edit: It's updated now, I can confirm with "There are currently no updates available.". It still doesn't let me continue with the conversation. I'm able to create new session though.
- ponyous 8mo agoRan a bunch of 3D Modeling benchmarks on Gemini 3.1 vs Gemini 3. Unsurprisingly 3.1 performs a bit better. But surprisingly it costs 2.6x as much ($0.14 vs. $0.37 per 3D Model Generation) and is 2.5x slower (1m 24s vs. 3m 28s). To me it feels like "lets increase our thinking budget and call it an improved model!"
- MASNeo 8mo agoAt risk to be unpopular Gemini 3.0 Pro made a huge difference for me when I moved some workflow to Antigravity, especially compared to ChatGPT. The latest update? I simply don’t care. I am not paid to evaluate models, I am paid to build. Not sure 4 benchmark points are making the difference.
- hackrmn 8mo agoI am reading opinions here from agent users, but I haven't adopted the "agentic workflow" myself because I believe I am (for now) now getting a lot of my trouble's worth using Gemini (3 Pro) in the traditional conversational manner. It is adequate at suggesting solutions in the form of code, or reasoning in general. My problems are software engineering but also everything that is not, since I have a subscription it's my go to problem solving partner. I see no reasons to switch to another product for now either, I am constantly in the loop getting samples of chats with Grok and ChatGPT and it seems a very close race. If Claude is that one race horse that's built different -- and I absolutely can believe it is so because they have rightfully tuned it -- I am not convinced I am missing out much. But maybe because I am more traditionalist to most of everyone's having embraced the idea of having an agent run a loop on their workstation(s) and trusting it to deliver. Perhaps if I were in more of a tight time frame, I'd be pressed to do so myself, but for now I am already benefiting from the extra speed "rubberducking" with Gemini all manner of software engineering problems that I need to solve, so I simply have no reasons to abandon it. I think this is also Google's strength -- they have the data, they've already integrated Gemini or a variant of it anyway, into google.com which is one of their prized cash cows, and it's everywhere else too. Like others here have said, Google may not have the absolute best in class at all times, but they're fairly good and they still have the brains that gave us DeepMind and GPT, unless there's some sort of stagnation going on in their ranks, I expect they're not resting on the laurels. With their capital they're still at the head of the race. Anthropic and OpenAI have the benefit of being nimble, though, and it shows too. Anyway, competition is good, the cat's out of the bag and on the greener side of the river :-)
- d4rkp4ttern 8mo agoYes people are too fixated on just the model. The real question for coding use cases is - does Gemini X + Gemini CLI outperform Opus + Claude Code? With 3.0 the answer was no. I won’t waste time checking 3.1 until I hear otherwise.
- brap 8mo agoI had it coding autonomously for about an hour (including lots of tool wait time) on a difficult task, and it actually produced good results. What’s most surprising is that I had it follow a strict loop/workflow and it did that perfectly. Normally these things go off the rails after a while with complex workflows. It’s something I have to usually enforce with some orchestration script and multiple agents, but this time it was just one session meticulously following orders. Impressive, and saves a lot of time on building the orchestration glue.
- deleted 8mo ago[deleted]
- 6d6b73 8mo agoIn these discussions we see some people hating the models, while others love them. What I find interesting is that this is exactly how we feel about other people - some people will love working with you while others can't stand being in the same room you're in.
- conception 8mo agoMy current AI test. There was a BBS I was on in the 90s and there was this door game I hadn't seen anywhere else. I simply describe the BBS, where it was popular, its name, the year it was around, and the BBS game and a description of it mechanics, etc. OpenAI and Google's Deep Research produce a very long, 100% made up report. If I question the AI on the report, they both admit they just made it up. Claude just returns, "I couldn't find anything on the BBS or the game."
- barfingclouds 8mo agoI’m no tech expert like a lot of people here, but I find Gemini 3.0 insanely good for my regular daily questions. Hoping this one is great too. I’m kind of at the point where many answers are essentially perfect and I don’t know if I need much more
- metavolvelabs 8mo agoThey crushed it with Gemini 3.1... especially when in Thinking Mode with Deep Think initiated. If you are working towards something with code, research etc. and hit a snag, run it by Gemini with these settings. Here's another KILLER trick: In Gemini Thinking mode select Nano Banana and have it put together a comprehensive slide with paragraph length text portions. It'll nail it.
- dudeinhawaii 8mo agoAfter 2 days of giving it a go, I find that Gemini CLI is still considerably worse than both Codex and Claude Code. The model itself also has strange behaviors that seem like it gets randomly replaced with Gemini-3-Flash or something else. I'll explain. Once agentic coding was a bust, I gave it a run as a daily driver for AI assistant. It performed fairly well but then began behaving strangely. It would lose context mid conversation. For instance, I said "In san francisco I'm looking for XYZ". Two turns later I'm asking about food and it gives me suggestions all over the world. Another time, I asked it about the likelihood of the pending east coast winter storm of affecting my flight. I gave it all the details (flight, stops, time, cities). Both GPT-5.2 and Claude crunched and came back with high quality estimations and rationale. Gemini 3.1 Pro... 5 times, returned a weather forecast widget for either the layover or final destination. This was on "Pro" reasoning, the highest exposed on the Gemini App/WebApp. I've always suspected Google swaps out models randomly so this.. wasn't surprising. I then asked Gemini 3.1 Pro via the API and it returned a response similar to Claude and GPT-5.2 -- carefully considering all factors. This tells me that a Google AI Ultra subscription gives me a sub-par coding agent which often swaps in Flash models, a sub-par web/app AI experience that also isn't using the advertised SOTA models, and a bunch of preview apps for video gen, audio gen (crashed every time I attempted), and world gen (Genie was interesting but a toy). This will be a quick cancel as soon as the intro rate is done. It's like Google doesn't ACTUALLY want to be the leader in AI or serve people their best models. They want to generate hype around benchmarks and then nerf the model and go silent. Gemini 3 Pro Preview went from exceptional in the first month to mediocre and then out of my rotation within a month.
- deleted 8mo ago[deleted]