12 ms·
Gemini 3.0 spotted in the wild through A/B testing
- Topfi 1y agoHas been ongoing for roughly a month now, with a variety of checkpoints along the usual speculation. As it stands, I'd just wait for the official announcement, prior to making any judgement. What their release plans are, whether a checkpoint is a possible replacement for Pro, Flash, Flash Lite, a new category of model, won't be released at all, etc. we cannot know. More importantly, because of the way AIStudio does A/B testing, the only output we can get is for a single prompt and I personally maintain that outside of getting some basic understanding on speed, latency and prompt adherence, output from one single prompt is not a good measure for performance in the day-to-day. It also, naturally, cannot tell us a thing about handling multi file ingest and tool calls, but hype will be hype. That there are people who are ranking alleged performance solely by one-prompt A/B testing output says a lot about how unprofessionally some evaluate model performance. Not saying the Gemini 3.0 models couldn't be competitive, I just want to caution against getting caught up in over-excitement and possible disappointment. Same reason I dislike speculative content in general, it rarely is put into the proper context cause that isn't as eyecatching.
- tuesdaynight 1y agoI understand that hyping is the career of a lot of people, but it's a little annoying how every Twitter link posted here is full of "IT'S A GAME CHANGER!!! NOTHING IS THE SAME ANYMORE!!! BRACE FOR IMPACT!!!" energy. The examples look great, but it's hard to ignore the unprofessional evaluation that you described.
- cactusplant7374 1y agoThe example in this case is an SVG of a video game controller.
- jmkni 1y agoI might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent at UI/UX web development, in my experience Very excited to see what 3.0 is like
- swalsh 1y agoWe've moved to it for our clinical workflow agents. Great quality, better pricing and performance compared to Anthropic.
- OsrsNeedsf2P 1y agoWhat's your use case? We've found Gemini to work well with large context windows, but it sucks at calling MCPs and is worse at writing code
- jmkni 1y agoBuilding out user interfaces in html and scss (mainly in Angular) You need to give it detailed instructions and be willing to do the plumbing yourself, but we've found it to be very good at it
- moffkalast 1y agoAngular is probably what sets your use case apart. It has a very rigidly defined style which Gemini can't break, so you avoid the main downside of it, i.e. completely refactoring everything for no reason.
- cj 1y agoI use LLMs a lot for health related things (e.g. “Here are 6 bloodwork panels over the past 12 months, here’s a list of medical information, please identify trends/insights/correlations [etc]”) I default to using ChatGPT since I like the Projects feature (missing from Gemini I think?). I occasionally run the same prompts in Gemini to compare. A couple notes: 1) Gemini is faster to respond in 100% of cases (most of my prompts kick ChatGPT into thinking mode). ChatGPT is slow. 2) The longer thinking time doesn’t seem to correlate with better quality responses. If anything, Gemini provides better quality analyses despite shorter response time. 3) Gemini (and Claude) are more censored than ChatGPT. Gemini/Claude often refuse medical related prompts, while ChatGPT will answer.
- msp26 1y agoRumour is a release on the 22nd I believe
- FergusArgyll 1y agoBet on it! https://manifold.markets/ItsMe/gemini-3-releases-october-22 https://manifold.markets/ItsMe/gemini-3-releases-october-22
- smusamashah 1y agoIt's based on leaked p photo of a deck.
- CSMastermind 1y agoPretty sure everyone said that's an old date and that's no longer the timeline but hopefully that's just misinformation and we'll get it on the 22nd.
- smusamashah 1y agohttps://x.com/chetaslua https://x.com/chetaslua is experimenting a lot with Gemini 3 and posting its results (various web desktops, a vampire survivor clone which is actually very playable, voxel 3d models, other game clones, SVG etc). They look really good, specially when they are one-shot.
- joshhug 1y agoThis was cool: https://codepen.io/ChetasLua/pen/yyezLjN https://codepen.io/ChetasLua/pen/yyezLjN Somewhat amusing 4th wall breaking if you open Python from the terminal in the fake Windows. Examples: 1. If you try to print something using the "Python" print keyword, it opens a print dialog in your browser. 2. If you try to open a file using the "Python" open keyword, it opens a new browser tab trying to access that file. That is, it's forwarding the print and open calls to your browser.
- joshhug 1y agoAh, that's because the "python" is actually just using javascript evals. } else if (mode === 'python') { if (cmd === 'exit()') { mode = 'sh'; } else { try { // Safe(ish) eval for demo purposes. // In production, never use eval. Use a JS parser library. // Mapping JS math to appear somewhat pythonesque let result = eval(cmd); if (result !== undefined) output(String(result)); } catch (e) { output(`Traceback (most recent call last):\n File "<stdin>", line 1, in <module>\n${e.name}: ${e.message}`, true); } }
- solarkraft 1y agoI hope they are going to solve the looping problem. It’s real and it’s awful. It’s so bad that the CLI has a loop detection which I promptly ran into after a minute of use. In the Gemini app 2.5 Pro also regularly repeats itself VERBATIM after explicitly being told not to multiple times to the point of uselessness.
- kristofferR 1y agoI hope Gemini 3.0 will also be free, like Gemini 2.5 Pro is if you use the CLI or the right subdomain.
- floppyd 1y ago2.5 Pro is limited to 100 request per day every where I think. My Gemini CLI is authed through the Google Account (not API key) and after 100 requests it switches to Flash, API keys are also limited to 100 requests each (and I think there's a limit on free keys now as well)
- SweetSoftPillow 1y agoAnd there are some wild examples: https://news.ycombinator.com/item?id=45578346 https://news.ycombinator.com/item?id=45578346
- incomingpain 1y agoThis is super exciting. Gemini 2.5 pro was starting to feel like it's lagging behind a little bit; or at least it's still near the best but 3.0 had to be coming along. It's my goto coder; it just jives better with me than claude or gpt. Better than my home hardware can handle. What I really hope for 3.0. Their context length is real 1 million. In my experience 256k is the real limit.
- jedberg 1y ago> Gemini 3.0 is one of the most anticipated releases in AI at the moment because of the expected advances in coding performance. Based on what I'm hearing from friends who work at Google and are using it for coding, we're all going to be very disappointed. Edit: It sound like they don't actually have Gemini 3 access, which would explain why they aren't happy with it.
- phendrenad2 1y agoWhich should surprise no one. LLMs are reaching diminishing returns, unless we find a way to build GPUs more cheaply.
- mwest217 1y agoGemini 3.0 isn't broadly available inside Google. There's are "Gemini for Google" fine-tuned versions of 2.5 Pro and 2.5 Flash, but there's been no broad availability of any 3.0 models yet. Source: I work at Google (on payments, not any AI teams). Opinions mine not Google's.
- kridsdale3 1y agoHate to spoil this excitement, but we at Google do not have Gemini 3 available to us for use in Vibecoding.
- andrewstuart 1y agoChatGPT is great at analysis and problem solving but often gets lost and loses code and ends up in a tangle when trying to write the code. So I get ChatGPT to spec out the work as a developer brief including suggested code then I give it to Gemini to implement.
- deepanwadhwa 1y agoGemini2.5 Pro has assisted me better in every aspect of AI as compared to ChatGPT5. I hope they don't screw up Gemini 3 like OpenAI screwed ChatGPT with GPT5.
- adjbsibdunhe 1y agoAdjhe
- grej 1y agoMy strange observation is that Gemini 2.5 Pro is maybe the best model overall for many use cases, but starting from the first chat. In other words, if it has all the context it needs and produces one output, it's excellent. The longer a chat goes, it gets worse very quickly. Which is strange because it has a much longer context window than other models. I have found a good way to use it is to drop the entire huge context of a while project (200k-ish tokens) into the chat window and ask one well formed question, then kill the chat.
- CaptainOfCoit 1y ago> The longer a chat goes, it gets worse very quickly. This has been the same for every single LLM I've used, ever, they're all terrible at that. So terrible that I've stopped going beyond two messages in total. If it doesn't get it right at the first try, its more and more unlikely to get it right for every message you add. Better to always start fresh, iterate on the initial prompt instead.
- grej 1y agoYes agree, but it seems gemini drops off more quickly than other foundation models for some reason.
- TurboSkyline 1y agoHey, this has been my experience, too! I like Gemini because I’ve told it the tone and style I like my answers in and the first answer is very, very on point with that. But several times I’ve noticed that if I ask follow-up questions, the style immediately changes for the worse, often no longer following my preferences. I’ve also noticed that in follow-ups it makes really bad analogies that are not suitable at all for the kind of audience that the first response is catered to. I’ve been clicking the thumbs-down button every time I’ve seen this and commenting on the change in style and quality, so hopefully the training process will ingest that at some point.
- simonw 1y agoThis is a very good pelican. I'm really looking forward to trying out Gemini 3 myself. https://x.com/cannn064/status/1978779247930953885 https://x.com/cannn064/status/1978779247930953885
- jacquesm 1y agoThat's good? Looks like complete crap to me.
- recallingmemory 1y agoHave you seen the current SVG art that LLMs generate? It's pretty comical what they output.
- OtherShrezzing 1y agoI like the pelican riding a bike test, but my standards for what’s “good” seem higher than generally expected by others. The models can generate hyper realistic renders of pelicans riding bikes in png format. They also have perfect knowledge of the SVG spec, and comprehensive knowledge of most human creative artistic endeavours. They should be able to produce astonishing results for the request. I don’t want to see a chunky icon-styled vector graphic. I want to see one of these models meticulously paint what is unambiguously a pelican riding what is unambiguously a bicycle, to a quality on-par with Michelangelo, using the SVG standard as a medium. And I don’t just want it to define individual pixels. I want brush strokes building up a layered and textured birds wing.
- scrollaway 1y agoIt’s not true agi until it can recreate the emotional state of Van Gogh when he cut his ear and express the pain through the brush, in svg format.
- paintbox 1y ago>I like the pelican riding a bike test, but my standards for what’s “good” seem higher than generally expected by others. If you train for your first marathon, is your goal to run it under 2h? We are all looking forward to perfect results, but our standards are reasonable. We know what the results were last month, and judge the improvement velocity. Nobody thinks that's a good SVG of a pelican riding a bike - on it's own. But it's a lot better compared to all the other LLM-generated SVGs of a pelican riding a bike. We judge relative results - you judge absolute results. Confusion ensues.
- jjcm 1y agoThere are a lot more of these Gemini 3 examples out on twitter right now. After seeing them, I bought Google stock. What shocks me about its output is it actually feels like it's producing net new creative designs, not just regurgitated template output. Its extremely hard to design in code in a way that produces consistent, beautiful output, but it seems to be achieving it. That combined with Google being the only one in the core model space that is fully vertically integrated with their own hardware makes me feel extremely bullish on their success in the AI race.
- bl4ckneon 1y agoI'm no financial advisor but I can tell you that it's not a financially sound decision to buy stock based off of speculative hype Twitter posts. But you do you if you have "fun money" to throw around!
- drcode 1y agobuy on the rumor, sell on the news
- weatherlite 1y agoI agree, though the time to buy was 6 months ago when everyone hated the stock. I think it can still appreciate nicely in the coming 1-3 years, search isn't really going anywhere and their other pieces (Youtube, Cloud, A.I subscriptions) will do good. If this bull market continues 4 trillion market cap is reasonable.
- butlike 1y agoAfter looking at the Gemini 2.5 iterations under Appendix: “Gemini 3.0” A/B result versus the Gemini 2.5 Pro model, I couldn't help but think: It's like a child who's given up on their homework out of frustration. Iteration 1 is way off, 2-3 seem to be improvements, then it starts to veer wildly off-track until essentially everything is changed in iteration 10. E.g. "HERE, IS THIS WHAT YOU WANT?!" Which led me to hypothesize that context pollution could be viewed as a defense mechanism of sorts. Pollute the context until the prompter (perturber) stops perturbing.
- cindyllm 1y ago[dead]
- smusamashah 1y agoThe vampire survivor clone which is very playable https://x.com/cannn064/status/1977542849848823845 https://x.com/cannn064/status/1977542849848823845 https://codepen.io/jules064/pen/bNErYKX https://codepen.io/jules064/pen/bNErYKX With more work https://x.com/cannn064/status/1977882763832201643 https://x.com/cannn064/status/1977882763832201643 https://codepen.io/jules064/pen/PwZKMQq https://codepen.io/jules064/pen/PwZKMQq
- jwithington 1y agogrok 4's controller lol
- lampreyface 1y agoHere's your controller, bro.
- ofek 1y agoThe sentiment in this thread surprises me a great deal. For me, Gemini 2.5 Pro is markedly worse than GPT-5 Thinking along every axis of hallucinations, rigidity in its self-assured correctness and sycophancy. Claude Opus used to be marginally better but now Claude Sonnet 4.5 is far better, although not quite on par with GPT-5 Thinking. I frequently ask the same question side-by-side to all 3 and the only situation in which I sometimes prefer Gemini 2.5 Pro is when making lifestyle choices, like explaining item descriptions on Doordash that aren't in English. edit: It's more of a system prompt issue but I despise the verbosity of Gemini 2.5 Pro's responses.
- Diggsey 1y agoI've found Gemini to be much better at completing tasks and following instructions. For example, let's say I want to extract all the questions from a word document and output them as a CSV. If I ask ChatGPT to do this, it will do one of two things: 1) Extract the first ~10-20 questions perfectly, and then either just give up, or else hallucinate a bunch of stuff. 2) Write code that tries to use regex to extract the questions, which then fails because the questions are too free-form to be reliably matched by a regex. If I ask Gemini to do the same thing, it will just do it and output a perfectly formed and most importantly complete CSV.
- arresin 1y agoMy honest belief is that they’re are bots. I also find 2.5 worse.
- cageface 1y agoFor writing code at least this has been exactly my experience. GPT5 is the best but slow. Sonnet 4.5 is a few notches below but significantly faster and good enough for a lot of things. I have yet to get a single useful result from Gemini.
- coffeeaddict1 1y agoYep, I agree. Gpt 5 thinking is by far the best reasoning model ime. Gemini 2.5 pro is worse in pretty much everything.
- 1y ago
- 1oooqooq 1y agoit is wild to me that people will see that invisible change in output they have zero insight, opinion, let alone control... and say "perfect! let's build a business on top of it!"
- nextworddev 1y agoMy friends at Google hate AI coding with passion. I have some theories as to why. But anyone here venture a guess?
- ares623 1y agoTraining their replacements?
- speedgoose 1y agoConservatisme, resistance to change, fear of losing the skills and becoming irrelevant.
- fauigerzigerk 1y agoPossibly, but I think something else could be happening at large companies that are fearful of missing a sea change. Managers will be wary of exactly the sort of motivations for resistance that you mentioned. So they will try to counteract that by putting in place quantitative metrics to incentivise or even force AI use where it doesn't necessarily make sense. This could cause resentment and fear irrespective of the real benefits that AI undboutedly brings. This is complete speculation on my part where Google specifically is concerned. It's just something I think will inevitably happen at some companies.
- botanical76 1y agoAI coding is in many ways antithetical to great software engineering. It is the current spear-edge of the investor pressure to ship products faster, and monetize users more aggressively, all at the cost of quality, reliability, ethics, security. If you, as a software engineer, once held an ideal about programming as an art or craft, AI coding flies in the face of all that. It turns out that maximising for short-term profit leaves many other objectives behind in its wake.
- dudeinhawaii 1y agoIt's very interesting, and also quite frustrating that no two AI experiences are the same. Scrolling through the threads here and they're all seemingly contradictory. I've had the Gemini 3.0 (presumably) A/B test and been unimpressed. It's usually on fairly novel questions. I've also gotten to the point where I often don't bother with getting Gemini's opinion on something because it's usually the worst of the bunch. I have a Claude Pro and OpenAI Pro sub and use Gemini 2.5 Pro via key. The most glaring difference is the very low quality of web search it performs. It's the fastest of the three by far but never goes deep. Claude and Gemini seemingly take a problem apart and perform queries as they walk through it and then branch from those. Gemini feels very "last year" in this regard. I do find it to be top notch when it comes to writing oriented tasks and sounding natural. I also find it to be fairly good about "keeping the plot" when it comes to creative writing. Claude is a great writer but makes a bit too many assumptions or changes. OpenAI is just flat out poor at creative writing currently due to the issues with "metaphorical language". On speculative tasks -- e.g., "let's rank these polearms and swords in a tier list based on these 5 dimensions" -- Gemini does well. On code work, Gemini is GOOD so long as it's not recent APIs. It tends to do poorly for APIs that have changed. For instance, "do XYZ in Stripe now that the API surface has changed, lookup the docs for the most recent version". GPT-5 has consistently amazed me with its ability to do this -- though taking an eternity to research. It's generally performed great with single-shot code questions (analyze this large amount of code and resolve X or fix Y). On the Agentic front - it's a nonstarter. Both the CLI toolset and every integration I've used as recently as Monday have been sub-par when compared to Codex CLI and Claude Code. On troubleshooting issues (PC/Software but not code), it tends to give me very generic and non-useful answers. "update your drivers, reset your PC". GPT-5 was willing to go more speculative dive deeper, given the same prompt. On factual questions, Gemini is top notch. "Why were medieval armies smaller than Roman era armies" and that sort of thing. On product/purchase type questions, Gemini does great. These are questions like "help me find a 25" stone vanity counter top with sink that has great reviews and from a reputable company, price cap $1000, prefer quality where possible". Unfortunately, like all of the other AI models, there's a non-zero chance that you'll walk through links and find that the product is not as described, not in-stock, or just plain wrong. One last thing I'll note is that -- while I can't put my finger on it -- I feel like the quality of Gemini 2.5 Pro has declined over time while the model has also sped up dramatically. As a pay-per-token user, I do not like this. I'd rather pay more to get higher quality. This is my subjective set of experiences as one person who uses AI everyday as a developer and entrepreneur. You'll notice that I'm not asking math questions or typical homework style questions. If you're using Gemini for college homework, perhaps it's the best model.
- deleted 1y ago[deleted]
- starchild3001 1y ago1. I find Gemini 2.5 Pro's text very easy and smooth to read. Whereas GPT5 thinking is often too terse, and has a weird writing style. 2. GPT5 thinking tends to do better with i) trick questions ii) puzzles iii) queries that involve search plus citations. 3. Gemini deep research is pretty good -- somewhat long reports, but almost always quite informative with unique insights. 4. Gemini 2.5 pro is favored in side by side comparisons (LMsys) whereas trick question benchmarks slightly favor GPT5 Thinking (livebench.ai). 5. Overall, I use both, usually simulatenously in two separate tabs. Then pick and choose the better response. If I were forced to choose one model only, that'd be GPT5 today. But the choice was Gemini 2.5 Pro when it first came out. Next week it might go back to Gemini 3.0 Pro.
- ripped_britches 1y agoAll I can hope for is that the “effective context window” (some level before competency plummets) is like 1m+ tokens. I would give a finger to just put my entire codebase into a model every time I want to talk to it. For now I’m still only talking to parts of the codebase, so to speak.
- chrsw 1y agoHave you tried Claude Code, Cursor, Codex CLI, Gemini CLI, etc?
- ripped_britches 1y agoYes mostly cursor for last 1.5 years but as of this month I am 100% codex CLI. Very freakin good
- aitchnyu 1y agoDo the models evaluate SVGs by "eye" and iterate it? Or we hoping the one-shot result is perfect?
- simonw 1y agoMy benchmark only gives them one chance. I've also tried a variant where the vision models get fed a rendered version and have up to three attempts to make it better. It didn't seem to produce better results, to my surprise.
- nurettin 1y agoHopefully this one will learn to edit files like claude instead of trying ten times consecutively and then shitting the bed.
- elcomet 1y agoI don't understand all the hype for generating SVG with LLM. The task is not really useful, doesn't seem that interesting in single shot as it's really hard, and no human could do it (it would be more useful if the model has visual feedback and could correct the result). And also, since it becomes a popular task, companies will add the examples in their training set, so you're just benchmarking who has the better text to SVG training set, not the overall quality of the model.
- Lucasoato 1y agoOne of my co-founders lost the SVG of our startup logo, and the designer who helped us was away on vacation. I really wanted to experiment with some logo animations for an upcoming demo, so I decided to take matters into my own hands. I grabbed a high-quality PNG, gave it to ChatGPT, and managed to recreate the SVG from the image, after quite a bit of prompting and tweaking. But it worked out great!
- bertylicious 1y agoBut isn't this something Inkscape can do since forever?
- hennell 1y agoMy take is no one really cares about generating SVG, but it's a structured "code" format with very direct visual results. I can't look at 3 piles of code and instantly tell which is best (assuming minimum competence) , but I can judge the SVG outputs very easily. As a quick shot it gets a point across faster and with easier comparison. As a technical comparison it's not so strong, but thats harder to do and judge and less fun to read.
- Topfi 1y agoIt goes back to Sparks of AGI [0] unless I am mistaken. Can recommend the talk, one that has stayed in the back of my mind since I first saw it two years ago. Personally, still have major reservations about throwing claims of intelligence or understanding around, but I do agree that SVG code generation can be a very effective source to get a quick and easy to present understanding of a models ability to output code with a rather open ended prompt that needs a high degree of coherence and were a lot of layers depend/build on each other. Helps that these are eye catching (literally as the output is visual) and easy to grasp. Same reason a lot of hype is created around the web desktops. [0] https://youtu.be/qbIk7-JPB2c?si=_TNRrxN-_5FOlfy5&t=1342 https://youtu.be/qbIk7-JPB2c?si=_TNRrxN-_5FOlfy5&t=1342
- antirez 1y ago"SVG generation as a quality proxy" No need to read further.
- blauditore 1y agoThat doesn't really look like an actual XBox controller. Yes, it's impressive what it can generate, but not really on par with what professional humans could do. As usual, the model can get like 95% close to the gold standard, but the last few percent are the hardest ones. I honestly think that most dream scenarios of AI applications will remain dreams for exactly that reason, and the AI bubble will burst badly. Yes, there are real use cases for the current generation of LLMs and generative models, but they make up only a small fraction of what some of the big companies would like to believe.
- suminjs 1y agoWhile the speed and terseness of models like GPT-5 are great for simple coding tasks or short answers, the verbosity of Gemini is a massive asset for high-stakes tasks where depth matters.
- sd9 1y agoI find verbosity annoying. I prefer depth/accuracy/structure without extra words. Maybe Gemini still wins on that front anyway.
- bgwalter 1y agoPeople were also raving about Gemini 2.5. Allegedly it powers Google's "AI mode", which is the worst model I have tested. EDIT: The religious downvotes are pretty useless. Does the post contain a factual error? Is Google "AI mode" (which has a separate button and is distinct from the "AI" summaries"!) not powered by Gemini 2.5? Then say so. Do you doubt that the "AI" chat that you enter via the separate button is bad? Then say so, but you'll be quite alone with your opinion outside of "AI" echo chambers.
- nprateem 1y agoGemini has developed an annoying habit of writing blog posts or news articles in response to questions. That and continually blowing smoke up my ass. When I tell it I don't need its validation it just replies "Yes, you've got me. That is the sharpest comment you could have made", etc etc
- ethanpark 1y agoI've been switching between Gemini and Claude depending on the task. Gemini 2.5 Pro is incredibly fast and handles large context really well, but I've noticed it can get stuck in loops during longer conversations. Claude is more reliable for iterative coding work. Really curious to see if Gemini 3.0 fixes the context issues, that would be a game changer for my workflow.
- Aissen 1y agoWhy wouldn't this be just the result of a different seed? Is Gemini behaving deterministically by default?