4 ms·
I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other provider
by Gareth321 10d ago
I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.
This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.
- KeplerBoy 10d ago[dead]
- otabdeveloper4 10d ago> you're just not using the latest model, bro Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world. P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.
- belowavgiq 10d agoWhere did he claim otherwise? Take your meds.
- Gareth321 10d ago> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem. The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.
- vintermann 10d agoThe tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price. It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.
- Gareth321 10d agoI think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.
- vintermann 10d agoI would run into the use limits very quickly, and (for Anthropic) have to switch frameworks. By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.
- Gareth321 10d agoNo argument that DeepSeek is much cheaper, but you certainly get what you pay for.
- irthomasthomas 10d agoartificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash. Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.
- Gareth321 10d agoIt's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-gpt-6-astra https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...
- 074287432996 10d agoImpressive how feverishly you defend the steaming pile of shit that Gemini is.
- TacticalCoder 10d ago> I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models. I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Google/GMail) for anything that is not coding. I find Gemini better/quicker/more polished for basically every single subject out there that is not "write me lines of code".
- Gareth321 10d agoTo be fair, it has been about six months since I tried a paid Google model. I will give it another go to compare. Hallucinations were the main issue back then but perhaps it has come a long way.
- mathgeek 10d agoMuch respect for being able to admit that you haven’t used a model in a while after saying folks probably haven’t used the models you use.
- Gareth321 10d agoOkay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible. One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc. 3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info. I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy. But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt-6-astra https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...) This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.