4 ms·
I swear every time a new model is released it's great at first but then performance gets worse over time. I figured they were fine-tuning it to get rid of bad o
by lispisok 1y ago
I swear every time a new model is released it's great at first but then performance gets worse over time. I figured they were fine-tuning it to get rid of bad output which also nerfed the really good output. Now I'm wondering if they were quantizing it.
- codr7 1y ago[flagged]
- bboygravity 1y agoBut OpenAI breathes honesty. They're open source! They would never do such a thing. /s
- daseiner1 1y agoIt's still a very competitive marketplace
- mathgradthrow 1y agohonestly refreshing take.
- nabla9 1y agoIt seems that least Google is overselling their compute capacity. You pay monthly fee, but Gemini is completely jammed 5-6 hours when North America is working.
- baq 1y agoGemini is simply that good. I’m trying out Claude 4 every now and then and go back to Gemini to fix its mess…
- fasterthanlime 1y agoFunny, I have the exact opposite experience! I use Claude to fix Gemini’s mess.
- symfoniq 1y agoMaybe LLMs just make messes.
- hgomersall 1y agoI heard that, but I'm getting consistent garbage from Gemini.
- dayjah 1y agoFor code? Use the context7 mcp.
- energy123 1y agoGemini is the best model in the world. Gemini is the worst web app in the world. Somehow those two things are coexisting. The web devs in their UI team have really betrayed the hard work of their ML and hardware colleagues. I don't say this lightly - I say this after having paid attention to critical bugs, more than I can count on one hand, that persisted for over a year. They either don't care or are grossly incompetent.
- thorum 1y agoTry AI Studio if you haven’t already: https://aistudio.google.com/ https://aistudio.google.com/
- koakuma-chan 1y agohttps://ai.dev https://ai.dev
- nabla9 1y agoWell said. Google is best in pure AI research, both quality and volume. They have sucked at productization for years. Not not just AI but other products as well. Real mystery.
- energy123 1y agoI don't understand why they can't just make it fast and go through the bug reports from a year ago and fix them. Is it that hard to build a box for users to type text into without it lagging for 5 seconds or throwing a bunch of errors?
- baq 1y agoIf it doesn’t make sense, it makes sense. Nobody will get their promo by ‘fixing bugs’.
- edzitron 1y agoWhen you say "jammed," how do you mean?
- solfox 1y agoI have seen this behavior as well.
- mhitza 1y agoThat was my suspicion when I first deleted my account, when it felt the output got worse in ChatGPT and I found highly suspicious when I saw an errand davinci model keyword in the chatgpt url. Now I'm feeling similarly with their image generation (which is the only reason I created a paid account two months ago, and the output looks more generic by default).
- beering 1y agoAre you able to quantify how quickly your perception gets skewed by how long you use the models?
- mhitza 1y agoI can't quantity it for my past experience, that was more than a year ago, and I wasn't using ChatGPT daily at the time either. This time around it felt pretty stark. I used ChatGPT to create at most 20 different image compositions. And after a couple of good ones at first, it felt worse after. One thing I've noticed recently is that when working on vector art compositions, the results start more simplistic, and often enough look like clipart thrown together. This wasn't my experience first time around. Might be temperature tweaks, or changes in their prompt that lead to this effect. Might be some random seed data they use, who knows.
- Tiberium 1y agoI've heard lots of people say that, but no objective reproducible benchmarks confirm such a thing happening often. Could this simply be a case of novelty/excitement for a new model fading away as you learn more about its shortcomings?
- 85392_school 1y agoI think it's an illusion. People have been claiming it since the GPT-4 days, but nobody's ever posted any good evidence to the "model-changes" channel in Anthropic's Discord. It's probably just nostalgia.
- herval 1y agothere's definitely measurements (eg https://hdsr.mitpress.mit.edu/pub/y95zitmz/release/2 https://hdsr.mitpress.mit.edu/pub/y95zitmz/release/2 ) but I imagine they're rare because those benchmarks are expensive, so nobody keeps running them all the time? Anecdotally, it's quite clear that some models are throttled during the day (eg Claude sometimes falls back to "concise mode" - with and without a warning on the app). You can tell if you're using Windsurf/Cursor too - there are times of the day where the models constantly fail to do tool calling, and other times they "just work" (for the same query). Finally, there's cases where it was confirmed by the company, like Gpt-4o's sycopanth tirade that very clearly impacted its output (https://openai.com/index/sycophancy-in-gpt-4o/ https://openai.com/index/sycophancy-in-gpt-4o/)
- drewnick 1y agoI feel this too. I swear some of the coding Claude Code does on weekends is superior to the weekdays. It just has these eureka moments every now and then.
- herval 1y agoClaude has been particularly bad since they released 4.0. The push to remove 3.7 from Windsurf hasn’t helped either. Pretty evident they’re trying to force people to pay for Claude Code… Trusting these LLM providers today is as risky as trusting Facebook as a platform, when they were pushing their “opensocial” stuff
- JamesBarney 1y agoI'm pretty sure this is just a psychological phenomenon. When a new model is released all the capabilities the new model has that the old model lacks are very salient. This makes it seem amazing. Then you get used to the model, push it to the frontier, and suddenly the most salient memories of the new model are it's failures. There are tons of benchmarks that don't show any regressions. Even small and unpublished ones rarely show regressions.
- JoshuaDavid 1y agoI suspect what's happening is that lots of people have a collection of questions / private evals that they've been testing on every new model, and when a new model comes out it sometimes can answer a question that previous models couldn't. So that selects for questions where the new model is at the edge of its capabilities and probably got lucky. But when you come up with a new question, it's generally going to be on the level of the questions the new model is newly able to solve. Like I suspect if there was a "new" model which was best-of-256 sampling of gpt-3.5-turbo that too would seem like a really exciting model for the first little bit after it came out, because it could probably solve a lot of problems current top models struggle with (which people would notice immediately) while failing to do lots of things that are a breeze for top models (which would take people a little bit to notice).
- beering 1y agoIt’s easy to measure the models getting worse, so you should be suspicious that nobody who claims this has scientific evidence to back it up.