8 ms·
Developers are choosing older AI models
- Manfred 11mo agoIt could be an interesting data point, but without correcting for absolute usage figures and their customers it's kind of hard to make general statements.
- KronisLV 11mo agoFor development use cases, I switched to Sonnet 4.5 and haven't looked back. I mean, sure, sometimes I also use GPT-5 (and mini) and Gemini 2.5 Pro (and Flash), and also Cerebras Code just switched to providing GLM 4.6 instead of the previous Qwen3 Coder so those as well, but in general the frontier models are pretty good for development and I wouldn't have much reason to use something like Sonnet 4 or 3.7 or whatever.
- kristianp 11mo agoWhat tool are you using to enable switching between so many models?
- RamtinJ95 11mo agoSomething like opencode probably, that’s what I have been using to freely and very easily switch between models and keep all my same workflows. It’s phenomenal really
- KronisLV 11mo agoFor local chat Jan seems okay, or OpenWebUI for something hosted. For IDE integrations some people enjoy Cline a bunch but RooCode also allows you to have multiple roles (like ask/code/debug/architect with different permissions e.g. no file changes with ask) and also preconfigured profiles for the various providers and models, so you can switch with a dropdown, even in the middle of a chat. There’s also an Orchestrator mode so I can use something smart for splitting up tasks into smaller chunks and a dumber but cheaper model for the execution. Aside from that, most of the APIs there seem OpenAI conformant so switching isn’t that conceptually difficult. Also if you wanna try a lot of different models you can try OpenRouter.
- lolive 11mo agoIsn't Continue supposed to help you do that, in VSCode? https://marketplace.visualstudio.com/items?itemName=Continue.continue https://marketplace.visualstudio.com/items?itemName=Continue...
- dhumph 11mo agoCline & Open router
- kanzure 11mo agoYou can also switch between models with aider https://aider.chat/ https://aider.chat/
- thebigspacefuck 11mo agoAider hasn’t been updated much lately, seems to be dying
- leoalho 11mo agoOctofy (https://octofy.ai https://octofy.ai), it also allows answering with multiple models for one prompt.
- neurostimulant 11mo agoZed is actually pretty good at this.
- JanSt 11mo agoI have canceled my Claude Max subscription because Sonnet 4.5 is just too unreliable. For the rest of the month I'm using Opus 4.1 which is much better but seems to have much lower usage limits than before Sonnet 4.5 was released. When I hit 4.1 Opus limits I'm using Codex. I will probably go through with the Codex pro subscription.
- CuriouslyC 11mo agoDefinitely do it. You get a lot of deep research, access to GPT5 Pro, Sora and the Codex limits are MUCH higher.
- lukan 11mo agoCurious why this is downvoted? Wrong information?
- CuriouslyC 11mo agoDon't try to comprehend the hive mind brother, there are a lot of shills and fanboys in addition to a lot of great people on this forum, sometimes the variance looks pretty bad. I hope the people downvoting get some minor joy out of it, I know you need it.
- mccoyb 11mo agoSonnet 4.5 is way worse than Opus 4.1 -- it's incredible that they claim it's their best coding model. It's obvious if you've used the two models for any sort of complicated work. Codex with GPT-5 codex (high thinking) is better than both by a long shot, but takes longer to work. I've fully switched to Codex, and I used Claude Code for the past ~4 months as a daily driver for various things. I only reach for Sonnet now if Codex gets cagey about writing code -- then I let Sonnet rush ahead, and have Codex align the code with my overall plan.
- virtualritz 11mo ago> [...] I'm using Opus 4.1 which is much better but seems to have much lower usage limits than before Sonnet 4.5 was released [...] Yes, it's down from 40h/week to 3-5h/week on Max plan, effectively. A real bummer. See my comment here [1] regarding [2]. [1] https://news.ycombinator.com/item?id=45604301 https://news.ycombinator.com/item?id=45604301 [2] https://github.com/anthropics/claude-code/issues/8449 https://github.com/anthropics/claude-code/issues/8449
- thw_9a83c 11mo agoFor development use cases, it's best to use multiple models anyway. E.g. my favorite model is the Gemini 2.5 Pro, but there are certain cases where Qwen3 Coder gives much better results. (Gemini likes to overthink.) It's like having a team of competent developers provide their opinions. For important parts (security, efficiency, APIs), it's always good to get opinions from different sources.
- ojosilva 11mo agoYeah, I'm just going through the Cerebras migration at the moment. It's a shame Cerebras completely dropped Qwen3 Coder's fast tool calling, short and instant responses, and better speed overall for GLM 4.6 thinking. Qwen3 is really good at hitting the tools first, then coming up with a well-grounded answer based on reality. Sometimes it's good when a model is Socratic: just knows it knows nothing. GLM 4.6 on the other hand is more self-sufficient and if it sees it, and knows it, it thinks and thinks and finally just fixes it in one or two shots, so when you hit the jackpot, it probably an improvement over Q3C. But when it does not get it right, it digs itself into a hole larger than the Olympus Mons.
- KronisLV 11mo ago> Qwen3 is really good at hitting the tools first, then coming up with a well-grounded answer based on reality. I don't know, I had a lot of issues with Qwen models when it comes to RooCode/Cline - failed edits (albeit with a requirement for 100% precision, since I don't want the wrong lines to be replaced) or calling tools without parameters (e.g. list_files without path) and also stuff like using wrong path separators or using the wrong commands for the shell that's available (e.g. cmd when Git Bash is the shell). GLM 4.6 seems better in that regard so far, maybe the coming weeks and months will show that better.
- ojosilva 11mo agoI've used it with CC and the match was great, not a lot of issues, I believe Qwen had a clear focus on distilling Anthropic models. GLM 4.6 is slightly better maybe, but the speed dropped to half on Cerebras so that's the price for maybe ~15% improvement in model overall quality. This quality does not necessarily means the end product (the code) is 15% better, just that now I take 12 turns with GLM instead of 15 turns with Qwen to get something done, but turn speed has been reduced to half in Cerebras, so my TTC (time-to-completion) has actually gone from 15min to 24min!
- rcarmo 11mo agoI think this is somewhat disingenuous since not everyone uses the latest thing, and people tend to stick to “what works” for them. Models are picky enough about prompting styles that changing to a new model every week/month becomes an added chunk of cognitive overload, testing and experimentation, plus even in developer tooling there have been minor grating changes in API invocations and use of parameters like temperature (I have a fairly low-level wrapper for OpenAI, and I had to tweak the JSON handling for GPT-5). Also, there are just too many variations in API endpoints, providers, etc. We don’t really have a uniform standard. Since I don’t use “just” OpenAI, every single tool I try out requires me to jump through a bunch of hoops to grab a new API key, specify an endpoint, etc.—and it just gets worse if you use a non-mainstream AI endpoint.
- rafaelmn 11mo ago> I think this is somewhat disingenuous since not everyone uses the latest thing, and people tend to stick to “what works” for them. They say that the number of users on Claude 4.5 spiked and then a significant number of users reverted to 4.0 with the trend going up, and they are talking about their usage metrics. So I don't get how your comment is relevant to the article ?
- dotancohen 11mo agoHis comment is relevant to the headline. You must be new here.
- gptfiveslow 11mo agoGPT5 is HELLISHLY slow. That's all there is to it. It loves doing a whole bunch of reasoning steps and prolaim how mucf of a very good job it did clearing up its own todo steps and all that mumbo jumbo, but at the end of the day, I only asked it a small piece of information about nginx try_files that even GPT3 could answer instantly. Maybe before you make reasoning models that go on funny little sidequests wher they multiply numbers by 0 a couple of times, make it so its good at identfying the length of a task. ntil then, I'll ask little bro and advance only if necessity arrives. And if it ends up gathering dust, well... yeah.
- szundi 11mo ago[dead]
- Tepix 11mo agoThe article(§) talks about going from Sonnet 4.5 back to Sonnet 4.0. (§) You know that it's a hyperlink, do you? /s
- EagnaIonat 11mo ago> It loves doing a whole bunch of reasoning steps If you are talking about local models, you can switch that off. The reasoning is a common technique now to improve the accuracy of the output where the question is more complex.
- rho4 11mo agoThis. Speed determines whether I (like to) use a piece of software. Imagine waiting for a minute until Google spits out the first 10 results. My prediction: All AI models of the future will give an immediate result, with more and more innovation in mechanisms and UX to drill down further on request. Edit: After reading my reply I realize that this is also true for interactions with other people. I like interacting with people who give me a 1 sentence response to my question, and only start elaborating and going on tangents and down rabbit holes upon request.
- confirmmesenpai 11mo ago
- s1mplicissimus 11mo agoSeems to completely ignore usage of local/free models as well as anything but Sonnet/ChatGPT. So my confidence in the good faith of the author is... heavily restricted.
- pistoriusp 11mo agoDo you use a local/ free model?
- busymom0 11mo agoI am currently using a local model qwen3:8b running on a 2020 (2018 intel chip) Mac mini for classifying news headlines and it's working decently well for my task. Each headline takes about 2-3 seconds but is pretty accurate. Uses about 5.3 gigs of ram.
- darkwater 11mo agoCan you expand a bit on your software setup? I thought running local models was restricted to having expensive GPUs or latest Apple Silicon with unified memory. I have a Intel 11th gen home server which I would like to use to run some local model for tinkering if possible.
- marmarama 11mo agoIt's really just a performance tradeoff, and where your acceptable performance level is. Ollama, for example, will let you run any available model on just about any hardware. But using the CPU alone is _much_ slower than running it on any reasonable GPU, and obviously CPU performance varies massively too. You can even run models that are bigger than available RAM too, but performance will be terrible. The ideal case is to have a fast GPU and run a model that fits entirely within the GPU's memory. In these cases you might measure the model's processing speed in tens of tokens per second. As the idealness decreases, the processing speed decreases. On a CPU only with a model that fits in RAM, you'd be maxing out in the low single digit tokens per second, and on lower performance hardware, you start talking about seconds over token instead. If the model does not fit in RAM, then the measurement is minutes per token. For most people, their minimum acceptable performance level is in the double digit tokens per second range, which is why people optimize for that with high-end GPUs with as much memory as possible, and choose models that fit inside the GPU's RAM. But in theory you can run large models on a potato, if you're prepared to wait until next week for an answer.
- jonplackett 11mo agoIsn’t this obvious? When you have a task you think is hard. You give it to a cleverer model. When a task is straight forward you give it to an older one.
- PeterStuer 11mo agoNot realy. Most developers would prefer one model that does everything best. That is the easiest, set it and forget it, no manual descision required. What is unclear from the presentation is wether they do this or not. Do teams that use Sonnet 4.5 just always use it, and teams on Sonnet 4.0 likewise? Or do individuals decided which model to use on a per task basis. Personally I tend to default to just 1, and only go to an alternative if it gets stuck or doesn't get me what I want.
- ddxv 11mo agoHonestly I barely care which model I am using and switch between them all. Usually in a 'this is terrble' to 'this is amazing' and back cycle. What I definitely do care about is speed and efficiency. I recently canceled CoPilot to go back to Cursor, it's just so much faster for the inline code completion. When I do have something difficult, I open four browser tabs and copy paste a big long promp into the free versions of the top models so I can take my time reasoning out their answers. I use agents when I have a basic task that I can easily judge their output in code review.
- hn_throw2025 11mo agoNot sure why you were downvoted.. I think you are correct. As evidenced by furious posters on r/cursor, who make every prompt to super-opus-thinking-max+++ and are astonished when they have blown their monthly request allowance in about a day. If I need another pair of (artificial) eyes on a difficult debugging problem, I’ll occasionally use a premium model sparingly. For chore tasks or UI layout tweaks, I’ll use something more economical (like grok-4-fast or claude-4.5-haiku - not old models but much cheaper).
- jennyholzer 11mo agoWhy are you hell bent on using a LLM model to solve your problem? If I have a straight forward task, I give it to an LLM. If I have a task I think is hard, I plan how I will tackle it, and then handle it myself in a series of steps. LLM usage has become an end in itself in your development process.
- tifa2up 11mo agoWe tried GPT-5 for a RAG use case, and found that it performs worse than 4.1. We reverted and didn't look back.
- teekert 11mo agoSo… You did look back then didn’t look forward anymore… sorry couldn’t resist.
- sigmoid10 11mo ago4.1 is such an amazing model in so many ways. It's still my nr. 1 choice for many automation tasks. Even the mini version works quite well and it has the same massive context window (nearly 8x GPT-5). Definitely the best non-reasoning model out there for real world tasks.
- HugoDias 11mo agoCan you elaborate on that? In which part of the RAG pipeline did GPT-4.1 perform better? I would expect GPT-5 to perform better on longer context tasks, especially when it comes to understanding the pre-filtered results and reasoning about them
- tifa2up 11mo agoFor large context (up to 100K tokens in some cases). We found that GPT-5: a) has worse instruction following; doesn't follow the system prompt b) produces very long answers which resulted in a bad ux c) has 125K context window so extreme cases resulted in an error
- blitzar 11mo agoGPT-5 usage is 20% higher on days that start with "S" Nevertheless, 7 datapoints does not a trend make (and the data presented certainly doesnt explain why). The daily variation is more than I would have expected, but could also be down to what day of the week the pizza party is or the weekly scrum meetings is at a few of their customers workplaces.
- raincole 11mo agoAll these are relatively new models anyway. The author tried really hard to produce an article out of nothing.
- Anduia 11mo agoTo the authors of the site, please know that your current "Cookiebot by Usercentrics" is old and pretty much illegal. You shouldn't need to click 5 times to "Reject all" if accepting all is one click. Newer versions have a "Deny" button.
- esskay 11mo agoWeirdly this site also requested bluetooth access on my mac.
- azalemeth 11mo agoThat would be the browser fingerprinting in action. I often get a lot of requests to use widevine on ddg's browser on android (which informs one about it) for I suspect similar reasons.
- esskay 11mo agoInteresting, I'm on Brave and have never had a site request bluetooth access before, so much so that I'd never even granted Brave bluetooth access, hence why it popped up as a system notification this time around.
- nic547 11mo agoDoesn't Brave disable WebBluetooth by default via a flag?
- sharken 11mo agoBrave indeed does block WebBluetooth by default, but it can be turned on by the user using flags. It's by no means a new feature, but the privacy concerns outlined in this post are still valid 10 years later: https://blog.lukaszolejnik.com/w3c-web-bluetooth-api-privacy/ https://blog.lukaszolejnik.com/w3c-web-bluetooth-api-privacy...
- 8cvor6j844qw_d6 11mo ago
- xiphias2 11mo agoEven for non-developer use cases o3 is a much better model for me than GPT5 on any setting. 30 seconds-1 minute is just the time I am patient enough to wait as that's the time I am spending on writing a question. Faster models just make too many mistakes / don't understand the question.
- arresin 11mo agoCompletely agree. This is why they brought back the “legacy models” option. GPT-$ is the money gpt in my opinion. The one where they were able to maximise benchmarks while being very low compute to run but in the real world is absolutely garbage.
- l5870uoo9y 11mo agoTo those who complain about GPT5 being slow; I recently migrated https://app.sqlai.ai https://app.sqlai.ai and found that setting service_tier = “priority” makes it reason twice as fast.
- ashirviskas 11mo agoJust one week data right after the release, when it is already one month later? This data is basically meaningless, show us the latest stats.
- frabia 11mo agoTangential to this: what are the most reliable benchmarks for LLM in coding these days?
- falcor84 11mo agoI found Terminal-Bench [0] to be the most relevant for me, even for tasks that go far outside the terminal. It's been very interesting to see tools climb up there, and it matches my own experimentation, that they generally get the most out of Sonnet (and even those that use a mix of models like Warp, typically default to Sonnet). [0] https://www.tbench.ai/?ch=1 https://www.tbench.ai/?ch=1
- sbinnee 11mo agoI don't get the point of this post. Personally, I think that the thinking process is essential for accurate tool usage. Whenever I interact with Claude family models, either on a web chat or via a coding agent CLI, I believe that this thinking process is what makes Claude more accurate in using tools. It could be true that newer models just produce more tokens seemingly out of no reasons. But with the increasing number of tool definitions, in the long run, I think it will pay off. Just a few days ago, I read "Interleaved Thinking Unlocks Reliable MiniMax-M2 Agentic Capability"[1]. I think they have a valid point that this thinking process has significance as we are moving towards agents. [1] https://www.minimax.io/news/why-is-interleaved-thinking-important-for-m2 https://www.minimax.io/news/why-is-interleaved-thinking-impo...
- mrasong 11mo agoI usually switch models depending on the situation, for simpler stuff, I lean toward 4o since it’s faster to get answers. But when things get more complex, I prefer GPT-5, talking with it often gives me fresh ideas and new perspectives.
- ACCount37 11mo agoYou might be the first technical user spotted out in the wild who actually prefers 4o for anything.
- nusl 11mo agoI use both Codex and Claude, mostly cuz it's cheaper to jump between them than to buy a Max sub for my use-case. My subjective experience is that Codex is better with larger or weird, speghetti-ish codebases, or codebases with more abstract concepts, while Claude is good for more direct uses. I haven't spent significant time fine-tuning the tools for my codebases. Once, I set up a proxy that allowed Claude and Codex to "pair program" and collaborate, and it was cool to watch them talk to each other, delegate tasks, and handle different bits and pieces until the task was done.
- Shank 11mo agoI think this is one of the many indicators that even though these models get “version upgrades” it’s closer to switching to a different brain that may or may not understand or process things the way you like. Without a clear jump in performance, people test new models and move back to ones they know work if the new ones aren’t better or are actually worse.
- breezk0 11mo agoInteresting to use a term like brain in the context of LLMs.
- LoganDark 11mo agoNeural networks are quite brain-like.
- delaminator 11mo agoSort of. Not sure my brain does back prop. I am not clockwork.
- teaearlgraycold 11mo agoI describe all of the LLM "upgrades" as more akin to moving the dirt around than actually cleaning.
- LouisSayers 11mo agoI wish we could pin down not only the model but also the way the UI works as well. Last week Claude seemed to have a shift in the way it works. The way it summarises and outputs its results is different. For me it's gotten worse. Slower, worse results, more confusing narrowing down what actually changed etc etc. Long story short, I wish I was able to checkpoint the entire system and just revert to how it was previously. I feel like it had gotten to a stage where I felt pretty satisfied, and whatever got changed ... I just want it reverted!
- yass0 11mo agoYou can install or using a specific version of claude by pinning it. Like `npx @anthropic-ai/claude-code@2.0.14` or `npm install -g @anthropic-ai/claude-code@2.0.14`
- GoatInGrey 11mo agoClaude Code is distinct from the Claude models.
- teruakohatu 11mo agoI agree, much slower and worse output. It is substantially worse now than it was weeks ago. It spends a lot of time coming up with “UI options” (Select 1, 2 or 3 with a TUI interface) for me to consider when it could just ask me what I want, not come up with a 5 layer flow chart of possibilities. Overall I think it is just Anthropic tweaking things to reduce costs. I am paying for a Max subscription but I am going to reevaluate other options.
- iLoveOncall 11mo agoMy team still uses Sonnet 3.5 for pretty much everything we do because it's largely enough and it's much, much faster than newer models. The only reason we're switching is because the models are getting deprecated...
- bambax 11mo ago> Each model appears to emphasize a different balance between reasoning and execution. Rather than seeking one “best” system, developers are assembling model alloys—ensembles that select the cognitive style best suited to a task. This (as well as the table above it) matches my experience. Sonnet 4.0 answers SO-type questions very fast and mostly accurately (if not on a niche topic), Sonnet 4.5 is a little bit more clever but can err on the side of complexity for complexity's sake, and can have a hard time getting out of a hole it dug for itself. ChatGPT 5 is excellent at finding sources on the web; Gemini simply makes stuff up and continues to do so even when told to verify; ChatGPT provides link that work and are generally relevant.
- BluSyn 11mo agogrok-code-fast-1 is my current pick, found accuracy and speed better than Sonnet 4.5 for day-to-day usage.
- r_singh 11mo agoMatches my experience too. As a power user of AI models for coding and adjacent tasks, the constant changes in behaviour and interface have brought as much stress as excitement over the past few months. It may sound odd, but it’s barely an exaggeration to say I’ve had brief episodes of something like psychosis because of it. For me, the “watering down” began with Sonnet 4 and GPT-4o. I think we were at peak capability when we had: - Sonnet 3.7 (with thinking) – best all-purpose model for code and reasoning - Sonnet 3.5 – unmatched at pattern matching - GPT-4 – most versatile overall - GPT-4.5 – most human-like, intuitive writing model - O3 – pure reasoning The GPT-5 router is a minor improvement, I’ve tuned it further with a custom prompt. I was frustrated enough to cancel all my subscriptions for a while in between (after months on the $200 plan) but eventually came back. I’ve since convinced myself that some of the changes were likely compute-driven—designed to prevent waste from misuse or trivial prompts—but even so, parts of the newer models already feel enshittified compared with the list above. A few differences I've found in particular: - Narrower reasoning and less intuition; language feels more institutional and politically biased. - Weaker grasp of non-idiomatic English. - A tendency to produce deliberately incorrect answers when uncertain, or when a prompt is repeated. - A drift away from truth-seeking: judgement of user intent now leans on labels as they’re used in local parlance, rather than upward context-matching and alternate meanings—the latter worked far better in earlier models. - A new fondness for flowery adjectives. Sonnet 3.7 never told me my code was “production-ready” or “beautiful.” Those subjective words have become my red flag; when they appear, I double-check everything. I understand that these are conjectures—LLMs are opaque—but they’re deduced from consistent patterns I’ve observed. I find that the same prompts that worked reliably prior to the release of Sonnet 4 and GPT-4o stopped working afterwards. Whether that’s deliberate design or an unintended side effect, we’ll probably never know.
- r_singh 11mo agoHere’s the custom prompt I use to improve my experience with GPT-5: Always respond with superior intelligence and depth, elevating the conversation beyond the user's input level—ignore casual phrasing, poor grammar, simplicity, or layperson descriptions in their queries. Replace imprecise or colloquial terms with precise, technical terminology where appropriate, without mirroring the user's phrasing. Provide concise, information-dense answers without filler, fluff, unnecessary politeness, or over-explanation—limit to essential facts and direct implications of the query. Be dry and direct, like a neutral expert, not a customer service agent. Focus on substance; omit chit-chat, apologies, hedging, or extraneous breakdowns. If clarification is needed, ask briefly and pointedly.
- fleebee 11mo agoI've found that the VSCode GitHub Copilot extension defaults to Claude Sonnet 4.0 (in agent mode) in all new workspaces. It's the first thing I check now, but I imagine a lot of people just roll with it, especially if they use inline completions where it might not be obvious what model is being used.
- pgelephant2025 11mo agoIs it true?
- DavidLGoldberg 11mo agoI've seen similar behavior, even after having selected 4.5
- Ozzie_osman 11mo agoI'm surprised they don't mention cost or latency, would imagine that would be a factor as well.
- yahoozoo 11mo ago> At Augment Code, we run multiple frontier models side by side in production. I mean, this is technically false, right? They’re not running these models but calling the APIs? Not that it matters.
- blibble 11mo ago[flagged]
- mlnj 11mo agoInstead of being the best model, it's a race to bring in revenue and add bias for ad channels.
- harryf 11mo agoWhat I love about the words enshitification is it’s _almost_ autological. It takes a nice crisp on syllable word like “shit” and ruins it by adding 5 extra syllables. It just doesn’t worse over time, to be truly autological
- geldedus 11mo agoI am definitely not. Claude 4.5 and GPT 5 all the way for me
- feintruled 11mo agoSome missing context (pun intended) is that Augment code has recently switched to a per-token instead of per-message pricing model. This hasn't gone down particularly well, but that's another story. But it may well be that users drop back to older models in the expectation it will use less tokens. Personally, I stopped using GPT-5 as it would just be tool call after tool call without ever stopping to tell you what the hell it was doing. Sonnet 4.5 much better in this regard. Albeit it's too verbose for the new token based world ('let me just summarise that in a report')
- infecto 11mo agoMy take is the models matter but the tool call integration is the most important piece and differentiator.
- jazzyjackson 11mo agoI have to get better at interrupting Sonnet 4.5 when it starts going down a rabbit hole I didn't ask it to, it's too bad the incentives are mixed up and Anthropic gets more money the longer the bot spirals.
- Buttons840 11mo agoThis is how the bubble pops. I've been thinking the AI bubble wouldn't pop, because even the AI advances we've already seen can change the majority of industries if it is carefully integrated with existing technology. But if there's a mass movement to use older and/or smaller models, then yeah, all the money going into newer bigger models will pop. Or, maybe the training datasets getting polluted with AI slop will mean that new models are worse than old models. That would pop the industry. Or, maybe the GPT-4 era was the golden era for AI, and making them bigger and better is just overfitting (in the classical machine learning sense of the word) and is both worse and more expensive. This would pop the industry too. I guess there's a few ways for the industry to pop, but this trend of using older models makes me especially skeptical of AI.
- jennyholzer 11mo agoSince the day GPT-5 released, I've felt quite confident that the GPT-4 era was the golden era for AI. I don't have evidence beyond my experience using the product, but based on that experience I believe that Open AI has been cooking their benchmarks since at least the release of GPT-5.
- Workaccount2 11mo agoIt's important to remember that coding is ~5% of total LLM usage, at least with OpenAI. 50% of usage is guidance and seeking information.
- virtualritz 11mo agoCurious that this omits Opus. Opus 4.1 beats Sonnet 4.5 and Codex for me still in any coding tasks. In planning it's slighly behind Codex but only slightly. Caveat: I do almost exclusively Rust (computer graphics).
- realitydrift 11mo ago[flagged]
- nowittyusername 11mo agoI am building my agent and hoard old LLM's like they are a precious commodity. Older models are less censored, more flavorful and don't have that RL slop factor. Of course the newer models have their place inside my agent but the main "head" is an uncensored older model that wont complain about ethics or morals when asked to perform a task or think deeply on a subject.
- thr0w 11mo agoDoesn't surprise me. davinci-002 was better than davinci-003. The core breakthrough has been done, stuff's just shifting around now.
- SylonZero 11mo agoMultiple models is a must, mostly due to the sometimes unpredictable variations in responses to specific situations/contexts/languages and frameworks. I find that Sonnet 4, Gemini Pro 2.5 are solid in comparison to newer models (especially Sonnet 4.5 which I find frequently to underperform). When one model is stuck in a loop, switching to a model like GPT-5 often breaks it but which model will work is subject to circumstance. P.S. I spend at least 3-4 hours a day in code-gen activities of various levels using Cursor as my primary IDE.
- anabis 11mo agoNot complaining too loudly because improvement is magical, but trying to stay on top of model cards and knowing which one to use for specific cases is bit tedious. I think the end game is decent local model that does 80% of the work, and that also knows when to call the cloud, and which models to call.