3 ms·
This is too dismissive. There is a massive difference in outputted code quality between models. The recent Google Gemini 2.5 models have dethroned the Anthropic
by dinfinity 1y ago
This is too dismissive. There is a massive difference in outputted code quality between models. The recent Google Gemini 2.5 models have dethroned the Anthropic ones. OpenAI o3 is the only other one that is even worth considering; all the rest of the OpenAI models are trash in comparison.
In addition to that we've also seen that the way you prompt (amount of use of expert language), what context you provide through instructions, and tool use make a huge difference on the outcome.
At this point, if your coding experience with LLMs sucks I'd say there is an 80% chance that you're just doing it wrong.
- creesch 1y agoThe differences can't be that massive given that the hype already made many of these promises well before these models were ever a thing. Basically you are just adding to it with "You just need to use the latest model, anything before that is trash". Ignoring that the same was said just a few months prior when those models were cutting edge.
- dinfinity 1y ago> The differences can't be that massive given that the hype already made many of these promises well before these models were ever a thing. I'm not talking about other people's incorrect promises _and_ I mentioned a number of things in which proper usage today is different from what people were doing before. > Basically you are just adding to it with "You just need to use the latest model, anything before that is trash". That's not what I said. Don't put words into my mouth. I said that the older models are trash in comparison, not that they are trash. The older models require more work in prompting to get decent results. > Ignoring that the same was said just a few months prior when those models were cutting edge. You are conveniently ignoring the other parts I mentioned to make my claim. Here: "In addition to that we've also seen that the way you prompt (amount of use of expert language), what context you provide through instructions, and tool use make a huge difference on the outcome."
- creesch 1y ago> I'm not talking about other people's incorrect promises _and_ I mentioned a number of things in which proper usage today is different from what people were doing before. Alright? What you replied to and the context of this entire thread is about promises that have been made for a while now. In fact, we are approaching the point where we can safely talk about years of hype now. For reference, I am using the gpt-4 release as a significant ramp-up in the hype around LLMs. > That's not what I said. Don't put words into my mouth. That might not have been your intention. But, again, given the context of where you are replying it does read like. Even with the qualifier of "in comparison". Like it or not, your comment does add extra qualifiers to the list. Look, I am not saying that there isn't any progress being made here. I also agree that LLMs can be useful tools as part of a developer toolkit. What I personally don't agree with is that they can do the same job, even less so in real world scenarios. Even the latest models, including Gemini 2.5 and 03 struggle with moderately complex code base. And yes, the argument is always to let them work on small isolated bits of code. Or that if your requirements are tight enough they produce very good code. Which is entirely true, but I also envy the developers who work in structured environments where their code base is that clean and requirements that well-defined. So, in my experience, coding with these models still sucks. Using them as interactive rubber duckies, replacements for some of the things I used to spend hours googling, debugging small snippets of code, etc. Sure, there they are very useful tools to me. But, to me, that is not coding with LLMs. That is having LLMs available as a tool whenever I need them.
- dinfinity 1y ago> And yes, the argument is always to let them work on small isolated bits of code. Or that if your requirements are tight enough they produce very good code. You're getting there. The most valuable change is using software (like Cursor) that runs the models in agentic mode so they: 1. find the context themselves (with a basic document with context to point them in the right general direction for the current task). 2. can run commands and specifically tests. 3. iterate on their own output and changes to ensure you don't need to point out what they did wrong manually each time. Just let it find out itself, correct and keep moving until it is a good solution. Think about how any human developer would approach an issue. For a lot of the 'moderately complex code bases', a developer new to it would also need a lot of pointing in the right direction, a lot of trying stuff out and then correcting themselves. Forget treating LLMs like one-shot magic solution givers, but instead as junior devs that you have to provide with all kinds of things to be successful.