4 ms·
> I'm not talking about other people's incorrect promises _and_ I mentioned a number of things in which proper usage today is different from what people were do
by creesch 1y ago
> I'm not talking about other people's incorrect promises _and_ I mentioned a number of things in which proper usage today is different from what people were doing before.
Alright? What you replied to and the context of this entire thread is about promises that have been made for a while now. In fact, we are approaching the point where we can safely talk about years of hype now. For reference, I am using the gpt-4 release as a significant ramp-up in the hype around LLMs.
> That's not what I said. Don't put words into my mouth.
That might not have been your intention. But, again, given the context of where you are replying it does read like. Even with the qualifier of "in comparison".
Like it or not, your comment does add extra qualifiers to the list.
Look, I am not saying that there isn't any progress being made here. I also agree that LLMs can be useful tools as part of a developer toolkit. What I personally don't agree with is that they can do the same job, even less so in real world scenarios. Even the latest models, including Gemini 2.5 and 03 struggle with moderately complex code base. And yes, the argument is always to let them work on small isolated bits of code. Or that if your requirements are tight enough they produce very good code. Which is entirely true, but I also envy the developers who work in structured environments where their code base is that clean and requirements that well-defined.
So, in my experience, coding with these models still sucks. Using them as interactive rubber duckies, replacements for some of the things I used to spend hours googling, debugging small snippets of code, etc. Sure, there they are very useful tools to me. But, to me, that is not coding with LLMs. That is having LLMs available as a tool whenever I need them.
- dinfinity 1y ago> And yes, the argument is always to let them work on small isolated bits of code. Or that if your requirements are tight enough they produce very good code. You're getting there. The most valuable change is using software (like Cursor) that runs the models in agentic mode so they: 1. find the context themselves (with a basic document with context to point them in the right general direction for the current task). 2. can run commands and specifically tests. 3. iterate on their own output and changes to ensure you don't need to point out what they did wrong manually each time. Just let it find out itself, correct and keep moving until it is a good solution. Think about how any human developer would approach an issue. For a lot of the 'moderately complex code bases', a developer new to it would also need a lot of pointing in the right direction, a lot of trying stuff out and then correcting themselves. Forget treating LLMs like one-shot magic solution givers, but instead as junior devs that you have to provide with all kinds of things to be successful.
- creesch 1y agoRight, so now we have moved to "well yeah, but you need to use LLM agents". Do you truly not see how you are actually continuing the trend of shifting the goalpost every time someone is critical about the use of LLMs? Critical about parts that just a few short months ago had the same promise you now moved to the use of agents? Not to mention that with each iteration the amount of tokens needed goes up substantially. Having worked with LLM APIs and their pricing there simply is no way that Cursor is breaking even on that $20 or $40 per month if everyone uses it fully. Not even close. This very much hints at the costs being hidden right now, subsidized if you will, by VC money. Also, once you have brought junior developers up to speed and guided them they are now slightly more capable developers who can more easily on board on future projects. With LLMs you need to effectively babysit them on each project again. And there are a lot more caveats, prerequisites and moving targets involved that make the promise and reality for many people and companies not something they can actually be met. And again, I am not discounting that there are specific areas where people see benefits from using LLMs in agentic from. But those areas are not as ubiquitous as the hype train leads us to believe. And to start using them you need to set up a lot more in the way of infrastructure and due process as well.
- dinfinity 1y ago> Do you truly not see how you are actually continuing the trend of shifting the goalpost every time someone is critical about the use of LLMs A late reply, but the point is that the old models are harder to get good results with, even in agentic mode. Therein lies the "massive difference in outputted code quality between models" I mentioned. So using the newest models with the right approach quite easily leads to good results, which is what I meant with: "At this point, if your coding experience with LLMs sucks I'd say there is an 80% chance that you're just doing it wrong." This as opposed to the original final claim made by the person I was replying to: "Just roll the dice correctly." That implies that it is just about getting lucky with almost random output. It hasn't been that way for a long time, unless you really try to get the LLM to fail. > With LLMs you need to effectively babysit them on each project again. Only if you're doing it badly. Document and point the AI at documentation properly. Your onboarding process and documentation should not rely on people anyway. Write Once, Read Many.