4 ms·
I don’t like the term taste, but the problem that I have is that LLMs don’t seem to work “good enough”. They seem to be able to solve the immediate problem, but
by boron1006 2mo ago
I don’t like the term taste, but the problem that I have is that LLMs don’t seem to work “good enough”. They seem to be able to solve the immediate problem, but stacking this on the scale of 3-4 devs over 6 months or so doesn’t seem to produce anything.
One thing that I’m particularly frustrated with is the writing quality of LLMs. Like this is the thing that they should be able to do, but I would say almost everything they write has almost no signal.
Over a mid sized AI generated codebase that means I’m reading like 500 words to figure out what a module is even doing.
- ryoshu 2mo agoI like judgement. Taste is a subset.
- vasco 2mo agoFor me taste is judgement about things that don't matter. Like you have taste in clothes. But a judge doesn't have bad taste when he makes a wrong decision, he has bad judgement.
- nemomarx 2mo agoSince they can summarize, you'd really think they would be better at condensing their own output and cutting out the filler after. I wonder if you could use a specialized second pass for it or something?
- a123b456c 2mo agoPrompting is a skill.
- neerajsi 2mo agoIn my recent experience it seems to be a contextual issue. Llms are constantly being dropped into new situations where they have to rediscover high level and cross cutting information about the system they're working on from contextual clues. The internal representations of this state and its projection back out to human language wouldn't be as concise as that of a practitioner or team that develop their own verbiage and ontology over time molded to their system. This verbosity might get better as we figure out better ways for agents to learn long term and use that knowledge to adapt to the users and projects over time. There might also be some good harness improvements we could consider like forked output streams or multiple long lived filter subagents to ensure that output appropriate for thinking is separate from code output and separate from output given to the user driving the session.
- notashelf 2mo agoI'll risk sounding antagonistic and ask you this: if LLMs are not "good enough", why are they still around? Obviously "good enough" was a poor pick of words of words on my behalf---and I'll gladly own it---but they should be "good enough" for something if a significant amount of resources keep getting allocated to them. Yes on a very personal level you look at LLM-generated code and think to yourself "wow this is garbage" but what about the people, as pointed out many times in this thread, that simply do not care? Does that not count as "good enough"?
- SCLeo 2mo ago> I'll risk sounding antagonistic and ask you this: if LLMs are not "good enough", why are they still around? To quote the parent commenter: > They seem to be able to solve the immediate problem, but not long term
- notashelf 2mo agoBut that means they're good enough for the short them. This is why I find it such an interesting question to wield, but also, it is why my wording was poor in the post.
- deleted 2mo ago[deleted]
- deathanatos 2mo agoShort term can, IME, be very short. I've seen people generate, say, a bash script with an LLM. It's generated: short term, the problem is "solved": we've generated a bash script. … but does it work? Someone comes along, reviews it, "this is garbage, and does not do what it says it purports to do". Perhaps it even gave an output: the script computed … something, but it's just GIGO. But that "check if this works" friction is the same friction that is what people try to avoid by generating it with an LLM in the first place. If you're too lazy to write the script, you're practically by definition too lazy to verify it.
- deleted 2mo ago[deleted]
- esikich 2mo agoThis is just a problem of articulating what you want. Most people can't write for shit, I'm sorry. Because they maybe took one writing class in college. It's a skill in it's own. LLMs write better than anyone I've ever worked with and the reasons are obvious. If it rambles or is too "purple", that's on the person prompting it, because they can't write for shit to begin with.
- crooked-v 2mo ago> LLMs write better than anyone I've ever worked with And yet every time I see AI-written documentation, there's paragraphs worth of throat clearing and space-wasters like the word "genuinely" for things that could have been outlined in a couple of bullet points.
- esikich 2mo agoRight, that's on the prompter being ok with that. It does what you tell it to do. I don't understand how you dorks on this site are so obtuse about this shit. Give it examples of what you want, authors, etc. Or don't and just cry. The world is your oyster.
- nunez 2mo agoI'd much rather see badly written text that someone invested time into instead of perfectly written and infinitely soulless AI slop
- esikich 2mo agoLol why, this is just silly
- kuschku 2mo agoThen why is it that ALL content generated by LLMs, no matter who prompted it, is of such low quality?
- lelanthran 2mo ago> LLMs write better than anyone I've ever worked with If you're talking about prose, this is obviously false. All the LLM prose I see is so mentally fatiguing to read that I usually just give up. There are no lack of examples of poorly written LLM prose.
- Urb_RS 2mo ago[flagged]
- KurSix 2mo ago[dead]
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- CM30 2mo agoI think there are really two types of use cases for LLMs here, and people that get value out of them. First are those that care very little about the quality of the work, or the process behind it, and just want a 'finished' product no matter what. These are many of the folks vibe-coding everything without even looking at the output, and a depressingly number of those getting hacked or what not. But hey, if you've got an idea for a complex app or website but no interest in actually building it, an LLM will get you... something vaguely like what you wanted. Second are small projects and small sections of existing ones, most of which easily fit inside the LLM's context window. If you're making a fairly basic WordPress plugin or React component or one page website, an LLM will probably be able to handle it just fine. Heck, it might not even look all that different from what a human might have created code wise. One of the big issues we have those is that plenty of people and companies are using these tools despite having requirements that aren't met by them in the slightest. If you're working on a Google/Microsoft/Meta scale product where performance and security and code quality are of the utmost importance, then an LLM isn't going to be a great fit. Unfortunately, that's exactly where this technology is being used, and the results are becoming more and more obvious.
- datsci_est_2015 2mo agoI’ve started requiring that unless absolutely necessary, all comments and docstrings be one-liners in our code 10k+ codebase(s). LLM-generated code has a tendency to “over-justify”, which is reasonable when first reviewing the code as a human. But then, once it’s passed the first human it should be considered human-to-human communication rather than LLM-to-human communication, and the long, winding, and often repeated justifications can be significantly condensed. Also you can’t let it run loose with unit tests unless you want to be blasted in the face with hundreds of lines of absurd test fixture preparation.
- thinkharderdev 2mo ago> LLM-generated code has a tendency to “over-justify” This is the thing that drives my nuts about LLM-generated code. I will see PRs that fix a bug and the entire bug fix is re-explained in 10 different places in a code comment.
- lowbloodsugar 2mo agoI get this is a cheap response but “you’re doing it wrong”. Early last year they were shite. Then they became good enough to write tests. Then they became good enough to write code. Currently they’re good enough to design APIs for “typical” applications. Are they good enough for vibecoding a production app? No. Not yet. Don’t let that confuse you. Many of us are using them effectively, and learning what works and what doesn’t. We’re also learning how to evaluate a new model: can it do more than the last one? Does it need the same level of detail as the last one? Can we go faster?