28 ms·
> What LLMs produce is often broken, hallucinated, or below codebase standards. With enough rules and good prompting this is not true. The code I generate is u
by jf22 1y ago
> What LLMs produce is often broken, hallucinated, or below codebase standards.
With enough rules and good prompting this is not true. The code I generate is usually better than what I'd do by hand.
The reason the code is better all the extra polish and gold plating is essentially free.
Everything I generate comes out commented great error handling, logging, SOLID, and united tested using established patterns in the code base.
- tobyhinloopen 1y agoI agree. With plenty of prompts (leave them in documents) you can get pretty good results.
- drgiggles 1y agoEasily 99% of comments generated by LLMs are useless.
- exe34 1y agoThey comment on the how, not the why.
- micromacrofoot 1y agoThey're often repetitive if you're reading the code, but they're useful context that feeds back into the LLM. Often once the code is clear enough I'll delete them before pushing to production.
- mrugge 1y agodo you have proof of this being useful for llm? wouldn't you rather it re-read the actual code it generated instead of assuming that the potentially wishful thinking or stale comment is going to lead it astray?
- micromacrofoot 1y agoit reads both, so with the comments it more or less parrots the desired outcome I explained... and it sometimes catches the mismatch between code and comment itself before I even mention it I read and understand 100% of the code it outputs, so I'm not so worried about falling too far astray... being too prescriptive about it (like prompting "don't write comments") makes the output worse in my experience
- the__alchemist 1y agoI've noticed this too. They are often restatements of the line in verbal form, or intended for me, the LLM-reader about the prompt, vice a code maintainer.
- shortrounddev2 1y agoI think its because LLMs are often trained with data from code tutorial sites and forums like stackoverflow, and not always production code
- ggregoire 1y agoThat's how I detect who is using LLMs at work. # loop over the images for filename in images_filenames: # download the image image = download_image(filename) # resize the image resize_image(image) # upload the image upload_image(image)
- rob_c 1y ago99% of comments are not needed as they just re-express what the code below does. I prefer to push for self documenting code anyway, never saw the need for docs other than for an API when I'm calling something like a black box.
- cjfd 1y agoVery often comments generated by humans are also useless. The reason for this are mandated comment policies, e.g., 'every public method should have a comment'. An utterly disgusting practice. One should only have a comment if one has something interesting to say. In a not-overly-complex code base there should maybe be a comment perhaps every 100 lines or so. In many cases it makes more sense to comment the unit tests than the code.
- drgiggles 1y agoI am pretty far to one end of the spectrum on need for comments. Very rarely is a comment useful to help you/another developer decipher the intent and function of a piece of code.
- skydhash 1y agoI think the rules for comments on public method is to use something like doxygen to extract the reference. And most IDE can display them upon hovering. And comments can remind the caller of pre- and post-conditions.
- gddgb 1y ago[dead]
- jf22 1y agoThen tell it to write better comments...
- timmytokyo 1y agoAh, so it's good enough to write code on its own without time-consuming, excessive hand-holding. But it's not good enough to write comments on its own.
- jf22 1y agoIf you put in the work to write rules and give good prompts you get good results just like every other tool created by mankind. How often do you use coding LLMs?
- ewoodrich 1y agoI can't speak to comments rules specifically but I am a heavy user of "agentic" coding and use rules files and while they help they are simply not that reliable. For something like comments that's probably not that big of a deal because some extra bad comments isn't the end of the world. But I have rules that are quite important for successfully completing a task by my standards and it's very frustrating when the LLM randomly ignores them. In a previous comment I explained my experiences in more detail but depending on the circumstances instruction compliance is 9/10 times at best, with some instructions/tasks as poor as 6/10 in the most "demanding" scenarios particularly as the context window fills up during a longer agentic run.
- wglb 1y agoNot what I have found with gemini. What is particularly useful is the comments about reasoning about new code added at my request.
- reactordev 1y agoLet’s talk about rules and docs, shall we? What makes a good rule for AI to keep it on task? What are your setups for docs and attaching them to the context (do you need to? Or just the location?) Let’s boil this down to an easy set of reproducible steps any engineer can take to wrangle some sense from their AI trip.
- rob_c 1y agoAka, let's train people how to use the tool...
- exe34 1y agoLet's check that the claim matches the evidence first!
- reactordev 1y agoYou seem to be against the idea. Yet you yourself were trained. Weird.
- jf22 1y agoThere are tons of these guides around the internet. I'm only using what other people have already published.
- icey 1y agoThe company I work at (https://getunblocked.com https://getunblocked.com) is built to give tools like Claude Code and Cursor context based on all your docs, issues, code, and chat threads from Slack and soon Teams. Happy to give you a demo sometime if you're interested!
- symfrog 1y agoDo you have a link to some of the code that you have produced using this approach? I am yet to see a public or private repo with non-trivial generated code that is not fundamentally flawed.
- rob_c 1y agoThat's disingenuous or naive. Almost nobody decides to expressly highlight the section of code (or whole files generated by ai) they just get on with the job when there's real deadlines and it's not about coding for the sake of the art form...
- symfrog 1y agoI am talking about correctness, not style, coding isn't just about being able to show activity (code produced), but rather producing a system that is correctly performing the intended task
- rob_c 1y agoYes, and frankly you should be spending time writing large integration tests correctly not microscopic tests that forgot how tools interact. It's not about lines of code or quality it's about solving a problem. If the problem creates another problem then it's bad code. If it solves the problem without causing that then great. Move onto the next problem.
- jakelazaroff 1y agoIf the generated implementation is not good, you're trading short-term "getting on with the job" and "real deadlines" for mid-to-long-term slowdown and missed deadlines. In other words, it matters whether the AI is creating technical debt.
- rob_c 1y agoIf you're creating technical debt, you're creating technical debt. That has nothing to do with AI/LLMs. If you can't understand what the tool spits out either; learn, throw it away, or get it to make something you can understand.
- the__alchemist 1y agoI am pattern matching your last statement with what I've seen with my teammates who are more AI-oriented: I suspect this is a matter of making the metrics the goal. I would rather maintain something that is simple, works, and have targeted comments than something messy that meets the metrics you list.
- echelon 1y agoI don't get all the prompt vibe coding going around. I don't use prompts to generate code. I use "tab-tab" auto complete to speed through refactorings and adding new fields / plumbing. It's easily a 3x productivity gain. On a good day it might be 10x. It gets me through boring tedium. It gets strings and method names right for languages that aren't statically typed. For languages that are statically typed, it's still better than the best IDE AST understanding. It won't replace the design and engineering work I do to scope out active-active systems of record, but it'll help me when time comes to build.
- recursive 1y agoI use tab auto complete, and i think it's a 5% productivity gain. On a good day, maybe 10%. I haven't put much effort into optimizing the setup or learning advanced usage patterns or anything. I'm using stock copilot, provided by my employer. If I had to pay for it, I wouldn't be using it, as it doesn't justify the cost.
- Bootvis 1y agoReally, what are you making that a 5% increase in productivity doesn’t justify a Copilot subscription?
- recursive 1y agoThat's not a rigorously measured number. The 5% is an increase in straight-ahead code speed. I spend a small fraction of my time typing code. Smaller than I'd like. And it very well might be an economically rational subscription. For me personally, I'm subscription averse based on the overhead of remembering that I have a subscription and managing it.
- abtinf 1y agoCan you point to an example repo with enough rules and good prompts?
- mrugge 1y agoFirst thing I do is tell llm to stop writing useless docstrings and comments and instead follow clean code principles where each variable is a noun and function call a verb.
- drgiggles 1y ago[flagged]
- deleted 1y ago[deleted]
- mrbungie 1y ago> The code I generate is usually better than what I'd do by hand. I'm always baffled by this. If you can't do it that well by hand, how can you discriminate its quality so confidently? I get there is a artist/art consumer analogy to be made (i.e. you can see a piece is good without knowing how to paint), but I'm not convinced it is transferrable to code. Also, not really my experience when dealing with IaC or (complex) data related code.
- tptacek 1y agoWhat an odd question. For the exact same reason people who write prose professionally usually have someone else edit their work: because editing your own work is harder, and everybody slips up sometimes.
- dingnuts 1y agoehhhhhhh yeah but this is like hiring Reddit to do your prose editing, considering generated code is slightly worse than what you'd find on r/programming
- tptacek 1y agoYou can believe that or not believe that without changing the implication of the previous question, which was that someone who routinely slips while writing code would be incapable of determining whether the LLM got it right. Obviously not.
- mrbungie 1y agoI'm not getting this analogy. Editors can't normally discriminate if the content itself is good (after all, the writer is the SME), but rather, only perfect its form (syntax, grammar, etc). Well-written bullshit in perfect prose is still bullshit.
- halfmatthalfcat 1y agoI didn't find it odd at all and it seems more odd to liken an LLM to a human editor.
- micromacrofoot 1y agoI've been finding actual human-written bugs and correcting them with Claude, so I find the "often broken" claims a load of nonsense... I've been fixing dozens of minor bugs in our codebase that no one's been arsed to fix for years due to bigger priorities (which tbh is generating more features and tech debt). It may change in the future, but AI is without a doubt improving our codebase right now. Maybe not 10X but it can easily 2X as long as you actually understand your codebase enough to explain it in writing.
- intended 1y agoCould you share an example ? These conversations on AI code good, vs AI code bad constantly keep cropping up. I feel we need to build a cultural norm to share examples places of succeeded, and failures, so that we can get to some sort of comparison and categorization. The sharing also has to be made non-contentious, so that we get a multitude of examples. Otherwise we’d get nerd-sniped into arguing the specifics of a single case.
- r3trohack3r 1y agoI do think a lot of the discourse in this space can be summed up as: people are arguing about two non-overlapping segments of a distribution having no idea the other segment even exists; instead they just assume the other side is [hype/pessimistic].
- micahscopes 1y agoIt makes me wince a little
- zoeysmithe 1y agoWhat a scary time it is for devs. We spent all this time learning this obscure skill and now when I play with claude or even chatgpt it makes really good code. I just asked it to write me a video game and it did it. Perfect godot code. I was stunned it didn't hallucinate and when I asked for clarification on a snippet of code, it perfectly answered. I think its only a matter of time until our roles are commoditized and vibe-coding becomes the norm in most industries. Vibe coding being a dismissive term on developing a new skillset. For example we'll be doing more planning and testing and such instead of writing code. The same way, say, sysadmins just spin up k8s instead of racking servers or car mechanics read diagnosis codes from readers and, often, just replace an electric part instead of hand-tuning carbs or gapping spark plugs and such. That is to say, a level of skill is being abstracted away. I think we just have to see this, most likely, as how things will get done going forward.
- tovej 1y agoCould you at least mention what the video game was, or why it was such a good implementation? Also, what was "perfect" about the code? "Perfect" is not a word I would ever use to describe code. This reads like empty hype to me, and there's more than one claim like this in these threads, where AI magically creates an app, but any description of the app itself is always conspicuously missing.
- zoeysmithe 1y agoYes Im exaggerating and its not writing a AAA game from a prompt but I asked it to make a game like Zelda and it figured it out and walked me through all the aspects of it. That's a lot more than I expected. I'm not a games programmer, so I'm probably a lot more impresse than I should be, but I went from not knowing anything about godot to having a framework up to build a 2d rpg-esque game fairly quickly and me learning as it gave me the code. Note, I used the new chatgpt study mode, so that's may be different than just regular prompts. I fully expected just broken code and random AI musings, but instead I got a very solid implementation of a game, albeit a simple one. Or at least as simple as I asked for, I imagine I can keep building out more with its help. I also have never used godot before, and I was surprised at how well it navigated and taught me the interface as well. At least the horror stories about "all the code is broken and hallucinations" isn't really true for me and my uses so far. If LLM's will succeed anywhere it will be in the overly logical and predictable worlds of programming languages, but that's just a guess on my part, but thus far whenever I reach out for code from LLM's, its been a fairly positive experience.
- tempodox 1y agoSo you're not a developer any more but a tenant who pays rent. Speaking of which, I have a bridge to sell you…
- jf22 1y agoI'm not even sure what this is supposed to say.
- neutronicus 1y agoYeah that fucker Claude is tireless when it comes to checking return types, checking for null, etc etc
- nurettin 1y agoHere's my workflow (if I feel like using claude) Me: Here's the relevant part of the code, add this simple feature. Opus: here's the modified code blah blah bs bs Me: Will this work? Opus: There's a fundamental flaw in blah bleh bs bs here's the fix, but I only generate part of the code, go hunt for the lines to make the changes yourself. Me: did you change anything from the original logic? Opus: I added this part, do you want me to leave it as it was? Me: closes chat
- NitpickLawyer 1y agoSorry to be that guy, but you're using it wrong. The best flows right now are architect -> act -> test. First you have a session in "architect" / "plan" mode (depending on your ide/tool) where you discuss, ask questions, etc. Then, when everything is clear in "chat" mode, you ask the model to make a plan. You verify the plan, and then you tell it to start implementing it. You still get to approve tools, calls, tests, etc. You can also provide feedback on the way if you missed something (i.e. use uv instead of pip, etc). Coding in a chat interface, and expecting the same results as with dedicated tools is ... 1-1.5 years old at this point. It might work, but your results will be subpar.
- nurettin 1y agoNah it's good thanks for your input. I saw people use plan.md and todo.md and ide/commandline for this before. manus.ai demonstrates this via its chat interface as well.
- apwell23 1y ago> With enough rules and good prompting this is not true. There are atleast 10 posts on HN these days with the same discussion in circle. 1. AI sucks at code 2. you are not using my magic prompting technique
- jf22 1y agoIt's not magic. The techniques are well established and widely shared.
- apwell23 1y ago[flagged]
- swader999 1y agoYeah there's so many now it's hard to settle on one. YouTube is littered with them. Agent OS, amp.code, BMAD. I'm probably trying BMAD in earnest next ...
- jf22 1y agoEach of the "tools" does things slightly differently but the techniques to use them effectively are largely the same now (rules, planning, context management, good prompting). You know like when the loom came out there were probably quite a few models but using it was similar. Like cars are now.
- miggy 1y agoIn my experience, unit tests and logging code generated by LLMs tend to be overly verbose, miss meaningful assertions, and often produce boilerplate that looks correct but doesn’t test or log anything useful. It’s easy to get misled by the surface structure.