5 ms·
> To this end, LLMs will be used extensively to deslop... So we're using LLMs to clean up the code that LLMs ruined in the first place? We’ve reached peak tech
by wsdn 2mo ago
> To this end, LLMs will be used extensively to deslop...
So we're using LLMs to clean up the code that LLMs ruined in the first place? We’ve reached peak tech in 2026.
- maccard 2mo agoI’d say I fall in the “AI skeptic but willing to use it” category. If LLMs can actually clean up after themselves it would actually be a game changer.
- figmert 2mo agoThey can. You just have to steer them
- maccard 2mo agoCan you share a transcript of an LLM generated feature/bug fix that you steered into being good quality?
- onlyrealcuzzo 2mo agoI'm about to release some tooling that's been very effective for me. Essentially, there's a few ways LLMs write "bad" code that is different from how people write "bad" code. We've got pretty good tooling to catch the ways people write bad code - it just happens to be much easier to do with static analysis (and is less noise prone). The ways LLMs write bad code is typically 1) bad architecture - hard to detect in the ways that are really important, 2) unnecessary state and control flow (and decisions based on state), 3) bad / inadequate tests. Methods to detect these problems have existed for ages, but they've never caught on because it's typically too difficult to tune them to have high signal / noise for humans, and AFAIK - no one else tried putting them all together and seeing how LLMs work with it. LLMs are great at sorting through signal / noise -> so you can help surface potential issues with metrics that would be too noisy for humans, but seems to work pretty well for LLMs to find the source of architectural problems and design better solutions (from my experience - may be biased, I built the tooling to literally solve this problem for the main project I'm working on).
- mirekrusin 2mo agoYou forgot to include link to your tool.
- onlyrealcuzzo 2mo agoAn LLM could probably figure out how to use it now. But it doesn't yet have a coherent UX unless you're me. Hopefully, I'll iron that out over the next week and I'll update you.
- mirekrusin 2mo agoI’ll keep checking.
- all2 2mo agoIf I had to do a 'code quality' checker I think I'd try to use some combination of syntax tree analysis, data flow analysis, and LOC changed. Something like that. Too many new nodes in the syntax tree, or a sub-tree that appears sufficiently similar to another sub-tree (for various definitions of similar), data-flow/side effects gets more convoluted, too many LOC, and so on.
- fragmede 2mo agohttps://ponytail.dev/ https://ponytail.dev/
- andai 2mo agoCan you elaborate on the tests thing? I'm new to testing, and AI agents recently voluntarily added an large number of tests to one of my projects. I've been learning a lot by reading them, and it seems like a great habit to develop. But they're also writing some very strange code and some very strange tests. I don't know what good practices look like here, so when something looks strange to me I can't trust my own judgment, whether it's actually smelly or just a pattern I'm not used to yet. -- On a side note, I recently had an agent implement a major architectural change. It turned out to have done it completely backwards, in a way that was pointless. (Improved nothing and actively made things worse.) However it had supplied generous tests for the new code, and of course all the tests passed... So it had "proven the correctness" of something which was completely incorrect. I later realized that even formal verification would not have prevented this. It would have just written a mathematical proof that the wrong code was correct.
- jerf 2mo agoLLMs can clean up after themselves if you steer them. It's not very hard. I often have a two-step dance I do where the LLM first outputs some code and then I prompt it to fix the types up to my standards. I haven't found a way to prompt it with any number of skills or CLAUDE.mds or anything else to get it to do it the way I want on the first pass, but it's not that hard to just fix it afterwards. It's a fast enough process that it feels fine using it. One can even make a case for it being a decent way to operate anyhow; make a sketch, then firm it up isn't entirely unlike my manual coding process was anyhow. The main problem is that pointing an LLM at a codebase and telling it to just "make it better" taps out pretty quickly. That is, not that there's zero juice to squeeze there, but there's not a ton. You can get a bit of improvement but it also rapidly starts changing things just to change things, which I'm not even going to complain about all that much because there's a sense in which it is simply doing as you asked. So you still need human taste and direction. This will be especially true for something the size of that codebase where you can only hold small fractions of the actual code in the context window at once. Summaries only get you so far.
- andai 2mo agoI asked an agent for some high level cleanup and refactoring. I found that I ended up disagreeing with most of the structural changes. Some of them were necessary, some were beneficial, but largely it turned straight line code into abstract factory manager type stuff. More broadly I've found that making code more elegant (e.g. by removing duplication) increases the cognitive load, because now you can't just read the code anymore but need to mentally "decompress" the higher level structures and indirection into the straight line code, the "code that actually runs."
- jerf 2mo agoIf it wasn't clear, I don't ask the agent to "clean up my code". I ask it to "take this map[string]string you pass around with constant keys and turn that into a structure" or "extract this API provider out into an interface and make everything using it use this interface instead" or other concrete instructions. LLMs, to a first approximation, already did as well as they could on the first pass. You can get a bit more out of them by asking them to just try harder, but not much. Whereas if you go in with specific changes they are pretty good at implementing them. This especially matters because my personal style deviates from the common practices. This may also be why I can't get it to just happen by prompting for it. But if you walk it through it's perfectly capable of transforming the code into a better style, and it's still way faster than trying to write it from scratch.
- owaislone 2mo agoYou have to treat it like a co-worker send you PRs for review. You review, ask to improve/cleanup stuff and it'll follow up. You can even ask it to remember so it follows the patterns in future. Even better, if you work with it to come up with a design first and then send it coding, you often don't need as much cleanup as the desired architecture is agreed upon early.
- maccard 2mo agoI would tell my co worker “this code has so many issues that I’m not actually reviewing it. Please make sure it meets the project standards before sending it back to me”if it happened twice I’d talk to their manager.
- owaislone 2mo agoSure - define your project standards in the repo or agent's skill and tell it just that.
- maccard 2mo agoSee my other comment here [0] - I’ve not found it possible to actually have the agents follow the explicit instructions. Even single shot tasks once the context window fills up (which it does, very very fast, and repeatedly if you use subagents) will ignore very explicit instructions. [0] https://news.ycombinator.com/item?id=49036222 https://news.ycombinator.com/item?id=49036222
- owaislone 2mo ago- discuss the plan, dos-and-donts, instructions and have them add it to a PLAN_FEATURE.md with the goal, non-goal, non-negotiables, technical details on the top. - come up with phases to implement your thing. P1, P2, ... P8.. whatever - in each phase add simple one line TODOs with as much explanantion of the task as needed e.g, - [ ] Update DB schema, - [ ] Generate migrations - Ask the agent to go one phase or todo at a time depending on complexity. - if the context fills up, /clear the window and just ask it to do the next line. - if you wrote up all you care about in the plan properly, it'll follow what you need it to do.
- nateb2022 2mo agoI'm assuming based on the special unicode angled quote marks that this is copy/pasted from an AI? I don't understand why people would type a lot more into an AI just to copy a comment this short, which would be at least 10x faster to just type by hand.
- defen 2mo agoNot sure about op but if I post here from my phone I get those. Example: “quoted string”
- nateb2022 2mo agohuh interesting, I've never realized this, is it iOS or android?
- defen 2mo agoiOS
- thejazzman 2mo agoIt’s a macOS thing too. I turn it off because it burns me when copying code
- maccard 2mo agoOP here - no AI used in the writing of any of my replies on this site. iOS does this when quoting [0]. I don’t think it’s in the spirit of this site to jump on me rather than the content of my comment either way. [0] https://apple.stackexchange.com/questions/299505/ios-11-default-quotation-mark-changed-to-and https://apple.stackexchange.com/questions/299505/ios-11-defa...
- axod 2mo agoLLMs are as good as the operator directing them.
- epolanski 2mo agoI'd say as thorough rather than good. Skilled engineers may still vibe and not care. Beginners may be thorough and experiment and ask till they get the design right even if they don't spot it immediately, just caring does a lot.
- left-struck 2mo agoIncidentally I’ve always thought that about working on cars, but more like skilled mechanics may still vibe and not care. Amateurs may be thorough but much slower, may need to do a lot of research before every step. Just caring does a lot.
- pseudocomposer 2mo agoI agree with you about “thoroughness” being key here, but I’d additionally advocate that we need to leave room for people to make some slop learn it. Part of the learning process, whether you’re a “skilled” or “unskilled” engineer (or not an engineer at all, and noting that going from one to the other is just a matter of learning), is being able to “vibe and not care.” Yes, it’s also helpful to identify things that might not work out theoretically in advance, using the things we learn in our CS programs. But coding with LLMs is a brand new modality, and we are all indeed just learning to work with them. I think alternating vibe coding with vibe-assisted DRYing/cleaning is ultimately a workflow enhancer that can make better software faster. But engineers have to be allowed to make some slop to learn it.
- lionkor 2mo agoThe error bars on that one are too large to fit on screen, but yes.
- baq 2mo agoOnly if you have a tiny screen. You can make them output what you would've written yourself, but that would be slower than writing it yourself, so no one does that - but it’s trivially true.
- epolanski 2mo agoThis is the present and future, willingly or not. Engineering is about design and plan and architectural integrity and qa. Not writing code anymore.
- onlyrealcuzzo 2mo ago> Not writing code anymore. If you were an L6+, it was already not really about that. It's kind of amazing to me that people think the only thing engineers do and the only value they bring is writing code.
- williamdclt 2mo agoWell so far we've been using humans to clean up the code that humans ruined in the first place, it's hardly a logical fallacy
- giancarlostoro 2mo agoThe pearl clutching about AI is insane to me, as if code quality was perfect before 2022. We have been seeing a decline in software quality for a long time coming now. A lot of AI hatred is not even based on seeing something genuinely bad, its just prejudice for the sake of prejudice. If I could view into two parallel worlds, one where I vibe code something, release it and tell people it's vibe coded, and another where I do the same but tell no one, codebases are both the same and I still reviewed the code, had the model fix issues with code quality and bugs, and genuinely invested hours of my time to raise up, in the universe where I don't mention AI, people will likely sing praises, in the universe where I do mention AI, people will not even review or view a single line of code, they will close the tab, hop back on HN and complain that it's AI slop, without any reasonable reason. Blind AI hatred is just insane to me and it goes against the spirit of HN, we're not supposed to just go "ITS BAD CODE" here, we are supposed to have actual engaging conversations. Too much of HN devolves into "ITS SLOP" instead of "Hey, I know you're leaning on the AI a bit, but here's what's wrong with the code: ...." which would be more constructive and fall better in line with what I've expected from HN and seen for years, until recently. AI fatigue and AI prejudice should not be confused for one another, sure you can have both, but most people who hate AI blindly just have AI prejudice at this point.
- yunwal 2mo ago> "Hey, I know you're leaning on the AI a bit, but here's what's wrong with the code: ...." I'm generally in agreement, but I think current corporate/productivity culture and AI is a bad mix that makes this quote/sentiment feel impossible. All of the current incentives are set up to push out as much sloppy code as possible, and AI is the perfect machine for doing that in huge quantities, to the point where reviewing all of it is impossible. It's not really AI's fault (imo). There's plenty of ways to use it to design more elegant, understandable, maintainable systems. But unfortunately that's going to take a cultural shift, and part of that cultural shift involves people pushing back against AI as it's used today.
- serial_dev 2mo agoYou have no idea, to slop-circle is getting more common. In 2025, sometimes I decided to put a slop LinkedIn post to ChatGPT (back then) to deslopify. Now, the slop circle is everywhere. Your PO uses automations to create tasks? Half of the comments are bot slop? Links to documents never edited or read by another human? A lazy three sentence description of a task requirement you would have gotten previously sounds like a dream. Now everything generates pages long texts and you can't possibly go through the slop without AI agents. You need your agents comb through the slop and demystify the task for you. Same with the other end... Slop PRs will be reviewed by your agents, hoping you catch 2-3 issues so that you can pretend you did a review. The author didn't do a review on their generated slop, but sure, let's pretend reviewers still review code.
- f17428d27584 2mo agoThis was always the plan. Code is no longer written by humans, code is LLLM output of broken slop that is mostly software-shaped, impossible for humans to read and make sense of it (because it makes no sense). So the only solution is that all code is intended to be read and written by machines. The human-in-the-loop era was always a stopgap.
- nullbio 2mo agoYeah, good luck with that one though. LLMs are terrible at deslop, or they wouldn't slop in the first place.
- fhd2 2mo agoNot entirely my experience. LLMs can be good at two things, but unable to be good at them _at the same time_. I've had some success with getting decent simplification suggestions out of them. Still ironic, of course.
- andai 2mo agoI came here to post the same thing. But this is addressed later in the same paragraph: > But hopefully better development practices, with a human in the driver’s seat, and a focus on reducing technical debt and writing idiomatic Zig, mean that in a few weeks or months there will be a presentable codebase that serves as a drop-in replacement for Rust Bun 1.4.0. So the intention here is simply more steering. Or perhaps better steering (code structure appears to be a matter of taste... I had a very perplexing chat with a friend yesterday who insisted that four backend processes were required to serve a single HTTP request...)
- jazzzooo 2mo ago[dead]