4 ms·
How on Earth do we solve this bloat and death-by-a-thousand-cuts issue with frontier LLMs? Are there any actual solutions or attempts at solutions to this probl
by nullbio 2mo ago
How on Earth do we solve this bloat and death-by-a-thousand-cuts issue with frontier LLMs? Are there any actual solutions or attempts at solutions to this problem that I can try? Any tools or frameworks? I've tried re-architecture skills, dedicated cleanup sessions, and a bunch of other stuff, but nothing really works well.
- impulser_ 2mo agoDon't have it write the whole thing in one session.
- embedding-shape 2mo agoDon't try to "run as fast as possible" and just add code willy-nilly, think about the design, refactor as you add new features/fixes and intentionally be a bit slower and more considerate. It's not a technical problem, if you steer the LLMs enough and actually review what they do, you can build proper and clean software with them.
- nullbio 2mo agoYeah, but that's boring. :)
- kooi 2mo agoAdd Sentrux to the LLM loop. https://github.com/sentrux/sentrux https://github.com/sentrux/sentrux It's an offline coding grader. It works well: You ask for a new module, LLM starts spitting out bloated crap, the code score goes down. LLM keep looping until code score is back up. Not a substitute for human code review, but keeps things within tight guard rails.
- KingMob 2mo agoAre you the author? The font size and contrast is kind of poor.
- nullbio 2mo agoHave you done any benchmarks on this approach? I wonder if it just ends up in reward-hacking, or if it actually substantially changes the quality of the code. Are these static analysis metrics enough on their own to cause a substantial improvement? I tried setting some constraints as part of my system instruction prompt to prevent God file creation, and it did end up breaking files up, but it would often just create a new type of mess and do it needlessly. Without the model doing additional reasoning and search for alternative approaches that were better fitted on top as it was doing this file splitting, it didn't actually seem to improve quality at all. So I'm curious to know if this would just result in more of that, or if it would actually result in different choices. Because there's more to code maintenance than just splitting things up into modules, etc. Like the main part of it that I suspect will not be captured is the part of thinking about the big picture, brainstorming multiple architectural approaches, and reasoning through the best one.
- kooi 2mo agoOnly anecdotal benchmarks for what ever that is worth. I notice the LLM reasoning things like: "The maintainability score went down, let me look.... I see I've duplicated existing code which already exists in this module...I see this logic could be consolidated". I have noticed file splitting is a strategy to improve the "score". In my experience, I've been doing more the architectural guidance. There is also a rules engine which I have not used, but looks pretty interesting: ``` [constraints] max_cycles = 0 max_coupling = "B" max_cc = 25 no_god_files = true [[layers]] name = "core" paths = ["src/core/"] order = 0 [[layers]] name = "app" paths = ["src/app/"] order = 2 [[boundaries]] from = "src/app/" to = "src/core/internal/" reason = "App must not depend on core internals" ```