13 ms·
It's starting to become obvious that if you can't effectively use AI to build systems it is a skill issue. At the moment it is a mysterious, occasionally fickle
by reedf1 7mo ago
It's starting to become obvious that if you can't effectively use AI to build systems it is a skill issue. At the moment it is a mysterious, occasionally fickle, tool - but if you provide the correct feedback mechanisms and provide small tweaks and context at idiosyncrasies, it's possible to get agents to reliably build very complex systems.
- Zafira 7mo ago> At the moment it is a mysterious, occasionally fickle, tool - but if you provide the correct feedback mechanisms and provide small tweaks and context at idiosyncrasies, it's possible to get agents to reliably build very complex. This sounds like arguing you can use these models to beat a game of whack-a-mole if you just know all the unknown unknowns and prompt it correctly about them. This is an assertion that is impossible to prove or disprove.
- rafaelmn 7mo agoNo it's more like if you knew how to build it before - LLM agents help you build it faster. There's really no useful analogy I can think of, but it fits my current role perfectly because my work is constantly interrupted by prod support, coordination, planning, context switching between issues etc. I rarely have blocks of "flow time" to do focused work. With LLMs I can keep progressing in parallel and then when I get to the block of time where I can actually dive deep it's review and guidance again - focus on high impact stuff instead of the noise. I don't think I'm any faster with this than my theoretical speed (LLMs spend a lot of time rebuilding context between steps, I have a feeling current level of agents is terrible at maintaining context for larger tasks, and also I'm guessing the model context length is white a lie - they might support working with 100k tokens but agents keep reloading stuff to context because old stuff is ignored). In practice I can get more done because I can get into the flow and back onto the task a lot faster. Will see how this pans out long term, but in current role I don't think there are alternatives, my performance would be shit otherwise.
- dkdbejwi383 7mo agoYou could probably replace LLM with "junior engineer" here as it sounds like you're basically a manager now. The big negative that LLMs have in comparison with junior engineers is that they can't learn and internalise new information based on feedback.
- lukan 7mo ago"The big negative that LLMs have in comparison with junior engineers is that they can't learn and internalise new information based on feedback." No, but they can take "notes" and can load those notes into context. That does work, but is of course not so easy as it is with humans. It is all about cleaning up and maintaining a tidy context.
- rafaelmn 7mo agoI don't like that analogy. If I had to work with a Claude like junior I would ask for them to get removed from my team - inability to learn stuff, completely unexpected/unrelatable faliure modes and performance. On the other hand Claudes tenacity, stamina and sustained speed is superhuman. The more capable models become the more valuable this is.
- reedf1 7mo agoThe same is true with human engineers - isn't this just what engineering is?
- threethirtytwo 7mo ago>This is an assertion that is impossible to prove or disprove. This is a joke right? There are complex systems that exist today that are built exclusively via AI. Is that not obvious? The existence of such complex systems IS proof. I don't understand how people walk around claiming there's no proof? Really?
- mmustapic 7mo agoThe assertion was "if you really know how to prompt, give feedback, do small corrections and fix LLM errors, then everything works fine". It is impossible to prove or disprove because if everything DOES NOT work fine you can always say that the prompts were bad, the agent was not configured correctly, the model was old, etc. And if it DOES work, then all of the previous was done correctly, but without any decent definition of what correct means.
- threethirtytwo 7mo ago>And if it DOES work, then all of the previous was done correctly, but without any decent definition of what correct means. If a program works, it means it's correct. If we know it's correct, it means we have a definition of what correct means otherwise how can we classify anything as "correct" or "incorrect". Then we can look at the prompts and see what was done in those prompts and those would be a "correct" way of prompting the LLM.
- satisfice 7mo agoYou don’t know it works. That you so glibly speak about products working is proof that your engineering judgment is impaired. You can’t infer the exact contents of a black box merely by looking at outside behavior. The fundamental fallacy you are exhibiting here is similar to saying that rolling a six sided die and getting a “6” means that you will always get a 6 any time you roll it. And that if you get a 6 and wanted a 6, you must have therefore rolled those dice “correctly” and had you not gotten a 6 that would have meant you rolled them “wrong.” You know that is not true.
- stavros 7mo agoI'd agree, I've been building a personal assistant (https://github.com/skorokithakis/stavrobot https://github.com/skorokithakis/stavrobot) and I'm amazed that, for the first time ever, LLMs manage to build reliably, with much fewer bugs than I'd expect from a human, and without the repo devolving to unmaintainability after a few cycles. It's really amazing, we've crossed a threshold, and I don't know what that means for our jobs.
- Grimblewald 7mo agoNo bugs means nothing if bugs get hidden and llms are great at hiding bugs and will absolutely fail to find some fairly critical ones. Your own repo, which is slop at best, fails to meet its core premise > Another AI agent. This one is awesome, though, and very secure. it isn't secure. It took me less than three minutes to find a vulnerability. Start engaging with your own code, it isn't as good as you think it is. edit: i had kimi "red team" it out of curiosity, it found the main critical vulnerability i did and several others Severity - Count - Categories Critical - 2 - SQL Injection, Path Traversal High - 4 - SSRF, Auth Bypass, Privilege Escalation, Secret Exposure Medium - 3 - DoS, Information Disclosure, Injection You need to sit down and really think about what people who do know what they're doing are saying. You're going to get yourself into deep trouble with this. I'm not a security specialist, i take a recreational interest in security, and llm's are by no means expert. A human with skill and intent would, i would gamble, be able fuck your shit up in a major way.
- reedf1 7mo agoBuild a redteam into your feedback mechanism. Seriously. You've identified the problem and even solved it. Now automate it.
- Grimblewald 7mo agoit sure can help, but it shouldn't be considered solved. Realistically you should treat it as vulnerable until it's received some attention from natural human folks who do know what they're doing. LLM's are great but they're not experts in any field as of yet.
- Grimblewald 7mo agoSay i buy into your mysticism based take, is it a useful tool if it blows up in damn well near every professionals face? lets say i accept you and you alone have the deep majiks required to use this tool correctly, when major platform devs could not so far, what makes this tool useful? Billions of dollars and environment ruining levels of worth it? I'd say the only real use for these tools to date has been mass surveillance, and sometimes semi useful boilerplate.
- reedf1 7mo agoHonestly people are in such a weird place with this shit. I'm not saying don't read the fucking code - but I managed to get my setup to write 100k lines of indistinguishable SWE code in a week or so. The main limitation was my reading speed. This is something like a 10x speedup for me.
- Grimblewald 7mo agoHow does one verify 100k lines in a week? Let alone evaluate it to being SWE equivallent? That's super human. I like to think I am pretty good at what I do, but really critically engaging with 100k lines in a week is beyond even 10 of me. Forgive my skepticism, but I'm going to hazard the guess that you don't know what the fuck you're doing. You've lost your goddamn mind if you think you're doing anything other than skim read at a rate of 42 lines a minute for your entire work day without a break.
- mikkupikku 7mo agoOn Saturday I had claude generate ~10k of lines of Lua code which uses the libASS subtitle format to build up nearly two dozen GUI widgets from subtitle drawing primitives, including nestable scrollable containers with clipping, drop down menus, animated tab bars, and everything else I could think of. I read probably about 100 lines of code myself that day, I "verified" the code only by testing out the demo claude was updating through the process. Then on Sunday I woke up and had claude bang out a series of half a dozen projects each using this GUI library. First, a script that simply offers to loop a video when the end is reached. Updated several of my old scripts that just print text without any graphical formatting. Then more adventurous, a playlist visualizer with support for drag to reorder. Another that gives a nice little control overlay for TTS reading normal media subtitles. Another that let's people select clips from whatever they're watching, reorder them and write out an edit decision list, maybe I'll turn this one into a complete NLE today when I get home from work. Reading every line of code? Why? The shit works, if I notice a bug I go back to claude and demand a "thoughtful and well reasoned" fix, without even caring what the fix will be so long as it works. The concepts and building blocks used for all of this is shit I've learned myself the hard way, but to do it all myself would take weeks and I would certainly take many shortcuts, like certainly skipping animations and only implementing the bare minimum. The reason I could make that stuff work fast is because I already broadly knew the problem space, I've probably read the mpv manpage a thousand times before, so when the agent says its going to bind to shift+wheel for horizonal scrolling, I can tell it no, mpv has WHEEL_LEFT and RIGHT, use those. I can tell it to pump its brakes and stop planning to load a PNG overlay, because mpv will only load raw pixel data that way. I can tell it that dragging UI elements without simultaneously dragging the whole window certainly must be possible, because the first party OSC supports it so it should go read that mess of code and figure it out, which it dutifully does. If you know the problem space, you can get a whole lot done very fast, in a way that demonstrably works. Does it have bugs? I'd eat a hat if it doesn't. They'll get fixed if/when I find them. I'm not worried about it. Reading every line of code is for people writing airliner autopilots, not cheeky little desktop programs.
- jamiemallers 7mo ago[dead]
- croes 7mo ago> if you provide the correct feedback And how do you define correct feedback? If the output is correct?
- reedf1 7mo agoI don't know if you deliberately cut-off the full point, but for the benefit of those with tired eyes I said 'feedback mechanisms', i.e. feedback in the control system sense.
- croes 7mo agoThe cut-off was not intended but what is the correct one? Is it wrong as soon the result is wrong?
- reedf1 7mo agoWe are still in early days and I'm sure this isn't the best way to approach this but here is what I do. 1. Agent context with platform/system idiosyncrasies, how to access tools, this is actually kept pretty minimal - and a line directing it to the plan document. 2. A plan document on how to make changes to the repo and work that needs to be done. This is a living document pruned by the orchestrating agent. Included in this document is a directive written by you to use, update the document after ever run. Here also is a guide on benchmarking, regression, unit tests that need to be performed every time. 2a. When an agent has a code change it is then analyzed by a council of subagents, each focused on a different area, some examples, security, maintainability, system architect, business domain expert. I encourage these to be adversarial "red team". We sit in the core loop until the code changes pass through the council. 2b. Additional subagents to create documentation, build architecture diagrams etc. 2c. A suggested workflow is created on how to independently invoke testing, and subagent, etc.
- zihotki 7mo agoAccording to the https://blog.katanaquant.com/p/your-llm-doesnt-write-correct-code https://blog.katanaquant.com/p/your-llm-doesnt-write-correct... previously discussed on HN, it may be at least partially true: > The vibes are not enough. Define what correct means. Then measure.
- mikkupikku 7mo agoTruth Nuke
- bartread 7mo ago> It's starting to become obvious that if you can't effectively use AI to build systems it is a skill issue. I think it's fair to say that you can get a long way with Claude very quickly if you're an individual or part of a very small team working on a greenfield project. Certainly at project sizes up to around 100k lines of code, it's pretty great. But I've been working startups off and on since 2024. My last "big" job was with a company that had a codebase well into the millions of lines of code. And whilst I keep in contact with a bunch of the team there, and I know they do use Claude and other similar tools, I don't get the vibe it's having quite the same impact. And these are very talented engineers, so I don't think it's a skill either. I think it's entirely possible that Claude is a great tool for bootstrapping and/or for solo devs or very small teams, but becomes considerably less effective when scaled across very large codebases, multiple teams, etc. For me, on that last point, the jury is out. Hopefully the company I'm working with now grows to a point where that becomes a problem I need to worry about but, in the meantime, Claude is doing great for us.
- ptak_dev 7mo ago[flagged]
- Greed 7mo agoWould definitely tend to agree. Whenever I read complaints about accuracy of LLMs with complex systems, it has generally been from those that aren't thinking very critically about how they're using them in the first place. If you were to replace that LLM with a real human junior, would you really walk away for a few weeks and then assume the solution given was correct by default when you got it back? Obviously not. So you identify and gatekeep the most critical parts ahead of time, make error correction part of the process, and chunk the Giant, Complex Thing into Smaller, Achievable, Verifiable Things. LLMs are proving to be very much force multipliers of the kind of developer you already are, and of those who report a 10x increase in productivity they're probably all being genuine. Whether that 10x is of careful, thoughtful choices or reckless rough-shod slop though is really an artifact of the developers themselves. I've been saying from the beginning that your effectiveness with LLMs is roughly equivalent to your ability to get effective results out of a real team of human contractors.
- yogthos 7mo agoThe point is that we can make life easier for both ourselves and the agents by structuring code in a way that's conducive towards their strengths.