23 ms·
LLM code generation may lead to an erosion of trust
- gblargg 1y agohttps://archive.is/5I9sB https://archive.is/5I9sB (Works on older browsers and doesn't require JavaScript except to get past CloudSnare).
- cheriot 1y ago> promises that the contributed code is not the product of an LLM but rather original and understood completely. > require them to be majority hand written. We should specify the outcome not the process. Expecting the contributor to understand the patch is a good idea. > Juniors may be encouraged/required to elide LLM-assisted tooling for a period of time during their onboarding. This is a terrible idea. Onboarding is a lot of random environment setup hitches that LLMs are often really good at. It's also getting up to speed on code and docs and I've got some great text search/summarizing tools to share.
- bluefirebrand 1y ago> Onboarding is a lot of random environment setup hitches Learning how to navigate these hitches is a really important process If we streamline every bit of difficulty or complexity out of our lives, it seems trivially obvious that we will soon have no idea what to do when we encounter difficulty or complexity. Is that just me thinking that?
- kmoser 1y agoThere will always be people who know how to handle the complexity we're trying to automate away. If I can't figure out some arcane tax law when filling out my taxes, I ask my accountant, as it's literally their job to know these things.
- bluefirebrand 1y ago> There will always be people who know how to handle the complexity we're trying to automate away This is not a given! If we automated all accounting, why would anyone still take the time to learn to become an accountant? Yes, there are sometimes people who are just invested in learning traditional stuff for the sake of it, but is that really what we want to rely on as the fallback when AI fails?
- kmoser 1y agoIt's highly unlikely that everybody will flock to LLMs, leaving absolutely nobody capable of stringing together a few lines of code on their own. Some devs may enjoy vibe coding, and may even be more productive that way, but there will always be use cases where it is preferable to produced deterministic code via a human dev.
- RunningDroid 1y ago> > Onboarding is a lot of random environment setup hitches > > Learning how to navigate these hitches is a really important process To add to this, a barrier to contribution can reduce low quality/spam contributions. The downside is that a barrier to contribution that's too high reduces all contributions.
- cheriot 1y agoSome people find a solution, think about it, and incorporate that into their understanding of the world. Some people ask for a coworker to do it for them, c/p stackoverflw, etc and never learn. AI makes the first group that much more effective. Setup should be a learning process not a hazing ritual.
- namenotrequired 1y ago> LLMs … approximate correctness for varying amounts of time. Once that time runs out there is a sharp drop off in model accuracy, it simply cannot continue to offer you an output that even approximates something workable. I have taken to calling this phenomenon the "AI Cliff," as it is very sharp and very sudden I’ve never heard of this cliff before. Has anyone else experienced this?
- sandspar 1y agoI'm not sure. Is he talking about context poisoning?
- Kuinox 1y agoI'm doing my own procedurally generated benchmark. I can make the problem input bigger as I want. Each LLM have a different thresholf for each problem, when crossed the performance of the LLM collapse.
- Paradigma11 1y agoIf the context gets to big or otherwise poisoned you have to restart the chat/agent. A bit like windows of old. This trains you to document the current state of your work so the new agent can get up to speed.
- bubblyworld 1y agoI've only experienced this while vibe coding through chat interfaces, i.e. in the complete absence of feedback loops. This is much less of a problem with agentic tools like claude code/codex/gemini cli, where they manage their own context windows and can run your dev tooling to sanity check themselves as they go.
- Syzygies 1y agoOne can find opinions that Claude Code Opus 4 is worth the monthly $200 I pay for Anthropic's Max plan. Opus 4 is smarter; one either can't afford to use it, or can't afford not to use it. I'm in the latter group. One feature others have noted is that the Opus 4 context buffer rarely "wears out" in a work session. It can, and one needs to recognize this and start over. With other agents, it was my routine experience that I'd be lucky to get an hour before having to restart my agent. A reliable way to induce this "cliff" is to let AI take on a much too hard problem in one step, then flail helplessly trying to fix their mess. Vibe-coding an unsuitable problem. One can even kill Opus 4 this way, but that's no way to run a race horse. Some "persistence of memory" harness is as important as one's testing harness, for effective AI coding. With the right care having AI edit its own context prompts for orienting new sessions, this all matters less. AI is spectacularly bad at breaking problems into small steps without our guidance, and small steps done right can be different sessions. I'll regularly start new sessions when I have a hunch that this will get me better focus for the next step. So the cliff isn't so important. But Opus 4 is smarter in other ways.
- beau_g 1y agoThe article opens with a statement saying the author isn't going to reword what others are writing, but the article reads as that and only that. That said, I do think it would be nice for people to note in pull requests which files have AI gen code in the diff. It's still a good idea to look at LLM gen code vs human code with a bit different lens, the mistakes each make are often a bit different in flavor, and it would save time for me in a review to know which is which. Has anyone seen this at a larger org and is it of value to you as a reviewer? Maybe some tool sets can already do this automatically (I suppose all these companies report the % of code that is LLM generated must have one if they actually have these granular metrics?)
- acedTrex 1y agoAuthor here: > The article opens with a statement saying the author isn't going to reword what others are writing, but the article reads as that and only that. Hmm, I was just saying I hadn't seen much literature or discussion on trust dynamics in teams with LLMs. Maybe I'm just in the wrong spaces for such discussions but I haven't really come across it.
- DyslexicAtheist 1y agoit's really hard using AI (not impossible) to produce meaningful offensive security to improve defense due to there being way too many guard rails. While on the other hand real nation-state threat actors would face no such limitations. On a more general level, what concerns me isn't whether people use it to get utility out of it (that would be silly), but the power-imbalance in the hand of a few, and with new people pouring their questions into it, this divide getting wider. But it's not just the people using AI directly but also every post online that eventually gets used for training. So to be against it would mean to stop producing digital content.
- davidthewatson 1y agoWell said. The death of trust in software is a well worn path from the money that funds and founds it to the design and engineering that builds it - at least the 2 guys-in-a-garage startup work I was involved in for decades. HITL is key. Even with a human in the loop, you wind up at Therac 25. That's exactly where hybrid closed loop insulin pumps are right now. Autonomy and insulin don't mix well. If there weren't a moat of attorneys keeping the signal/noise ratio down, we'd already realize that at scale - like the PR team at 3 letter technical universities designed to protect parents from the exploding pressure inside the halls there.
- tomhow 1y ago[Stub for offtopicness, including but not limited to comments replying to original title rather than article's content]
- michelsedgh 1y ago[flagged]
- incompatible 1y agoThe LLMs may be useful some day, if somebody can figure out how to stop them randomly generating complete garbage.
- petesergeant 1y agoIf you can't make LLMs do useful work for you today, that's on you.
- michelsedgh 1y agoI'm not trying to advertise anything here, but i'm telling you i can create apps that feel apple level quality and its not garbage. I built this in 2 weeks I had never known what swift is or anything. I wasn't a coder or app maker, but now I can build something like this in 2 weeks. Let me know if its garbage. I'll love criticism and I'll look for something else if its actually garbage.
- incompatible 1y agoSorry, I don't have an iphone or macos.
- incompatible 1y agoI'd be wary of LLM-generated code, even if it seems to work in test cases. Are you sure it doesn't go wrong in edge cases, or have security problems? If you don't even know the language, you can't do any meaningful code review.
- stavros 1y agoI don't understand the premise. If I trust someone to write good code, I learned to trust them because their code works well, not because I have a theory of mind for them that "produces good code" a priori. If someone uses an LLM and produces bug-free code, I'll trust them. If someone uses an LLM and produces buggy code, I won't trust them. How is this different from when they were only using their brain to produce the code?
- moffkalast 1y agoIt's easy to get overconfident and not test the LLM's code enough when it worked fine for a handful of times in a row, and then you miss something. The problem is often really one of miscommunication, the task may be clear to the person working on it, but with frequent context resets it's hard to make sure the LLM also knows what the whole picture is and they tend to make dumb assumptions when there's ambiguity. The thing that 4o does with deep research where it asks for additional info before it does anything should be standard for any code generation too tbh, it would prevent a mountain of issues.
- stavros 1y agoSure, but you're still responsible for the quality of the code you commit, LLM or no.
- moffkalast 1y agoOf course you are, but it's sort of like how people are responsible their Tesla driving on autopilot, which then suddenly swerves into a wall and disengages two seconds before impact. The process forces you to make mistakes you wouldn't normally ever do or even consider a possibility.
- JohnKemeny 1y agoTo add to devs and Teslas, you have journalists using LLMs writing summaries, lawyers using LLMs writing dispositions, doctors using LLMs writing their patient entries, and law enforcement using LLMs writing their forensics report. All of these make mistakes (there are documented incidents). And yes, we can counter with "the journalists are dumb for not verifying", "the lawyers are dumb for not checking", etc., but we should also be open for the fact that these are intelligent and professional people who make mistakes because they were mislead by those who sell LLMs.
- axegon_ 1y agoThat is already the case for me. The amount of times I've read "apologies for the oversight, you are absolutely correct" is staggering: 8 or 9 out of 10 times. Meanwhile I constantly see people mindlessly copy paying llm generated code and subsequently furious when it doesn't do what they expected it to do. Which, btw, is the better option: I'd rather have something obviously broken as opposed to something seemingly working.
- autobodie 1y agoIn my experience, LLMs are extremely inclined to modify code just to pass tests instead of meeting requirements.
- fwip 1y agoWhen they're not modifying the tests to match buggy behavior. :P
- devjab 1y agoAre you using the LLM's through a browser chatbot? Because the AI-agents we use with direct code-access aren't very chatty. I'd also argue that they are more capable than a lot of junior programmers, at least around here. We're almost at a point where you can feed the agents short specific tasks, and they will perform them well enough to not really require anything outside of a code review. That being said, the prediction engine still can't do any real engineering. If you don't specifically task them with using things like Python generators, you're very likely to have a piece of code that eats up a gazillion memory. Which unfortunately don't set them appart from a lot of Python programmers I know, but it is an example of how the LLM's are exactly as bad as you mention. On the positive side, it helps with people actually writing the specification tasks in more detail than just "add feature". Where AI-agents are the most useful for us is with legacy code that nobody prioritise. We have a data extractor which was written in the previous millennium. It basically uses around two hunded hard-coded coordinates to extact data from a specific type of documents which arrive by fax. It's worked for 30ish years because the documents haven't changed... but it recently did, and it took co-pilot like 30 seconds to correct the coordinates. Something that would've likely taken a human a full day of excruciating boredom. I have no idea how our industry expect anyone to become experts in the age of vibe coding though.
- atemerev 1y agoI am a software engineer who writes 80-90% code with AI (sorry, can't ignore the productivity boost), and I mostly agree with this sentiment. I found out very early that under no circumstances you may have the code you don't understand, anywhere. Well, you may, but not in public, and you should commit to understanding it before anyone else sees that. Particularly before sales guys do. However, AI can help you with learning too. You can run experiments, test hypotheses and burn your fingers so fast. I like it.
- pfdietz 1y agoThere was trust?
- OfficeChad 1y ago[dead]
- acedTrex 1y agoHi everyone, author here. Sorry about the JS stuff I wrote this while also fooling around with alpine.js for fun. I never expected it to make it to HN. I'll get a static version up and running. Happy to answer any questions or hear other thoughts. Edit: https://static.jaysthoughts.com/ https://static.jaysthoughts.com/ Static version here with slightly wonky formatting, sorry for the hassle. Edit2: Should work on mobile now well, added a quick breakpoint.
- konaraddi 1y agoGiven the topic of your post, and high pagespeed results, I think >99% of your intended audience can already read the original. No need to apologize or please HN users.
- pu_pe 1y ago> While the industry leaping abstractions that came before focused on removing complexity, they did so with the fundamental assertion that the abstraction they created was correct. That is not to say they were perfect, or they never caused bugs or failures. But those events were a failure of the given implementation a departure from what the abstraction was SUPPOSED to do, every mistake, once patched led to a safer more robust system. LLMs by their very fundamental design are a probabilistic prediction engine, they merely approximate correctness for varying amounts of time. I think what the author misses here is that imperfect, probabilistic agents can build reliable, deterministic systems. No one would trust a garbage collection tool based on how reliable the author was, but rather if it proves it can do what it intends to do after extensive testing. I can certainly see an erosion of trust in the future, with the result being that test-driven development gains even more momentum. Don't trust, and verify.
- acedTrex 1y ago> I think what the author misses here is that imperfect, probabilistic agents can build reliable, deterministic systems. No one would trust a garbage collection tool based on how reliable the author was, but rather if it proves it can do what it intends to do after extensive testing. > but rather if it proves it can do what it intends to do after extensive testing. Author here: Here I was less talking about the effectiveness of the output of a given tool and more so about the tool itself. To take your garbage collection example, sure perhaps an agentic system at some point can spin some stuff up and beat it into submission with test harnesses, bug fixes etc. But, imagine you used the model AS the garbage collector/tool, in that say every sweep you simply dumped the memory of the program into the model and told it to release the unneeded blocks. You would NEVER be able to trust that the model itself correctly identifies the correct memory blocks and no amount of "patching" or "fine tuning" would ever get you there. With other historical abstractions like say jvm, if the deterministic output, in this case the assembly the jit emits is incorrect that bug is patched and the abstraction will never have that same fault again. not so with LLMs. To me that distinction is very important when trying to point out previous developer tooling that changed the entire nature of the industry. It's not to say I do not think LLMs will have a profound impact on the way things work in the future. But I do think we are in completely uncharted territory with limited historical precedence to guide us.
- dirkc 1y agoI have a friend that always says "innovation happens at the speed of trust". Ever since GPT3, that quote comes to mind over and over. Verification has a high cost and trust is the main way to lower that cost. I don't see how one can build trust in LLMs. While they are extremely articulate in both code and natural language, they will also happily go down fractal rabbit holes and show behavior I would consider malicious in a person.
- acedTrex 1y agoAuthor here: I quite like that quote. A very succinct way of saying what took me a few paragraphs. This new world of having to verify every single thing at all points is quite exhausting and frankly pretty slow.
- Herring 1y agoSo get another LLM to do it. Judging is considerably easier [For LLMs] than writing something from scratch, so LLM judges will always have that edge in accuracy. Equivalently, I also like getting them to write tons of tests to build trust in correct behavior.
- acedTrex 1y ago> Judging is considerably easier than writing something from scratch I don't agree with this at all. Writing new code is trivially easy, to do a full in depth review takes significantly more brain power. You have to fully ascertain and insert yourself into someone elses thought process. Thats way more work than utilizing your own thought process.
- Herring 1y agoSorry, I should have been more specific. I meant LLMs are more reliable and accurate at judging than at generating from scratch. They basically achieve over 80% agreement with human evaluators [1]. This level of agreement is similar to the consensus rate between two human evaluators, making LLM-as-a-judge a scalable and reliable proxy for human judgment. [1] https://arxiv.org/abs/2306.05685 https://arxiv.org/abs/2306.05685 (2023)
- geor9e 1y agoThey changed the headline to "Yes, I will judge you for using AI..." so I feel like I got the whole story already.
- dr-detroit 1y ago[dead]
- satisfice 1y agoLLMs make bad work— of any kind— look like plausibly good work. That’s why it is rational to automatically discount the products of anyone who has used AI. I once had a member of my extended family who turned out to be a con artist. After she was caught, I cut off contact, saying I didn’t know her. She said “I am the same person you’ve known for ten years.” And I replied “I suppose so. And now I realized I have never known who that is, and that I never can know.” We all assume the people in our lives are not actively trying to hurt us. When that trust breaks, it breaks hard. No one who uses AI can claim “this is my work.” I don’t know that it is your work. No one who uses AI can claim that it is good work, unless they thoroughly understand it, which they probably don’t. A great many students of mine have claimed to have read and understand articles I have written, yet I discovered they didn’t. What if I were AI and they received my work and put their name on it as author? They’d be unable to explain, defend, or follow up on anything. This kind of problem is not new to AI. But it has become ten times worse.
- bobjordan 1y agoI see where you're coming from, and I appreciate your perspective. The "con artist" analogy is plausible, for the fear of inauthenticity this technology creates. However, I’d like to offer a different view from someone who has been deep in the trenches of full-stack software development. I’m someone who put in my "+10,000 hours" programming complex applications, before useful LLMs were released. I spent years diving into documentation and other people's source code every night, completely focused on full-stack mastery. Eventually, that commitment led to severe burnout. My health was bad, my marriage was suffering. I released my application and then I immediately had to walk away from it for three years just to recover. I was convinced I’d never pick it up again. It was hearing many reports that LLMs had gotten good at code that cautiously brought me back to my computer. That’s where my experience diverges so strongly from your concerns. You say, “No one who uses AI can claim ‘this is my work.’” I have to disagree. When I use an LLM, I am the architect and the final inspector. I direct the vision, design the system, and use a diff tool to review every single line of code it produces. Just recently, I used it as a partner to build a complex optimization model for my business's quote engine. Using a true optimization model was always the "right" way to do it but would have taken me months of grueling work before, learning all details of the library, reading other people’s code, etc. We got it done in a week. Do I feel like it’s my work? Absolutely. I just had a tireless and brilliant, if sometimes flawed, assistant. You also claim the user won't "thoroughly understand it." I’ve found the opposite. To use an LLM effectively for anything non-trivial, you need a deeper understanding of the fundamentals to guide it and to catch its frequent, subtle mistakes. Without my years of experience, I would be unable to steer it for complex multi-module development, debug its output, or know that the "plausibly good work" it produced was actually wrong in some ways (like N+1 problems). I can sympathize with your experience as a teacher. The problem of students using these tools to fake comprehension is real and difficult. In academia, the process of learning, getting some real fraction of the +10,000hrs is the goal. But in the professional world, the result is the goal, and this is a new, powerful tool to achieve better results. I’m not sure how a teacher should instruct students in this new reality, but demonizing LLM use is probably not the best approach. For me, it didn't make bad work look good. It made great work possible again, all while allowing me to have my life back. It brought the joy back to my software development craft without killing me or my family to do it. My life is a lot more balanced now and for that, I’m thankful.
- HardCodedBias 1y agoAll of this fighting against LLMs is pissing in the wind. It seems that LLMs, as they work today, make developers more productive. It is possible that they benefit less experienced developers even more than experienced developers. More productivity, and perhaps very large multiples of productivity, will not be abandoned due roadblocks constructed by those who oppose the technology due to some reason. Examples of the new productivity tool causing enormous harm (eg: bug that brings down some large service for a considerable amount of time) will not stop the technology if it being considerable productivity. Working with the technology and mitigating it's weaknesses is the only rational path forward. And those mitigation can't be a set of rules that completely strip the new technology of it's productivity gains. The mitigations have to work with the technology to increase its adoption or they will be worked around.
- ge96 1y agoIt is funny (ego) I remember when React was new and I refused to learn it, had I learned it earlier I probably would have entered the market years earlier. Even now I have this refusal to use GPT where as my coworkers lately have been saying "ChatGPT says" or this code was created by chatGPT idk, for me I take pride writing code myself/not using GPT but I also still use google/stackoverflow which you could say is a slower version of GPT.
- anthonypasq 1y agothis mindset does not work in software. My dad would still be programming with punchcards if he thought this way. instead he using copilot daily writing microservices and isnt some annoying dinosaur
- ge96 1y agoyeah it's pro con, I also hear my coworkers saying "I don't know how it works" or there are methods in the code that don't exist But anyway I'm at the point in my career where I am not learning to code/can already do it. Sure languages are new/can help there for syntax edit: other thing I'll add, I can see the throughput thing, it's like a person has never used opensearch before and it's a rabbithole, anything new there's that wall you have to overcome, but it's like we'll get the feature done, but did we really understand how it works... do we need to? Idk. I know this person can barely code but because they use something like chatGPT they're able to crap out walls of code and with tweaking it will work eventually -- I am aware this sounds like gatekeeping from my part Ultimately personally I don't want to do software professionaly/trying to save/invest enough then get out just because the job part sucks the fun out of development. I've been in it for about 10 years now, should have been plenty of time to save but I'm dumb/too generous. I think there is healthy skepticism too vs. just jumping on the bandwagon that everyone else is doing and really my problem is just I'm insecure/indecisive, I don't need everyone to accept me especially if I don't need money Last rant, I will be experimenting with agentic stuff as I do like Jarvis, make my own voice rec model/locally runs.
- okayoroof 1y ago[dead]
- observationist 1y agoThere's no reason to think AI will stop improving, and the rate of improvement is increasing as well, and no reason to think that these tools won't vastly outperform us in the very near future. Putting aside AGI and ASI, simply improving the frameworks of instructions and context, breaking down problems into smaller problems, and methodology of tools will result in quality multiplication. Making these sort of blanket assessments of AI, as if it were a singular, static phenomena is bad thinking. You can say things like "AI Code bad!" about a particular model, or a particular model used in a particular context, and make sense. You cannot make generalized statements about LLMs as if they are uniform in their flaws and failure modes. They're as bad now as they're ever going to be again, and they're getting better faster, at a rate outpacing the expectations and predictions of all the experts. The best experts in the world, working on these systems, have a nearly universal sentiment of "holy shit" when working on and building better AI - we should probably pay attention to what they're seeing and saying. There's a huge swathe of performance gains to be made in fixing awful human code. There's a ton of low hanging fruit to be gotten by doing repetitive and tedious stuff humans won't or can't do. Those two things mean at least 20 or more years of impressive utility from AI code can be had. Things are just going to get faster, and weirder, and weirder faster.
- christhecaribou 1y agoSure, if we all collectively ignore model collapse.
- ayakaneko 1y agoI think that, yes sure, there's no reason to think AI will stop improving. But I think that everyone is lossing trust not because there is no potential that LLMs could write good code or not, it's the trust to the user who uses LLMs to uncontrollable-ly generate those patches without any knowledge, fact checks, and verifications. (many of them may not even know how to test it.) In another word, while LLMs is potentially capable of being a good SWE, but the human behind it right now, is spamming, and doing non-sense works, and let the unpaid open source maintainers to review and feedback them (most of the time, manually).
- klabb3 1y ago
- I_Lorem 1y agoHe's making a good point on trust, but, really, doesn't the trust flow both directions? Should the Sr. Engineer rubber stamp or just take a quick glance at Bob's implementation because he's earned his chops, or should the Sr. Engineer apply the same level of review regardless of whether it's Bob, Mary, or Rando Calrissian submitting their work for review?
- eikenberry 1y agoThe Sr. Engineer should definitely give (presumably another Sr. Eng.) Bob's code a quicky review and approve it. If Mary or Rando are Sr. then they should get the same level as well. If anyone is a Jr. they should get a much more in-depth review as it's a teaching opportunity, whereas Sr. on Sr. reviews are done to enforce conventions and to be sure the PR has an audience (people take more care when they know other people will look at it).
- blurbleblurble 1y agoI bumped into this at work but not in the way you might expect. My colleague and I were under some pressure to show progress and decided to rush merging a pretty significant refactor I'd been working on. It was a draft PR but we merged it for momentum's sake. The next week some bugs popped up in an untested area of the code. As we were debugging, my colleague revealed his assumption that I'd used AI to write it, and expressed frustration at trying to understand something AI generated after the fact. But I hadn't used AI for this. Sure, yes I do use AI to write code. But this code I'd written by hand and with careful deliberate thought to the overall design. The bugs didn't stem from some fundamental flaw in the refactor, they were little oversights in adjusting existing code to a modified API. This actually ended up being a trust building experience over all because my colleague and I got to talk about the tension explicitly. It ended up being a pretty gentle encounter with the power of what's happening right now. In hindsight I'm glad it worked out this way, I could imagine in a different work environment, something like this could have been more messy. Be careful out there.
- kldg 1y agoIt can be a pretty a serious and offensive accusation for sure. When a dev voices their own characters in a game and has a flat affect and/or stilted speech pattern, it's inevitable to be called AI by someone. Art I don't understand or appreciate? Likely AI. Unimpressed by a Eurovision entry? Call it AI. Some people toss this around casually, but I wouldn't. I made myself known to be a big fool ~4 years ago. A local newspaper published an article on a particular person with outrageous claims primarily using photographs as proof. I challenged the editor directly via email, laying out my reasoning for why I was sure the images were manipulated. My arguments relied on misunderstandings on my part and the person claims were levied against showing zero deviation in position and stance while posing with multiple people during a meet-and-greet. The editor was offended and trolled me in response. I didn't let up, and he realized I was an idiot, not an agitator, and shared the full unpublished video from where the photos were taken with me, at which point I apologized deeply and made a donation. My ego was appropriately small for the following year. Before emailing him, I shared the photos with some level-headed friends for their opinion, specifically because I didn't want to make a false accusation. They came to the same conclusion that the images were most likely manipulated, so I was very confident going in. Now I trust this paper and people involved implicitly, but this was a lot of work to convince just one person.
- throwawayoldie 1y agoIMHO, s/may/has/
- benreesman 1y agoI'm currently standing up a C++ capability in an org that hasn't historically had one, so things like the style guide and examples folder require a lot of care to give a good start for new contributors. I have instructions for agents that are different in some details of convention, e.g. human contributors use AAA allocation style, agents are instructed to use type first. I convert code that "graduates" from agent product to review-ready as I review agent output, which keeps me honest that I don't myself submit code without scrutiny to the review of other humans: they are able to prompt an LLM without my involvement, and I'm able to ship LLM slop without making a demand on their time. Its an honor system, but a useful one if everyone acts in good faith. I get use from the agents, but I almost always make changes and reconcile contradictions.
- heisenbit 1y ago> The reality is that LLMs enable an inexperienced engineer to punch far above their proverbial weight class. That is to say, it allows them to work with concepts immediately that might have taken days, months or even years otherwise to get to that level of output. At the moment LLMs allow me to punch far above my weight class in Python where I do a short term job. But then I know all the concepts from decades dabbling in other ecosystems. Let‘s all admit there is a huge amount of accidental complexity (h/t Brooks‘s Silver-bullet) in our world. For better or worse there are skill silos that are now breaking down.
- mensetmanusman 1y agoAll this means is that the QC is going to be 10x more important.
- wg0 1y agoWe have seen those 10x engineers churning out PRs and huge PRs before anyone can fathom and make sense of the whole damn thing. Wondering what they would be producing with LLMs?
- lawlessone 1y agoOne trust breaking issue is we still can't know why the LLM makes specific choices. Sure we can ask it why it did something but any reason it gives is just something generated to sound plausible.
- mizzao 1y agoThe last section of this post seems to be quite predictive of a sibling post on the front page right now: https://news.ycombinator.com/item?id=44382752 https://news.ycombinator.com/item?id=44382752
- fhd2 1y agoA bit tangential, but I noticed quite a discrepancy between augmented coding done well, and augmented coding how I actually see it done in the wild. There's a lot of posts about how to do it well, and I like the idea of it, generally. I think GenAI has genuine applications in software development beyond as a Google/SO replacement. But then there's real world code. I constantly see: 1. Over engineering. People used to keep it simple because they were limited by how fast they can type. Well, those gloves sure did come off for a lot of developers. 2. Lack of understanding / memory. If I ask someone about how their code works, if they didn't write it (or at least carefully analyse it), it's rare for them to understand or even remember what they did there. The common answer to "how does this work?", went from "I think like this but let me double check" to "no idea". Some will be proud to tell you they auto generated documentation, too. If you have any questions about that, chances are you'll get another "no idea" response. If you ask an LLM how it works, that's very hit and miss for non-trivial systems. I always tell my devs I hire them to understand systems first and formost, building systems comes second. I feel increasingly alone with that attitude. 3. Bugs. So many bugs. It seems devs that generate code would need to do a lot more explicit testing than those who don't. There's probably just a missing feedback loop: When typing in code, you tend to have to test every little button action and so on at least once, it's just part of the work. Chances are you don't break it since you last tested it, so while this happens, manually written code generally has one time exhaustive manual testing built into the process naturally. If you generate a whole UI area, you need to do thorough testing of all kinds of conditions. Seems people don't. So while it could be great, from my perspective, it feels like more of a net negative in practice. It's all fun and games until there's a problem. And there always is. Maybe I have a bad sample of the industry. We essentially specialise on taking over technically disastrous projects and other kinds of tricky situations. Few people hire us to work on a good system with a strong team behind it. But still, comparing the questionable code bases I got into two years ago with those I get into now, there is a pretty clear change for the worse. Maybe I'm pessimistic, but I'm starting to think we'll need another software crisis (and perhaps a wee AI winter) to get our act together with this new technology. I hope I'm wrong.
- thedudeabides5 1y agodont.trust.machines
- helge9210 1y agoI checked with HR at my company and got an answer I'm not allowed to announce the following: anyone submitting the code or asking a question about the code without disclosing the fact that the code in question was generated by LLM would be cursed.
- archibaldJ 1y agoThis can be solved when the ARC puzzle is cracked (https://arcprize.org/play https://arcprize.org/play) so we can automate correctness-checking like in coq but for program synthesis.
- kordlessagain 1y agoSpending half the time building Claude Code tools (MCP servers) and half my time working on Gnosis, an AI powered oracle. What is an oracle? That's a system that: - Knows things - Has pre-crawled, indexed information about specific domains - Answers authoritatively - Not just web search, but curated, verified data - Connects isolated systems - Apps can query Gnosis instead of implementing their own crawling/search - May have some practical use for blockchain actions (typically a crypto "oracle" bridges web data with chain data. In this context the "oracle" is AI + storage + transactions on the chain. The Core Components: - Evolve: Our tooling layer - manages the MCP servers, handles deployment, monitors health. Agentic tools. - Wraith: Web crawler that fetches and processes content from URLs, handles JavaScript rendering, screenshots, and more. Agentic crawler. - Alaya: Vector database (streaming projected dimensions) for storing and searching through all the collected information. Agentic storage. - Gnosis-Docker: Container orchestration MCP server for managing these services locally. Agentic DevOps. There's more coming. https://github.com/kordless/gnosis-evolve https://github.com/kordless/gnosis-evolve https://linkedin.com/in/kordless https://linkedin.com/in/kordless https://github.com/kordless/gnosis-wraith https://github.com/kordless/gnosis-wraith (under heavy development) There's also a complete MCP inspection and debugging system for Python here: https://github.com/kordless/gnosis-mystic https://github.com/kordless/gnosis-mystic