12 ms·
As a programmer of over 20 years - this is terrifying. I'm willing to accept that I just have "get off my lawn" syndrome or something. But the idea of letting
by 0xFACEFEED 2y ago
As a programmer of over 20 years - this is terrifying.
I'm willing to accept that I just have "get off my lawn" syndrome or something.
But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible. Whenever I sit down to write some code, be it a large implementation or a small function, I think about what other people (or future versions of myself) will struggle with when interacting with the code. Is it clear and concise? Is it too clever? Is it too easy to write a subtle bug when making changes? Have I made it totally clear that X is relying on Y dangerous behavior by adding a comment or intentionally making it visible in some other way?
It goes the other way too. If I know someone well (or their style) then it makes evaluating their code easier. The more time I spend in a codebase the better idea I have of what the writer was trying to do. I remember spending a lot of time reading the early Redis codebase and got a pretty good sense of how Salvatore thinks. Or altering my approaches to code reviews depending on which coworker was submitting it. These weren't things I were doing out of desire but because all non-trivial code has so much subtlety; it's just the nature of the beast.
So the thought of opening up a codebase that was cobbled together by an AI is just scary to me. Subtle bugs and errors would be equally distributed across the whole thing instead of where the writer was less competent (as is often the case). The whole thing just sounds like a gargantuan mess.
Change my mind.
- bongodongobob 2y agoYou. Can. Write. Tests.
- ok_dad 2y agoTests haven’t saved us so far, humans have been writing tests that passed for software with bugs for decades.
- lanternfish 2y agoTests aren't a full solution for all the considerations of the above post.
- blitzar 2y agoJust let the LLM do that too.
- the_real_cher 2y agoEven better you can let the AI write tests.
- hakunin 2y agoHow do you write a test for code clarity / readability / maintainability?
- pnut 2y agoIf the code is a generated artefact and conforms to the test cases, who cares? What maintenance is required when the codebase is regenerated from scratch on each build?
- 0xFACEFEED 2y agoHow do tests account for cases where I'm looking at a 100 line function that could have easily been written in 20 lines with just as much, if not more, clarity? It reminds me of a time (long ago) when the trend/fad was building applications visually. You would drag and drop UI elements and define logic using GUIs. Behind the scenes the IDE would generate code that linked everything together. One of the selling points was that underneath the hood it's just code so if someone didn't have access to the IDE (or whatever) then they could just open the source and make edits themselves. It obviously didn't work out. But not because of the scope/scale (something AI code generation solves) but because, it turns out, writing maintainable secure software takes a lot of careful thought. I'm not talking about asking an AI to vomit out a CRUD UI. For that I'm sure it's well suited and the risk is pretty low. But as soon as you introduce domain specific logic or non-trivial things connected to the real world - it requires thought. Often times you need to spend more time thinking about the problem than writing the code. I just don't see how "guidance" of an LLM gets anywhere near writing good software outside of trivial stuff.
- canadianfella 2y ago[dead]
- Aeolun 2y ago> How do tests account for cases where I'm looking at a 100 line function that could have easily been written in 20 lines with just as much, if not more, clarity? That’s not a failure of the AI writing that 100 line monstrosity, it’s a failure of you deciding to actually use the thing. If you know what 20 lines are necessary and the AI doesn’t output that, why would you use it?
- pama 2y ago> How do tests account for cases where I'm looking at a 100 line function that could have easily been written in 20 lines with just as much, if not more, clarity? If the function is fast to evaluate and you have thorough coverage by tests, you couod iterate on an LLMs that aims to compress it down to a simpler / shorter version that behaves identical to the original function. Of course brevity for the sake of brevity can lead to less code that is not always more clear or simpler to understand than the original —LLMs are very good at mimicing code style, so show them a lot of your own code and ask them to mimic it and you may be surprized.
- TeMPOraL 2y agoMore importantly, you can read diffs. Depending on whether I'm using LLMs from my Emacs or via a tool like Aider, I either review and manually merge offered modifications as diffs (in editor), or review the automatically generated commits (Aider). Either way, I end up reading a lot of diffs and massaging the LLM output on the fly, and nothing that I haven't reviewed gets pushed to upstream. I mean, people aren't seriously pushing unreviewed LLM-generated code to production? Current models aren't good enough for that.
- fzeroracer 2y agoThe most common failure of TDD is that assuming just bolting on more tests will fix the problem of a poorly designed codebase.
- sbarre 2y agoI think it depends on the stakes of what you're building. A lot of the concerns you describe make me think you work in a larger company or team and so both the organizational stakes (maintenance, future changes, tech debt, other people taking it over) and the functional stakes (bug free, performant, secure, etc) are high? If the person you're responding to is cranking out a personal SaaS project or something they won't ever want to maintain much, then they can do different math on risks. And probably also the language you're using, and the actual code itself. Porting a multi-thousand line web SaaS product in Typescript that's just CRUD operations and cranking out web views? Sure why not. Porting a multi-thousand line game codebase that's performance-critical and written in C++? Probably not. That said, I am super fascinated by the approach of "let the LLM write the code and coach it when it gets it wrong" and I feel like I want to try that.. But probably not on a work project, and maybe just on a personal project.
- 0xFACEFEED 2y ago> A lot of the concerns you describe make me think you work in a larger company or team and so both the organizational stakes (maintenance, future changes, tech debt, other people taking it over) and the functional stakes (bug free, performant, secure, etc) are high? The most financially rewarding project I worked on started out as an early stage startup with small ambitions. It ended up growing and succeeding far beyond expectations. It was a small codebase but the stakes were still very high. We were all pretty experienced going into it so we each had preferences for which footguns to avoid. For example we shied away from ORMs because they're the kind of dependency that could get you stuck in mud. Pick a "bad" ORM, spend months piling code on top of it, and then find out that you're spending more time fighting it than being productive. But now you don't have the time to untangle yourself from that dependency. Worst of all, at least in our experience, it's impossible to really predict how likely you are to get "stuck" this way with a large dependency. So the judgement call was to avoid major dependencies like this unless we absolutely had to. I attribute the success of our project to literally thousands of minor and major decisions like that one. To me almost all software is high stakes. Unless it's so trivial that nothing about it matters at all; but that's not what these AI tools are marketing toward, are they? Something might start out as a small useful library and grow into a dependency that hundreds of thousands of people use. So that's why it terrifies me. I'm terrified of one day joining a team or wanting to contribute to an OSS project - only to be faced with thousands of lines of nonsensical autogenerated LLM code. If nothing else it takes all the joy out of programming computers (although I think there's a more existential risk here). If it was a team I'd probably just quit on the spot but I have that luxury and probably would have caught it during due diligence. If it's an OSS project I'd nope out and not contribute.
- ianbutler 2y agoI have 10 years professional experience and I've been writing code for 20 years, really with this workflow I just read and review significantly more code and I coach it when it structures or styles something in a way I don't like. I'm fully in control and nothing gets committed I haven't read its an extension of me at that point. Edit: I think the issues you've mentioned typically apply to people too and the answer is largely the same. Talk, coach, put hard fixes in like linting and review approvals.
- Aeolun 2y ago> Talk, coach, put hard fixes in like linting and review approvals. And sometimes, when all that doesn’t work? Just do it yourself :)
- geysersam 2y agoI'll take a stab at changing your mind. AIs are not able to write Redis. That's not their job. AIs should not write complex high performance code that millions of users rely on. If the code does something valuable for a large number of people you can afford humans to write it. AIs should write low value code that just repeats what's been done before but with some variations. Generic parts of CRUD apps, some fraction of typical frontends, common CI setups. That's what they're good at because they've seen it a million times already. That category constitutes most code written. This relieves human developers of ballpark 20% of their workload and that's already worth a lot of money.
- 0xFACEFEED 2y agoI can definitely see the value in letting AI generate low stakes code. I'm a daily CoPilot user and, while I don't let it generate implementations, the suggestions it gives for boilerplate-y things is top notch. Love it as a tool. My major issue with your position is that, at least in my experience, good software is the sum of even the seemingly low risk parts. When I think of real world software that people rely on (the only type I care about in this context) then it's hard to point a finger at some part of it and go "eh, this part doesn't matter". It all matters. The alternative, I fear, is 90% of the software we use exhibiting subtle goofy behavior and just being overall unpleasant to use. I guess an analogy for my concern is what it would look like if 60% of every film was AI generated using the models we have today. Some might argue that 60% of all films are low stakes scenes with simple exposition or whatever. And then remaining 40% are the climax or other important moments. But many people believe that 100% of the film matters - even the opening credits. And even if none of that were an issue: in my experience it's very difficult to assess what part of an application will/won't be low/high stakes. Imagine being a tech startup that needs to pivot your focus toward the low stakes part of the application that the LLM wrote.
- Aeolun 2y agoI think your concept of ‘what the AI wrote’ is too large. There is zero chance my one line copilot or three line cursor tab completions are going to have an effect on the overall quality of my codebase. What it is useful for is doing exactly the things I already know need to happen, but don’t want to spend the effort to write out (at least, not having to do it is great). Since my brain and focus aren’t killed by writing crud, I get to spend that on more useful stuff. If it doesn’t make me more effective, at least it makes my job more enjoyable.
- hansonkd 2y ago> The more time I spend in a codebase the better idea I have of what the writer was trying to do. This whole thing of using LLMs to Code reminds me a bit of when Google Translate came out and became popular, right around the time I started studying Russian. Yes, copying and pasting a block of Russian text produced a block of english text that you could get a general idea of what was happening. But translating from english to russian rarely worked well enough to fool the professor because of all the idioms, style, etc. Russian has a lot of ways you can write "compactly" with fewer words than english and have a much more precise meaning of the sentence. (I always likened russian to type-safe haskell and english to dynamic python) If you actually understood Russian and read the text, you could uncover much deeper and subtle meaning and connections that get lost in translation. If you went to russia today you could get around with google translate and people would understand you. But you aren't going to be having anything other than surface level requests and responses. Coding with LLMs reminds me a lot of this. Yes, they produce something that the computer understands and runs, but the meaning and intention of what you wanted to communicate gets lost through this translation layer. Coding is even worse because i feel like the intention of coding should never to be to belt out as many lines as possible. Coding has powerful abstractions that you can use to minimize the lines you write and crystalize meaning and intent.
- balder1991 2y ago> the intention of coding should never to be to belt out as many lines as possible That’s such an underrated statement. Especially when you consider the amount of code as a liability that you’ll have to take care later.
- mreid 2y agoI've heard a similar sentiment: "It's not lines of code written, it's lines of code spent." It also reminds me of this analogy for data, especially sensitive data: "it's not oil, it's nuclear waste."
- abadpoli 2y agoThis presumes that it will be real humans that have to “take care” of the code later. A lot of the people that are hawking AI, especially in management, are chasing a future where there are no humans, because AI writes the code and maintains the code, no pesky expensive humans needed. And AI won’t object to things like bad code style or low quality code.
- mithametacs 2y agoYou still use type systems, tests, and code review. For a lot of use cases it's powerful. If you ask it to build out a brand new system with a complex algorithm or to perform a more complex refactoring, it'll be more work correcting it than doing it yourself. But that malformed JSON document with the weird missing quotation marks (so the usual formatters break), and spaces before commas, and the indentation is wild... Give it to an LLM. Or when you're writing content impls for a game based on a list of text descriptions, copy the text into a block comment. Then impl 1 example. Then just sit back and press tab and watch your profits.
- girvo 2y agoThe (mostly useless boilerplate “I’m basically just testing my mocks”) tests are being written by AI too these days. Which is mildly annoying as a lot of those tests are basically just noise rather than useful tools. Humans have the same problem, but current models are especially prone to it from what I’ve observed And not enough devs are babysitting the AI to make sure the test cases are useful, even if they’re doing so for the original code it produced
- chillfox 2y agoThere are very few tutorials on how to do testing and I don't think I have ever seen one that was great. Compared to general coding stuff where there's great tutorials available for all the most common things. So I think quality testing is just not in the training data at anywhere close to the quantity needed.
- fullstackchris 2y agoTesting well is both an art and a science, and I mean, just look at the dev community on the topic, some are religious about TDD, some say unit tests only, some say the whole range to e2e etc. etc. hard to have good training data when there is no definition of what is "right" in the first place!
- sdesol 2y ago> But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible. I do think it is kind of crazy based on what I've seen. I'm convinced LLM is a game changer but I couldn't believe how stupid it can be. Take the following example, which is a spelling and grammar checker that I wrote: https://app.gitsense.com/?doc=f7419bfb27c8968bae&samples=5 https://app.gitsense.com/?doc=f7419bfb27c8968bae&samples=5 If you click on the sentence, you can see that Claude-3.5 and GPT-4o cannot tell that GitHub is spelled correctly most of the time. It was this example that made me realize how dangerous LLM can be. The sentence is short but Claude-3.5 and GPT-4o just can't process it properly. Having a LLM rewrite large swaths of code is crazy but I believe with proper tooling to verify and challenge changes, we can mitigate the risk. I'm just speculating, but I believe GitHub has come to the same conclusion that I have, which is, all models can be stupid, but it is unlikely that all will be stupid at the same time.
- outside1234 2y agoThe saying "You can delegate tasks but not responsibility" comes to mind. You are still responsible for the code AI is writing. It is just that writing code with AI is more like reviewing a PR now.
- OJFord 2y ago> As a programmer of over 20 years - this is terrifying. > > I'm willing to accept that I just have "get off my lawn" syndrome or something. > > But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible. My first thought was that I disagree (though I don't use or like this in-IDE AI stuff) because version control. But then the way people use (or can't use) SVC 'terrifies' me anyway, so maybe I agree? It would be fine correctly handled, but it won't be, sort of thing.
- com2kid 2y ago> But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible. People felt the same about compilers for a long time. And justifiably so, the idea that compilers are reliable is quite a new one, finding compilers bugs used to be pretty common. (Those experimenting with newer languages still get to enjoy the fun of this!) How about other code generation tools? Presumably you don't take much umbrage with schema generators? Or code generators that take a scheme and output library code (OpenAPI, Protocol buffers, or even COM)? Those can easily take a few dozen lines of input and output many thousands of LoC, and because they are part of an automated pipeline, even if you do want to fix the code up, any fixes you make will be destroyed on the next pipeline run! But there is also a LOT of boring boilerplate code that can be automated. For example, the necessary code to create a new server, attach a JSON schema to a POST endpoint, validate a bearer token, and enable a given CORS config is pretty cut and dry. If I am ramping up on a new backend framework, I can either spend hours learning the above and then copy and paste it forever more into each new project I start up, or I can use an AI to crap the code out for me. (Actually once I was setting up a new server and I decided to not just copy and paste and to do it myself, I flipped the order of two `use` directives and it cost me at least 4 hours to figure out WTF was wrong....) > As a programmer of over 20 years I'm almost up there, and my view is that I have two modes of working: 1. Super low level, where my intimate knowledge of algorithms, the language and framework I'm using, of CPU and memory constraints, all come together to let me write code that is damn near magical. 2. Super high level, where I am architecting a solution using design patterns and the individual pieces of code are functionally very simple, and it is how they are connected together that really matters. For #1, eh, for some popular problems AI can help (popular optimizations on Stack Overflow). For #2, AI is the most useful, because I have already broken the problem down into individual bite size testable nuggets. I can have the AI write a lot of the boilerplate, and then integrate the code within the larger, human architected, system. > So the thought of opening up a codebase that was cobbled together by an AI is just scary to me. The AI didn't cobble together the system. The AI did stuff like "go through this array and check the ID field of each object and if more than 3 of them are null log an error, increment the ExcessNullsEncountered metric counter, and return an HTTP 400 error to the caller" Edit: This just happened I am writing a small Canvas game renderer, and I am having an issue with text above a character's head renders off the canvas. So I had Cursor fix the function up to move text under a character if it would have been rendered above the canvas area. I was able to write the instructions out to Cursor faster than I could have found a pencil and paper to sketch out what I needed to do.
- eru 2y ago> But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible. Why? Presumably you let your coworkers move code around, too, and then you review it? (And vice versa.)
- epolanski 2y ago> Change my mind. Unit, integration, e2e, types and linters would catch most of the things you mention. Not every software is mission critical, often the most important thing is to go as fast and possible and iterate very quickly. Good enough is better than very good in many cases.
- almostdeadguy 2y ago> Unit, integration, e2e, types and linters would catch most of the things you mention. Who’s writing those?
- fullstackchris 2y agoLots of people. For certain types of software (ISO) they are required. But I'm in the boat (and also experienced many times first hand) all those tests you write will by definition, never test against that first production bug you get :)
- almostdeadguy 2y agoMy point was not to question that people would write tests, the point I'm making is that it's tempting to generate both code and tests once you start using an LLM to generate large swaths of code, and then the assurance that tests give you goes out the window. I'm not convinced that using AI as more than auto-complete is really a viable solution, because you can't shortcut an understanding of the problem domain to be assured of the correctness of code (and at that point the AI has mostly saved you some typing). The theory-crafting process of building software is probably the most important aspect of it. It not only provides assurance of the correctness of what you're building, it provides feedback into product development (restrictions, pushback that suggests alternate ways of doing things, etc.).
- dools 2y ago> But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible I heard a similar thing from a dude when I said I use it for bash scripts instead of copying and pasting things off StackOverflow. He was a bit "get off my lawny" about the idea of running any code you didn't write, especially bash scripts in a terminal. It is obviously the case that I didn't write most of the code in the world by a very large margin, but even not taking it to extremes if I'm working on a team and people are writing code how is it any different? Everyone makes mistakes, I make mistakes. I think it's a bad idea to run things that you don't at least understand what it's going to do but the speed with which ChatGPT can produce, for example, gcloud shell commands to manage resources is lightning fast (all of which is very readable, just takes a while if you want to look it up and compose the commands yourself). If your quality control method is "making sure there are no mistakes" then it's already broken regardless of where the code comes from. Me reviewing AI code is no different from me reviewing anyone else's code. Me testing AI code using unit or integration tests is no different from testing anyone else's code, or my own code for that matter.
- scubbo 2y ago> Me reviewing AI code is no different from me reviewing anyone else's code. I take your point, and on the whole I agree with your post, but this point is fundamentally _not_ correct, in that if I have a question about someone else's code I can ask them about their intention, state-of-mind, and understanding at the time they wrote it, and (subjectively, sure; but I think this is a reasonable claim) can _usually_ detect pretty well if they are bullshitting me when they respond. Asking AI for explanations tends to lead to extremely convincing and confident false justifications rather than an admission of error or doubt. However: > Me testing AI code using unit or integration tests is no different from testing anyone else's code, or my own code for that matter. This is totally fair
- mikeshi42 2y agoI'm assuming by bullshitting you mean differentiating between LLM hallucinations and a human with low confidence in their code. I've found that LLMs do sometimes acknowledge hallucinations. But really the check is much easier than a PR/questioning an author - just run the code given by the copilot and check that it works, just as if you typed it yourself.
- afro88 2y ago> But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible. I think this is where the bimodality comes from. When someone says "I used AI to refactor 3000 loc" some take it to mean they used AI in small steps as an accellerator, and others take it to mean a direct copy/paste, fix compile errors and move on. Treat AI like a mid level engineer that you are pair programming with, who can type insanely fast. Move in small steps. Read through it's code after each small iteration. Ask it to fix things (or fix them yourself if quick and easy). Brainstorm ideas with it etc etc.
- nwienert 2y agoIt’s really far from mid level. It’s a weird mix of expert at things it trained on, and complete misleading idiot at anything outside. For a bash script or the first steps of something simple it’s great. For anything complex at all it’s worse than nothing.
- afro88 2y agoFor anything complex, move in small steps. For anything truly novel, or on a codebase with a very bespoke in house architecture or DSL, yeah you won't get much out of it.
- nwienert 2y agoEven in small steps, it fails. I have two cases I test with, nothing special, just some TS generics in one instance and a schema-to-schema mapping tool in another. Both things that Junior devs could do given a couple days, even though they'd need to study and figure out various pieces. o1 can't get either, no matter how much I break it down, no matter how much prodding. In fact the more you try the worse it gets. And yes I do try starting new conversations and splitting it out. Simply does not help, at all. It's not to say it isn't really helpful for really simple things. Or even complex things but that are directly in the training set. But the second you go outside that, it's terrible.
- anonzzzies 2y ago
- AdieuToLogic 2y ago> Whenever I sit down to write some code, be it a large implementation or a small function, I think about what other people (or future versions of myself) will struggle with when interacting with the code. Is it clear and concise? Is it too clever? Is it too easy to write a subtle bug when making changes? Have I made it totally clear that X is relying on Y dangerous behavior by adding a comment or intentionally making it visible in some other way? > It goes the other way too. If I know someone well (or their style) then it makes evaluating their code easier. The more time I spend in a codebase the better idea I have of what the writer was trying to do. What I believe you are describing is a general definition of "understanding", which I am sure you are aware. And given your 20+ year experience, your summary of: > So the thought of opening up a codebase that was cobbled together by an AI is just scary to me. Subtle bugs and errors would be equally distributed across the whole thing instead of where the writer was less competent (as is often the case). Is not only entirely understandable (pardon the pun), but to be expected as algorithms employed lack the crucial bit which you identify - understanding. > The whole thing just sounds like a gargantuan mess. As it does to most whom envision having to live with artifacts produced by a statistical predictive text algorithm. > Change my mind. One cannot because understanding, as people know it, is intrinsic to each person by definition. It exists as a concept within the person whom possesses it and is defined entirely by said person.
- DeathArrow 2y ago> Whenever I sit down to write some code, be it a large implementation or a small function, I think about what other people (or future versions of myself) will struggle with when interacting with the code. Is it clear and concise? Is it too clever? Is it too easy to write a subtle bug when making changes? Have I made it totally clear that X is relying on Y dangerous behavior by adding a comment or intentionally making it visible in some other way? Over 20 years of experience, too, but I quit doing that for work. Nobody really really cares, all they care is about time to market and having features they've sold yesterday to customers being done today. As long as I follow some mental models and some rules, the code is reasonably well written and there is no need to procrastinate and think too much. When I write code for myself, or I am contributing to a small project with a small number of contributors, then things change. If I can afford and I like it I am not only willing to assure things are carefully thought out, but also I am willing to experiment and test until I am sure that I use the best variant I can come up with. Like going from 99% to 99.9%, even if it wouldn't matter in practice. Just for fun. As a manager, I wouldn't ask people to write perfect code, nor I would like them to ship buggy code very fast, but ship reasonably good code as fast as they can write reasonable good code.
- munksbeer 2y ago> Over 20 years of experience, too, but I quit doing that for work. Nobody really really cares, all they care is about time to market and having features they've sold yesterday to customers being done today. I don't recognise this. Or at least, I recognise that it can be that way but not always. In places I've worked, I tend to have worked with teams that care deeply about this. But we're not writing CRUD apps or web systems, or inventory management, or whatever. We're writing trading systems. I absolutely want to be working with code that we can understand in a hurry (and I mean, a real hurry) when things go wrong, and that we can change and/or fix in a hurry. So some of us really do care.
- DeathArrow 2y ago> We're writing trading systems If you write critical systems executing trades, managing traffic lights, landing a rover on the moon, then you should take your time and write the best possible version. Our code is both easy to read and easy to modify because that allows us to add features fast. It is not the very best possible version of what we can do, because that would cost us much more time. The code has few bugs, which are mostly caught by the QA teams, is reasonably fast. Maybe not the most elegant, not engineered to take into account future use cases and we push to eliminate from AC some very rare use cases,that will take too much time to implement. Maybe the code it's not the most resource efficient. But the key aspect is we focus on delivering the most features possible from what customers need in the limited amount of time we have and with the limited manpower we have. Company is owned by some private equity group and their focus is solely growing the customer base while paying as little as possible. Last year they fired 25% of personnel because their ARR was missing a few millions. Newertheless, most companies I worked before were in the hurry. With the exception of a very small company where I could work however I see fit.
- DeathArrow 2y ago>Change my mind. Nobody pays for splendid code that isn't in production. They will gladly pay for buggy code that is in production and solves their needs as long as the marketing team does a good job.
- KronisLV 2y ago> I think about what other people (or future versions of myself) will struggle with when interacting with the code. This feels like the sign of a good developer! On the other hand, sometimes you just need executable line noise that gets the job done by Thursday so you can ship it and think about refactoring later. As far as AI code goes, more often than not, it will read as something very generic, which is not necessarily a bad thing. When opening yet another Java CRUD project, I’d be more happy to see someone copy and pasting working code from tutorials or resources online (while it still works correctly), as opposed to seeing people develop bespoke systems on top of what a framework provides for every project.
- Etherlord87 2y ago> On the other hand, sometimes you just need executable line noise that gets the job done by Thursday so you can ship it and think about refactoring later. This is a problem too. ChatGPT enables you to write bad code. It's like Adobe Flash: flash using websites didn't have to be slow, but it was easy to make a slow website with it.
- JoeyJoJoJr 2y agoIt’s also like Macromedia Flash in that it is a highly creative force, but people who don’t get it or can’t make it work for them will complain.
- otikik 2y agoAgree that it is a mess. If I know that someone is using an llm to produce code, I think it is only fair that I use an LLM to review the code too. If you want me to put the work as a reviewer, you'd better put the work as a writer.
- Cthulhu_ 2y agoI don't know if you jest but this is likely the next stage, to be released within the next months, that is, AIs doing the first rounds of code reviews. It'll likely be from github / microsoft as they have one of the biggest code review datasets around.
- svieira 2y agoThis is already happening - I recently saw a resume which included "Added AI-driven code reviews" as an accomplishment bullet point (the person was working for a large consulting firm).
- lljk_kennedy 2y agoIt's like worrying about moving bits on a hard drive, or writing nice machine code. Eventually you just won't care. AI / LLMs interacting with code bases in future won't care about structure, clearness, conciseness etc. They'll tokenize it all the same.
- onion2k 2y agoThe whole thing just sounds like a gargantuan mess. Most apps are a gargantuan mess. It's just a mess that mostly works. In a typical large scale web app written in something like Node or PHP, I wouldn't be at all surprised if 95% of the code is brought in from libraries that the dev team don't review. They have no idea about the quality of the code they're running. I don't see why adding AI to the mix makes much of a difference.
- frankdenbow 2y ago"cobbled together by an AI" It will be as cobbled together as the thoughtfulness of the person in charge of the code. Same as if they wrote it themselves.
- HarHarVeryFunny 2y agoMaybe in the future when we have AGI, but not at the moment. Did you read yesterday's "How I code using Cursor" thread: https://news.ycombinator.com/item?id=41979203 https://news.ycombinator.com/item?id=41979203 The "Changes to my workflow" part is most relevant, and would be more accurately titled "How Cursor writes code differently to me [a senior developer]". For example: 1) Cursor/AI more likely to reinvent the wheel and write code from scratch rather than use support libraries. Good to avoid dependencies I suppose, but widely used specialized libraries are likely to be debugged, and mature - able to handle corner cases gracefully, etc. AI "writes code" by regenerating stuff from it's training data - akin to cut and pasting from Stack Overflow, etc. If you're using this for a throwaway prototype or personal project then maybe you don't care as long as it works most of the time, but for corporate production use this is a liability. 2) AI more likely to generate repetitive code rather than write reusable functions (which he spins as avoiding abstractions) means code that is harder to read, debug and maintain. It's like avoiding global symbolic constants and defining them multiple times throughout your code instead. This wouldn't pass typical human code review. When future you, or a co-worker, maybe using a different editor/IDE, fixes a bug, they may not realize that the same bug has been repeated multiple times throughout the code, rather than fixing it once in a function. We don't have human level AGI yet, far from it, and the code that today's AI generates reflects that. This isn't code that an experienced developer would write - this is LLM generated code, which means it's either regurgitated as-as from some unknown internet source, or worse yet (and probably more typical?) is a mashup of multiple sources, where the LLM may well have introduced it's own bugs in addition to those present in the original sources.
- TheNewsIsHere 2y agoI’m with you on this, at least in spirit. I’ve tried AI coding assistance tools multiple times. Notably with Ansible and then with some AWS stuff using Amazon Q. I decided I wanted to be curious and see how it went. With Ansible, it was an unusable mess. Ansible is a bit of an odd bird of a DSL though. You really need to be on top of versions and modules. It’s not well suited for AI because it requires a lot of nuanced understanding. There’s no one resource that will work (effectively) forever, like C or something. With AWS’s Amazon Q, I started out simple by having it write an IAM policy that I was having a hard time wrapping my head around. It wasn’t very helpful because the policies it provided used conditional keys that weren’t supported in the service that was addressed by the policy. I’ve found I can typically work with higher quality and less fixing by just writing it myself. I could babysit and teach an AI, but at that point what’s the point? I’m also unconvinced it’s worth the environmental impact, especially if I need to tutor it or mark up its output anyway. In any event, it’s easy enough to outsmart/out-clever oneself it a colleague (and vice versa). Adding AI to that just seems like adding a chaotic junior developer to that equation.
- Cthulhu_ 2y ago> Change my mind. Imagine the LLM is another developer and you're responsible for reviewing their code. Would you think of them the same thing? While I don't like AI either, I feel a lot of the fear around it is just that - fear and distrust that there will be bugs, some more subtle than others. But that's true for code written by anyone, isn't it? It's just that you're responsible for the code written by something or someone else, and that can be scary. But remember that no code should make it to production without human and automated review. Trust the system.
- shinycode 2y agoI have a colleague that did it also, moving parts of code and « writing » code quickly with copilot. Because it’s easier to overlook LLM updates he riddled the code with bugs. Subtle things that we undercover later, now that he’s gone. When you write everything yourself you are more keen to think deeply about changes. I read today that Google has 25% of their code written by AI. They have an history of trashing huge projects and the quality of their services is getting worse over time. Maybe the industry is going to move to « let’s trash the codebase and ask cGPT 8 to write everything with this new framework »… OP said he’s talking to AI like guiding an other dev. Isn’t he afraid that he will loose the ability to think about solutions for himself ? That’s a trained part of the brain that we can « loose » no ?
- peab 2y agoIt really depends what you're building. If you're building code that's going to go in some medical system, or a space shuttle, then yeah, you probably want to write every small function with great detail. If you're creating some silly consumer app like a "what will your baby look like in 5 years", then code quality doesn't matter you just need to ship fast. Most startups just need to ship fast to validate some ideas, 99% of your code will be deprecated within a few months
- nonethewiser 2y ago> But the idea of letting an LLM write/move large swaths of code seems so incredibly irresponsible. Whenever I sit down to write some code, be it a large implementation or a small function, I think about what other people (or future versions of myself) will struggle with when interacting with the code. Is it clear and concise? Is it too clever? I agree with our conclusion but not your supporting evidence. Not only can you read it to answer all these questions, but you can BETTER answer these questions from reading it. Because you are already looking at it from the perspective you are trying optimize for (future reader). What is less clear is if it handles all the edge cases correctly. In theory these should all be tested, but many of them cannot even be identified without thinking through the idiosyncrasies of the code which is a natural byproduct of writing code.
- thesz 2y agoYou are absolutely right. AI is used to generate new code, not to reduce its size through rewrite [1]. [1] https://www.folklore.org/Negative_2000_Lines_Of_Code.html https://www.folklore.org/Negative_2000_Lines_Of_Code.html The code is always a burden, a legacy. To support it, one need to look after it, take care of it, bear it's weight. I have 35+ years of experience on my shoulders. I always celebrate when I can reduce code - and, recently, I do that more and more often, being a support person for a large code base with a 20+ years of history. PS You were answering a comment that brags about migrating multithousand LOC code bases from language to the other. Recently I had to review a file that is more than (metric) megabyte (10^6 bytes) and more than 32 thousands LOC in size. I need to find a way to apply a fix for a problem triggered by single statement in several thousands (more that 1M of LOCs) of test cases.