3 ms·
It's the uncanny valley of AI. It's still not quite good enough yet that you can trust it blindly on a big codebase, so you still have to read and understand ev
by konschubert 29d ago
It's the uncanny valley of AI. It's still not quite good enough yet that you can trust it blindly on a big codebase, so you still have to read and understand everything - which is often harder than just writing it up yourself.
EDIT: don’t get me wrong. I still think AI is incredibly useful for a lot of tasks! But when implementing an architecturally hairy thing, I find it less stressful and equally quick to jump down to the editor level and use AI just for code completion.
- noir_lord 29d agoPretty much, The one thing I use it for is as a sanity check, pretty much "Look at <SomeFile>, point out issues you see, summarise them tersely" and it'll spot stuff a code review by a human might have spotted (in the mythical land where people actually do code reviews properly and don't just flag a spelling mistake to "show they looked at it"). Beyond that I don't trust it at all and I still write all my code the meat sack way. Trust is earned not given and it hasn't earned it yet.
- huurtehoog 29d agoIf anything, I think this hype cycle is fast exposing just how many people, teams, and companies just don't care about what is correct. They just wanna feel good about themselves and get paid. I for one welcome the fact that this whole thing has driven me back to books and deeper into the fundamentals. I have never read so much on math, hardware, and history as in the past 3 years or so.
- deadbabe 29d agoAre you doing those things for your own enjoyment though, or to eventually capitalize on it? And if it’s just for enjoyment, then doesn’t it make sense other people who want the same would just get a job where they can keep pushing things to an AI, feel good, get paid, then quickly get back to the hobbies they really love?
- huurtehoog 29d agoYou are advocating for the kind of compartmentalization that got us here in the first place.
- deadbabe 28d agoNothing wrong with compartmentalization. Some people like it, some people don’t. Just different ways to live your life.
- saltcured 28d agoI don't think "compartmentalization" does this topic justice. Note, I'm not accusing you, since I don't know your work. My post is about the general milieu and the rhetoric around it. This "compartment" terminology seems to whitewash a dimension that has diligence on one end, and fraud on the other. I've always worked in high-trust organizations where we depended on each other operating on the diligent side of things, making conservative choices to never wander into the murky area in between. It is horrifying to me how many people seem complacent about or even complicit in a different objective, which seems hell bent on wandering as far into the murk as one can without being caught. When the person who has a duty of diligence starts rubber-stamping AI outputs, they're veering off into that same murk. They accidentally or wantonly trust the agent as if they have delegated their duty of care. But the AI tool has no such duty and no capacity to care. I think this worker who has turned themselves into an outsourcing middleman needs to treat the results just like "found code" in a USB drive they found in the parking lot. Its origins and purpose are unclear. It could be flawed or obscurely inappropriate for the intended application, it could have legal entanglements, or it could even be subtly adversarial. The review task to figure this out is not simple. It is not something you do by skimming the result, or worse, asking some other AI tool to review and summarize. The person importing such code to a project needs a different kind of diligence to try to screen it. For a lot of people, I think this review may be impossible or at least no less laborious than doing the original work themselves with the required diligence. And, I think this importer needs to be fully liable and responsible for the outcome. But, instead, I think we're seeing frequent appeals to blame the machine and act like it is an honest mistake to let things pass because they've been rubber-stamping the imports. A lazy desire to claim credit for appearance of success, but shirk responsibility for detected failures.
- simonuv 29d ago[dead]
- 2sk21 29d agoMuch the same here - I have gone back to basics and am studying a lot more than I ever did.
- throwaway219450 28d agoMy advice is to try letting the agent fill in the gaps. You can probably architect better than it can. Write your class outlines, explicitly define the public facing bits and what you want APIs to look like. Write the key integration tests that you know ought to pass. The real advantage is that agents routinely write code without any silly copy/paste mistakes like accidentally accessing x twice on a coordinate operation instead of x and y. You can add some comments for what the function should do, throw in some real/pseudo code and let the LLM figure it out.
- neuronic 29d agoYour assumption is that LLMs will ever leave this uncanny valley. Maybe unforeseen breakthroughs and different architectures are achieved. Given LLM fundamental shortcomings grounded in mathematics and information theory, I highly doubt they will and we will always need to deal with these issues in some capacity.
- konschubert 29d agoI think that bigger context windows help, but I feel that for AI to cross this chasm, it needs to be able to encode more abstract context knowledge. I think the holy grail here is online learning.
- backlava12 29d ago[dead]
- bigstrat2003 28d agoAgreed. For AI to be something you can trust to operate autonomously, it needs to actually be able to understand the things it is working with and reason about them. LLMs cannot, by their very nature, do that. There can be no reliability with such a tool.
- itsalwaysgood 29d agoI'd say it's more about learning how to organize your work more efficiently. If you think about a product like marble: it's something that most be chiseled out of time. Some people can chisel better products: the AI is just a better chisel. Sometime still has to guide the chisel and judge the art/product. In our cases, the market judges products.
- konschubert 29d agoI think my point is that its sometimes (!) easier to use the manual chisel rather than go for the automatic chisel and then fix its mistakes.
- itsalwaysgood 29d agoThey say writing engages more of the brain and helps us to remember what's written more than if we just read it, or copy and paste. When you say it's easier to go manual, it seems you're talking about learning retention. And you're right. But seniors have learned enough that they're able to iterate quickly with AI. They know how to organize their work, manage change, tasks. They know how to break a problem down into smaller pieces. They're aware of context windows, token cost, estimated task lengths, etc. And most importantly, and to your point about ease: they have less to learn so retention isn't an issue. I have no opinion about whether we're in a good or bad situation, just making arguments from the toilet really.
- podocarp 29d agoIt's not just learning is it, why do people buy hand ground coffee when there quite literally isn't any difference? Or audiophile snake oil? As long as humans are still the consumers, some part of consumption will be emotional. Could be to support local artisans, could be gullibility, could be love, whatever. Maybe one day artisanal code will be a thing lol. Hand written like calligraphy. Those with refined tastes will have their favorite code artisans. And the plebs can continue with mass produced industrial junk.
- cyanydeez 29d agoI think the speed/context size of the large models is a threshold. I've been using a local model and watching it do killer stuff, and also shit out useless things; all in real time, requiring active steering.
- foobarbecue 29d agoI agree, except for the use of the word "yet" . I think what's missing is fundamental. I think the reason it sucks so much to work with LLM-generated code is that LLMs will never "know" what it's like to be human. They don't "understand" our frustrations and motivations, and they're missing the vast array of useful mental tactics we've evolved to cope with corporal existence. At this point I think progress towards a good colleague bot would require a new architecture which allows continuous leaning, and for the LLM to be raised as a human child (maybe in a simulation at 1000x speed or something).
- VCFundedGenYer 29d agoNah, It's not even good for small code changes. Try using it with Ansible. It spits back complete buffoonery.
- acedTrex 29d agoI was investigating an ansible playbook yesterday that had a 45 line comment to explain a single apt install command, completely and utterly useless. I am updating my neovim to just collapse all comments, the noise is unbearable.
- sevenseacat 29d agoI've seen people report cleaner code by forbidding agents from writing comments. Anecdotal, but interesting.
- cpeterso 28d agoMozilla's Firefox AGENTS.md begins with: Limit the amount of comments you put in the code to a strict minimum. You should almost never add comments, except sometimes on non-trivial code, function definitions if the arguments aren't self-explanatory, and class definitions and their members. Do not remove existing comments unless they are directly related to what you are changing. https://searchfox.org/firefox-main/source/AGENTS.md https://searchfox.org/firefox-main/source/AGENTS.md
- astrange 28d agoSomething like this seems unlikely to work for a behavior so burned into them. It'd be better to do a second pass to delete all the nonsense comments.
- weknowbetter 28d agoI promise you this doesn't work. Might help but not a lot.
- lelanthran 28d ago
- autoexec 28d ago> It's the uncanny valley of AI. It's still not quite good enough yet that you can trust it blindly on a big codebase, so you still have to read and understand everything - which is often harder than just writing it up yourself. You might as well have left it with: "It's still not quite good enough yet that you can trust it". That's the core of the issue. It doesn't matter what you ask it to do, it can't be trusted. Some things are just easier to verify and correct than others.
- lelanthran 28d ago> That's the core of the issue. It doesn't matter what you ask it to do, it can't be trusted. I dunno about that - whenever I ask it if I'm any good, I remain confident that it will assure me that I am!