11 ms·
AI can't stop making up software dependencies and sabotaging everything
- curiousgal 1y agoIf you get pwned by some AI code hallucination you deserve it honestly. They're code assistants not code developers.
- kazinator 1y agoIf you get pwned by external dependencies in any way, you deserve it. This idea of programs fetching reams of needed stuff from the cloud somewhere is a real scourge in programming.
- dijksterhuis 1y agoas ever, any task that has any sort of safety or security critical risks should never be left to a “magic black box”. human input/review/verification/validation is always required. verify the untrusted output of these systems. don’t believe the hype and don’t blindly trust them. — i did find the fact that google search’s assistant just parroted the crafted/fake READMEs thing particularly concerning - propagating false confidence/misplaced trust - although it’s not at all surprising given the current state of things. genuinely feel like “classic search” and “new-fangled LLM queries” need to be split out and separated for low-level/power user vs high-level/casual questions. at least with classic search i’m usually finding a github repo fairly quickly that i can start reading through, as an example. at the same time, i could totally see myself scanning through a README and going “yep, sounds like what i need” and making the same mistake (i need other people checking my work too).
- Arch485 1y ago> any task that has any sort of safety or security critical risks should never be left to a “magic black box”. > human input/review/verification/validation is always required. but, are humans not also a magic black box? We don't know what's going on in other people's heads, and while you can communicate with a human and tell them to do something, they are prone to misunderstanding, not listening, or lying. (which is quite similar to how LLMs behave!)
- dijksterhuis 1y agofrom my comment > at the same time, i could totally see myself scanning through a README and going “yep, sounds like what i need” and making the same mistake (i need other people checking my work too). yes, us humans have similar issues to the magic black box. i’m not arguing humans are perfect. this is why we have human code review, tests, staging environments etc. in the release cycle. especially so in safety/security critical contexts. plus warnings from things like register articles/CVEs to keep track of. like i said. don’t blindly trust the untrusted output (code) of these things — always verify it. like making sure your dependencies aren’t actually crypto miners. we should be doing that normally. but some people still seem to believe the hype about these “magic black box oracles”. the whole “agentic”/mcp/vibe-coding pattern sounds completely fucking nightmare-ish to me as it reeks of “blindly trust everything LLM throws at you despite what we’ve learned in the last 20 years of software development”.
- brookst 1y agoSounds like we just need to treat LLMs and humans similarly: accept they are fallible, put review processes in place when it matters if they fail, increase stringency of review as stakes increase. Vibe coding is all about deciding it doesn’t matter if the implementation is perfect. And that’s true for some things!
- dijksterhuis 1y ago> Vibe coding is all about deciding it doesn’t matter if the implementation is perfect. And that’s true for some things! i was going to say, sure yeah i’m currently building a portfolio/personal website for myself in react/ts, purely for interview showing off etc. probably a good candidate for “vibe coding”, right? here’s the problem - which is explicitly discussed in the article - vibe coding this thing can bring in a bunch of horrible dependencies that do nefarious things. so i’d be sitting in an interview showing off a few bits and pieces and suddenly their CPU usage spikes at 100% util over all cores because my vibe-coded personal site has a crypto miner package installed and i never noticed. maybe it does some data exfiltration as well just for shits and giggles. or maybe it does <insert some really dark thing here>. “safety and security critical” applies in way more situations than people think it does within software engineering. so many mundane/boring/vibe-it-out-the-way things we do as software engineers have implicit security considerations to bear in mind (do i install package A or package B?). which is why i find the entire concept of “vibe-coding” to be nightmarish - it treats everything as a secondary consideration to convenience and laziness, including basic and boring security practices like “don’t just randomly install shit”.
- ritabratamaiti 1y agoSlop in slop out
- croemer 1y agoThe article contains nothing new. Just opinions including a security firm CEO selling his security offerings. Read this instead, it's the technical report that is only linked to and barely mentioned in the article: https://socket.dev/blog/slopsquatting-how-ai-hallucinations-are-fueling-a-new-class-of-supply-chain-attacks https://socket.dev/blog/slopsquatting-how-ai-hallucinations-...
- dijksterhuis 1y agosocket article seems to mostly be a review of this arXiv preprint paper: https://arxiv.org/pdf/2406.10279 https://arxiv.org/pdf/2406.10279 there’s also some info from Python software foundation folks in the register article, so it’s not just a socket pitch article.
- kazinator 1y agoThat's not the technical report; it's also just a blog article which links to someone else's paper, and finishes off by promoting something: "Socket addresses this exact problem. Our platform scans every package in your dependency tree, flags high-risk behaviors like install scripts, obfuscated code, or hidden payloads, and alerts you before damage is done. Even if a hallucinated package gets published and spreads, Socket can stop it from making it into production environments."
- feross 1y agoHi — I’m the security firm CEO mentioned, though I wear a few other hats too: I’ve been maintaining open source projects for over a decade (some with 100s of millions of npm downloads), and I taught Stanford’s web security course (https://cs253.stanford.edu https://cs253.stanford.edu). Totally understand the skepticism. It’s easy to assume commercial motives are always front and center. But in this case, the company actually came after the problem. I’ve been deep in this space for a long time, and eventually it felt like the best way to make progress was to build something focused on it full-time.
- TheSwordsman 1y agoI'm waiting for the AI apologists to swarm on this post explaining how these are just the results of poorly written prompts, because AI could not make mistakes with proper prompts. Been seeing an increase of this recently on AI-critical content, and it's exhausting. Sure, with well written prompts you can have some success using AI assistants for things, but also with well-written non-ambiguous prompts you can inexplicably end up with absolute garbage. Until things become consistent, this sort of generative AI is more akin to a party trick than being able to replace or even supplement junior engineers.
- akdev1l 1y agoOne time some of our internal LLM tooling decided to delete a bunch of configuration and replace it with: “[EXISTING CONFIGURATION HERE]” Lmfaooo
- TheSwordsman 1y agoHahahaha. That's actually amazing.
- orbital-decay 1y agoHaha. That sounds like something Sonnet 3.6 would do, it learned to cheat that way and it's an absolute pain in the ass to make it produce longer outputs.
- bslalwn 1y agoYou are getting replaced, man. Burying your head in the sand won’t help.
- simonw 1y agoAs an "AI apologist", sorry to disappoint but the answer here isn't better prompting: it's code review. If an LLM spits out code that uses a dependency you aren't familiar with, it's your job to review that dependency before you install it. My lowest effort version of this is to check that it's got a credible commit and release history and evidence that many other people are using it already. Same as if some stranger opens a PR against your project introducing a new-to-you dependency. If you don't have the discipline to do good code review, you shouldn't be using AI-assisted programming outside of safe sandbox environments. (Understanding "safe sandbox environment" is a separate big challenge!)
- crazygringo 1y agoThis is a real problem, and AI is a new vector for it, but the root cause is the lack of reliable trust and security around packages in general. I really wonder what the solution is. Has there been any work on limiting the permissions of modules? E.g. by default a third-party module can't access disk or network or various system calls or shell functions or use tools like Python's "inspect" to access data outside what is passed to them? Unless you explicitly pass permissions in your import statement or something?
- matsemann 1y agoJava used to have Java Security Manager, which basically made it possible to set permissions for what a jar/dependency could do. But deprecated and no real good alternative anymore.
- fpoling 1y agoJava could have really nice security if it provided access to OS API via interfaces with main function receiving the interface for the real implementation. It would be possible then to implement really tight sandboxes. But that ship sailed 30 years ago…
- jruohonen 1y ago> This is a real problem, and AI is a new vector for it, but the root cause is the lack of reliable trust and security around packages in general. I agree. And the problem has intensified due to the explosion of dependencies. > Has there been any work on limiting the permissions of modules? With respect to PyPI, npm, and the like, and as far as I know: no. But regarding C and generally things you can control relatively easily yourself, see for instance: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=pledge%20openbsd&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- exitb 1y agoIt would be useful to have different levels of restrictions for various modules within a single process, which I don’t think pledge can do.
- 1970-01-01 1y agoA few days ago: https://news.ycombinator.com/item?id=43644880 https://news.ycombinator.com/item?id=43644880
- kazinator 1y agoThank you, AI, for exposing the idiocy of package-driven programming, where everything is a mess of churning external dependencies.
- brookst 1y agoWhat’s the alternative? Statically linked binaries?
- Isamu 1y agoPeople have made the point many times before, that “hallucination” is mainly what generative AI does and not the exception, but most of the time it’s a useful hallucination. I describe it as more of a “mashup”, like an interpolation of statistically related output that was in the training data. The thinking was in the minds of the people that created the tons of content used for training, and from the view of information theory there is enough redundancy in the content to recover much of the intent statistically. But some intent is harder to extract just from example content. So when generating statically similar output, the statistical model can miss the hidden rules that were a part of the thinking that went into the content that was used for training.
- alan-crowe 1y agoThere is an interesting phenomenon with polynomial interpolation called Runge Spikes. I think "Runge Spikes" offers a better metaphor than "hallucination" and argue the point: https://news.ycombinator.com/item?id=43612517 https://news.ycombinator.com/item?id=43612517
- Hackbraten 1y ago> the statistical model can miss the hidden rules that were a part of the thinking that went into the content that was used for training. Makes sense. Hidden rules such as, "recommending a package works only if I know the package actually exists and I’m at least somewhat familiar with it." Now that I think about it, this is pretty similar to cargo-culting.
- totetsu 1y agoCan’t we just move to have package managers point to a curated list of packages by default, with the option to enable an uncurated one if you know what your doing , ala Ubuntu source lists?
- miohtama 1y agoYes. But that would mean someone needs to work harder.
- ozim 1y agoThen you are stuck on whatever passes the gates. It is shitloads of work to maintain. Getting new package from 0 to any Linux distribution is close to impossible. Debian sucks as no one gets on top of reviewing and testing. „Can we just” is not just there is loads of work to be done to curate packages no one is willing to pay for it. There is so far no model that works where you can have up to date cutting edge stuff reviewed. So you are stuck with 5 year old crap because it was reviewed.
- tasuki 1y agoSo many good packages made it into Debian relatively recently! Eg: fzf, fd-find, ripgrep, jq, exa, nvim, ...
- mvdtnz 1y agoI was excited to see you mention fzf in Debian so I just installed it and it's an ancient version that doesn't support shell integration. Not a great look.
- pabs3 1y agofzf 0.60.3-1 from March 3rd is the latest in Debian, 0.61.1 from April 4 the latest upstream, so Debian is not exactly ancient. I guess you are using stable Debian? That doesn't get newer upstream releases, if you want the latest Debian fzf you can just install the build from unstable, which should work fine since Go does static linking, and there are often backports for when unstable packages can't be installed on stable. https://github.com/junegunn/fzf/tags https://github.com/junegunn/fzf/tags https://tracker.debian.org/pkg/fzf https://tracker.debian.org/pkg/fzf
- JohnCClarke 1y agoIME when AI "hallucinates" API endpoints or library functions that just aren't there it's almost always the case that they should be. In other words the AI has based it's understanding on the combined knoweledge of hundreds(?) of other APIs and libraries and is geenrating an obvious analogy. Turning this around: a great use case is to ask AI to review documents, APIs, etc. AI is really great for teasing out your blindspots.
- esafak 1y agoThe next step could be to ask it to generate the missing function.
- croes 1y agoIf the training data contains useless endpoints the AI will also hallucinate those useless endpoints. The wisdom of the crowd only works for the end result not if you consider every given answer, then you get more wrong answers because you fall to the average.
- jmaker 1y agoHard no. I’ve had that in lots of cases, it just applied some symmetry pattern to come up with purposefully absent public endpoints that exist for employees only. It gets dangerous the moment you put as much trust in it as you’re suggesting.
- keybored 1y agoI have no other gear than polemic on the topic of AI-for-code-generation so ignore this comment if you don’t like that. I think people in software envy real-engineering too much. Software development is what it is. If it does not live up to that bar then so be it. But AI-for-code-generation (“AI” for short now) really drops any kind of pretense. I got into software because it was supposed to be analytic, even kind of a priori. And deterministic. What even is AI right now? It melds the very high tech and probabilistic (AI tech) with the low tech of code generation (which is deterministic by itself but not with AI). That’s a regression both in terms of craftmanship (code generation) and so-called engineering (deterministic). I was looking forward to higher-level software development: more declarative (better programming languages and other things), more tool-assisted (tests, verification), more deterministic and controlled (Nix?), and less process redundancies (e.g. less redundancies in manual/automated testing, verification, review, auditing). Instead we are mining the hard work of the past three decades and spitting out things that have the mandatory label “this might be anything, verify it yourself”. We aren’t making higher-level tools—we[1] are making a taller tower with less support beams, until the tower reaches so high that the wind can topple it at any moment. The above was just for AI-for-code-generation. AI could perhaps be used to create genuinely higher level processes. A solid structure with better support. But that’s not the current trajectory/hype. [1] ChatGPT em-dash alert. https://news.ycombinator.com/item?id=43498204 https://news.ycombinator.com/item?id=43498204
- Tainnor 1y agoThis resonates with me. There are so many tools and techniques that many developers refuse to adopt because they're not hyped and take maybe a week to learn (e.g. take something like TLA+ which could be used to reason about distributed systems), but instead of improving the craft of programming we're just using LLMs to spam bad quality software at a faster rate.
- dep_b 1y agoIt's true, hallucinations in LLM's can be so consistent that I warn the LLM up front about stuff like "do not use NSCacheDefault, it does not exist, and there is no default value" and then keeping my fingers crossed it doesn't find a roundabout way to introduce it anyway. Can't really remember what is was exactly anymore, something in Apple's Vision libraries that just kept popping up if I didn't explicitly say to not use it.
- nurettin 1y agoWhen I suspect that it will make stuff up, I tell it to cite the docs that contain the functions it used. It causes more global warming, but it works fine.
- bdw5204 1y agoThis is just another reason why dependencies are an anti-pattern. If you do nothing, your software shouldn't change. I suspect that this style of development became popular in the first place because the LGPL has different copyright implications based on whether code is statically or dynamically linked. Corporations don't want to be forced to GPL their code so a system that outsources libraries to random web sites solves a legal problem for them. But it creates many worse problems because it involves linking your code to code that you didn't write and don't control. This upstream code can be changed in a breaking way or even turned into malware at any time but using these dependencies means you are trusting that such things won't happen. Modern dependency based software will never "just work" decades from now like all of that COBOL code from the 1960s that infamously still runs government and bank computer systems on the backend. Which is probably a major reason why they won't just rewrite the COBOL code. You could say as a counterargument that operating systems often include breaking changes as well. Which is true but you don't update your operating system on a regular basis. And the most popular operating system (Windows) is probably the most popular because Microsoft historically has prioritized backward compatibility even to the extreme point of including special code in Windows 95 to make sure it didn't break popular games like SimCity that relied on OS bugs from Windows 3.1 and MS-DOS[0]. [0]: https://www.joelonsoftware.com/2000/05/24/strategy-letter-ii-chicken-and-egg-problems/ https://www.joelonsoftware.com/2000/05/24/strategy-letter-ii...
- maleldil 1y agoWhat are you advocating for? Zero external dependencies? Write a new YAML parser from scratch? Rolling your own crypto?
- xrd 1y agoThe only real solution I see is lint and ci tooling that prevents non approved packages from getting into your repo. Even with this there is potential for theft on localhost. There are a dozen new YC startups visible in just those two sentences.
- sausagefeet 1y agoWho do you think is going to be writing those linting rules after the first person that cared about it the most finishes?
- xrd 1y agoGood point. And, surely a naughty LLM will get hacked and say "Hey, you've heard about this great thing called linting. Let me configure the newest and bestest one for you, it's called Rm-RF-Star..."
- CamperBob2 1y agoUsually, when the model hallucinates a dependency, the subject of the hallucination really should exist. I've often thought that was kind of interesting in itself. It can feel like a genuine glimpse of emergent creativity.
- mdp2021 1y agoChildren may invent the world as they do not know it well yet. Adults know that reality is not what you may expect. We need to deal with reality, so...
- onionisafruit 1y agoI’m not measuring it, but it seems like copilot suggests fewer imports than it used to. It could be that it has more context to see that I rarely import external packages and follows suit. Or maybe I’m using it subtlety different than I used to.
- Lvl999Noob 1y agoCould the AI providers themselves monitor any code snippets and look for non-existent dependencies? They could then ask the LLM to create that package with the necessary interface and implant an exploit in the code. Languages that allow build scripts would be perfect as then the malicious repo only needs to have the interface (so that the IDE doesn't complain) and the build script can download a separate malicious payload to run.
- ezst 1y agoThe AI providers already write the code, on the whole crazy promise that humans need not to care/read about it. I'm not sure that it changes anything at that point to add one weak level of indirection. You are already compromised.
- WhitneyLand 1y agoSeems to also especially love making up options and settings for command line tools.
- diggan 1y agoIf the LLM is "making up" APIs that don't exists, I'm guessing they've been introduced as the model tried to generalize from the training set, as that's the basic idea? These invented APIs might represent patterns the model identified across many similar libraries, or other texts people have written on the internet, wouldn't that actually be a sort of good library to have available if it wasn't already? Maybe we could use these "hallucinations" in a different way, if we could sort of know better what parts are "hallucination" vs not. Maybe just starting points for ideas if nothing else.
- brookst 1y agoBack in GPT3 days I put together a toy app that let you ask for a python program, and it hooked __getattr__ so if the LLM generated code called a non-existent function it could use GPT3 to define it dynamically. Ended up with some pretty wild alternate reality python implementations. Nothing useful though.
- OtherShrezzing 1y agoIn my experience, what's being made up is an incorrect name for an API that already exists elsewhere. They're especially bad at recommending deprecated methods on APIs.
- skydhash 1y agoThe average of the internet is heavily skewed towards the mediocre side.
- deleted 1y ago[deleted]
- marcosdumay 1y ago> wouldn't that actually be a sort of good library to have available if it wasn't already I for one do not want my libraries APIs defined by the median person commenting about code of making questions on Stack Overflow. Also, every time I see people using LLMs output as a starting point for software architecture the results became completely useless.
- Eggpants 1y ago
- aspbee555 1y agoI am constantly correcting the AI code it gives me, and all I get for it is "oh your right! here is the corrected code" then it gives me more hallucinations correcting the latest hallucination results in it telling me the first hallucination
- Trasmatta 1y agoI have this same experience. Vibe coding is literally hell.
- dwringer 1y agoIME it is rarely productive to ask an LLM to fix code it has just given you as part of the same session context. It can work but I find that the second version often introduces at least as many errors as it fixes, or at least changes unrelated bits of code for no apparent reason. Therefore I tend to work on a one-shot prompt, and restart the session entirely each time, making tweaks to the prompt based on each output hoping to get a better result (I've found it helpful to point out the AI's past errors as "common mistakes to be avoided"). Doing the prompting in this way also vastly reduces the context size sent with individual requests (asking it to fix something it just made in conversation tends to resubmit a huge chunk of context and use up allowance quotas). Then, if there are bits the AI never quite got correct, I'll go in bit by bit and ask it to fix an individual function or two, with a new session and heavily pruned context.
- Jcampuzano2 1y agoI agree with this, you will almost always get better results by simply undoing and rewording you prompt vs trying to coerce it to fix something it already did. Most of the time when I do use it, I almost always use just a couple prompts before starting a completely new one because it just falls off a cliff in terms of reliability after the first couple messages. At that point you're better off fixing it yourself than trying to get it to do it a way you'll accept.
- aspbee555 1y ago
- gchamonlive 1y agoPeople who solely code and are not good software architects will try and fail to delegate coding to LLM. What we are doing in practice when delegating coding to LLMs is climbing up the abstraction level ladder. We can compensate bad software architecture because we understand deeply the code details and make indirect couplings in the code. When we don't understand deeply the code, we need to compensate it with good architecture. That means thinking about code in terms of interfaces, stores, procedures, behaviours, actors, permissions and competences (what the actors should do, how they should behave and the scope of action they should be limited to). Then these details should reflect directly in the prompts. See how hard it is to make this process agentic, because you need user input in the agent inner workings. And after running these prompts and with luck successfully extracting functioning components, you are the one that should be putting these components together to make working system.
- ra0x3 1y ago> What we are doing in practice when delegating coding to LLMs is climbing up the abstraction level ladder. 100%. I like to say that we went from building a Millennium Falcon out of individual LEGO pieces, to instead building an entire LEGO planet made of Falcon-like objects. We’re still building, the pieces are just larger :)
- selfhoster 1y ago"What we are doing in practice when delegating coding to LLMs is climbing up the abstraction level ladder." Except that ladder is built on hallucinated rungs. Coding can be delegated to humans. Coding cannot be delegated to AI, LLM or ML because they are not real nor are they reliable.
- vbezhenar 1y agoI still think that main issue of hallucination is bad AI wrapper tools. AI must have every available public API with documentation, preloaded in the context. And explicit instructions to avoid using any API not mentioned in the context. LLM is like a developer without internet or docs access, who needs to write code on the paper. Every developer would hallucinate in that environment. It's a miracle that LLM does so much with so limited environment.
- alganet 1y ago> "What a world we live in: AI hallucinated packages are validated and rubber-stamped by another AI that is too eager to be helpful." That's actually hilarious.
- perrygeo 1y agoMy favorite is when the LLM hallucinates some function or an entire library and you call it out for the mistake. A likely response is "Oh, I'm sorry, You're right. Here's how you would implement function_that_does_not_exist()" and proceeds to write the library it hallucinated in the first place. It's quirks like these that prove LLMs are a long long way from AGI.
- unoti 1y agoWhen using AI, you are still the one responsible for the code. If the AI writes code and you don't read every line, why did it make its way into a commit? If you don't understand every line it wrote, what are you doing? If you don't actually love every line it wrote, why didn't you make it rewrite it with some guidance or rewrite it yourself? The situation described in the article is similar to having junior developers we don't trust committing code and us releasing it to production and blaming the failure on them. If a junior on the team does something dumb and causes a big failure, I wonder where the senior engineers and managers were during that situation. We closely supervise and direct the work of those people until they've built the skills and ways of thinking needed to be ready for that kind of autonomy. There are reasons we have multiple developers of varying levels of seniority: trust. We build relationships with people, and that is why we extend them the trust. We don't extend trust to people until they have demonstrated they are worthy of that trust over a period of time. At the heart of relationships is that we talk to each other and listen to each other, grow and learn about each other, are coachable, get onto the same page with each other. Although there are ways to coach llm's and fine tune them, LLM's don't do nearly as good of a job at this kind of growth and trust building as humans do. LLM's are super useful and absolutely should be worked into the engineering workflow, but they don't deserve the kind of trust that some people erroneously give to them. You still have to care deeply about your software. If this story talked about inexperienced junior engineers messing up codebases, I'd be wondering where the senior engineers and leadership were in allowing that to mess things up. A huge part of engineering is all about building reliable systems out of unreliable components and always has been. To me this story points to process improvement gaps and ways of thinking people need to change more than it points to the weak points of AI.
- jmaker 1y agoThe pace differs though. A junior would need a week for a feature an LLM can produce in an hour. And you’re expected to validate that just as quickly. And LLMs are trained to appeal to the reader, unlike an average junior dev. Devs will only get lazy the more they rely on LLMs. It’s like you’re at the university and there’s no homework anymore, just lectures. You’re just passively ingesting data, not getting trained on real problems because you’ve got AI to do that for you. So you’re no longer challenged to grow anymore in your domain. What’s left are hard problems that the AI will mislead you on because it’s unfamiliar with them, and your opportunity to learn was lost to delegating to AI. In the end the pressure will grow at work, more features will be expected in shorter time frames. You’ll get even less time to learn and grow as a developer or engineer.
- VladVladikoff 1y agoWhy can’t pypy / npm / etc just scan all newly uploaded modules for typical malware patterns before the package gets approved for distribution?
- simonw 1y agoBecause doing so is computationally expensive and would be making false promises. False positives where it incorrectly flagged a safe package would result in the need for a human review step, which is even more expensive. False negatives where malware patterns didn't match anything previously would happen all the time, so if people learned to "trust" the scanning they would get caught out - at which point what value is the scanning adding? I don't know if there are legal liability issues here too, but that would be worth digging into. As it stands, there are already third parties that are running scans against packages uploaded to npm and PyPI and helping flag malware. Leaving this to third parties feels like a better option to me, personally.
- VladVladikoff 1y ago>Leaving this to third parties feels like a better option to me, personally. Seems too late to me. At this point the module/package was already added into the ecosystem, it could potentially be some time (months?) before it is flagged by third party and removed.
- 12_throw_away 1y ago> Why can’t [X] just [Y] first? The word "just" here always presumes magic that does not actually exist.
- jruohonen 1y ago> The word "just" here always presumes magic that does not actually exist. The magic here is, yes, AI. If you look at the mobile app stores, they've all become much better, although false positives occur, of course.
- 1y ago
- stego-tech 1y agoThis was my chief critique when my company forced us to use their AI tooling. I was trying to stitch together our CMDB, two different VMware products, and the corporate technology directory into a form of product tenancy for our customers. At one point I was trying to move and transform data from our CRM into mongoDB, and figured "eh, let's knock these mandatory agent queries out of the way by asking the chatbot to help." I wrote a few prompts to try and explain what I was trying to accomplish (context), and how I'd like it done (instruction). The bot hallucinated a non-existent mongoDB Powershell cmdlet, complete with documentation on how it works, and then spat out a "solution" to the problem I asked. Every time I reworked the prompt, cut it up into smaller chunks, narrowed the scope of the problem, whatever I tried, the chatbot kept flatly hallucinating non-existent cmdlets, Python packages, or CLI commands, sometimes even providing (non-working) "solutions" in languages I didn't explicitly ask for (such as bash scripting instead of Powershell). This was at a large technology company, no less, one that's "all-in" on AI. If you're staying in a very narrow line with a singular language throughout and not calling custom packages, cmdlets, or libraries, then I suspect these things look and feel quite magical. Once you start doing actual work, they're complete jokes in my experience.
- vjerancrnjak 1y agoMost of the code is badly written. Models are doing what most of their dataset is doing. I remember, fresh out of college, being shocked by the amount of bugs in open source.
- simonw 1y agoMore recent models are producing much higher quality code than models from 6/12/18 months ago. I believe a lot of this is because the AI labs have figured out how to feed them better examples in the training - filtering for higher quality open source code libraries, or loading up on code that passes automated tests. A lot of model training these days uses synthetic data. Generating good code synthetic data is a whole lot easier than any other category, as you can at least ensure the code you're generating is gramatically valid and executes without syntax errors.
- jccooper 1y agoThe dataset isn't making up fake dependencies.
- jaco6 1y ago[dead]
- jmaker 1y agoI can’t speak to how others apply the LLMs, but in my coding experience they’re mostly in the way. A couple days ago VSCode released an agentic workflow similar to Cursor’s. And I must say, the YouTube demo was persuasive and I thought I’d just have to move to VSCode off JetBrains IDEs. So I gave it a spin, and after the past couple days, it’s been the most terrible IDE experience so far. The LLMs are always in the way, I’ve got Claude 3.5, 3.7, o1, o3-mini, o4, Gemini 2.0-flash, 2.5-pro, with/without reasoning, own models. Embedded Copilot is bugged, editor/agentic Copilot is bugged - it breaks your code if you selectively reject suggestions, your file buffer gets mangled, need to revert everything completely even if something was useful. Sidebar chat can get just as confusing as before. Typescript, Python, Java, Kotlin, Go. Rust won’t even compile, and don’t get me started on C++ codebases. Never had it type-check with mypy and pylance. In many cases even with codebases, extra MCP servers and fetching remote docs, it’s just not up to the task of making code even type check. Sometimes it just times out or fails on network errors. Very fragile, unreliable, misleading. I don’t know what people vibe code, but for a variety of codebases I’ve had to work on, it’s just in the way, injecting nonsense or outright garbage, which I need to reject every time. It’s useful as a sed alternative, but a less reliable one, a regex is often faster than 3+ prompts and waiting. It breaks my flow of conscience, I lose creativity and need to check everything after it. Dunning-Kruger maybe. To me that workflow is but a tailored and integrated StackOverflow, with snippets adapted to your code. Not sure how productive it is to let snippet insertions interfere with your flow, but very helpful when you forget or stumble. The more people rely on it, the more surprise awaits around the corner the moment AI fails. Now devs rely more on the network link instead of their brain, just like when people used to vibe code using StackOverflow. Creativity is at the stake. It must be kept in tone to stay productive. There’s a lot of work that’s a waste of time. If the goal is to replace devs, such companies will lose money in the end. If the goal is to assist devs and make them more productive, the LLMs need to adapt to take over such tasks reliably, e.g., scaffolding, standard algorithms, “best practices,” simulating and questioning design/architecture, and the UX must improve.