9 ms·
LLMs are the ultimate demoware
- bgwalter 1y agoAnd the article is kicked to the third page because it is well written and the demoware metaphor is so powerful that it needs to be suppressed.
- eitland 1y agoAgain and again people keep saying this while many of us keep using LLMs to create value.
- jdiff 1y agoCountless people in comments say this, but other people fail to see evidence of that in the wild. As has been said in response to this point many times in the past: Where's the open source renaissance that should be happening right now? Where are the actual, in-use dependencies and libraries that are being developed by AI? The only times I've personally seen LLMs engaged in repos has been handling issues, and they made an astounding mess of things that hurt far more often than it helped for anything more than automatically tagging issues. And I don't see any LLMs allowed off the leash to be making commits. Not in anything with any actual downstream users.
- simonw 1y agoLet's look at every PR on GitHub in public repos (many of which are likely to be under open source licenses) that may have been created with LLM tools, using GitHub Search for various clues: GitHub Copilot: 247,000 https://github.com/search?q=is%3Apr+author%3Acopilot-swe-agent%5Bbot%5D&type=pullrequests&s=created&o=desc https://github.com/search?q=is%3Apr+author%3Acopilot-swe-age... - is:pr author:copilot-swe-agent[bot] Claude: 147,000 https://github.com/search?q=is%3Apr+in%3Abody+%28%22Generated+with+Claude+Code%22+OR+%22Co-Authored-By%3A+Claude%22+OR+%22Co-authored-by%3A+Claude%22%29&type=pullrequests&s=created&o=desc https://github.com/search?q=is%3Apr+in%3Abody+%28%22Generate... - is:pr in:body ("Generated with Claude Code" OR "Co-Authored-By: Claude" OR "Co-authored-by: Claude") OpenAI Codex: ~2,000,000 (over-estimate, there's no obvious author reference here so this is just title or bid containing "codex"): https://github.com/search?q=is%3Apr+%28in%3Abody+OR+in%3Atitle%29+codex&type=pullrequests https://github.com/search?q=is%3Apr+%28in%3Abody+OR+in%3Atit... - is:pr (in:body OR in:title) codex Suggestions for improvements to this methodology are welcome!
- pjc50 1y agoWhat's the acceptance rate on such PRs?
- simonw 1y agoAdd is:merged to see. For Copilot I got 151,000 out of 247,000 = 61% For Claude 124,000 / 147,000 = 84% For Codex 1.7m / 2m = 85%
- simonw 1y ago... I just found out there's an existing repo and site that's been running these kinds of searches for a while: https://prarena.ai/ https://prarena.ai/ and https://github.com/aavetis/PRarena https://github.com/aavetis/PRarena
- pizlonator 1y agoThe main problem with your search methodology is that maybe AI is good at generating a high volume of slop commits. Slop commits are not unique to AI. Every project I’ve worked on had that person who has high commit count and when you peek at the commits they are just noise. I’m not saying you’re wrong btw. Just saying this is a possible hole in the methodology
- miltonlost 1y agoThat's a denominator of total. How many are actually useful?
- lm28469 1y agoHN people: lines of code and numbers of PRs are irrelevant to determine the capabilities of a developer. Also HN people: look at the magic slop machine, it made all these lines of codes and PRs, it is irrefutable proof that it's good and AGI
- simonw 1y agoBoth of these things can be true at the same time: 1. Counting lines of code is a bad way to measure developer productivity. 2. The number of merged PRs on GitHub overall that were created with LLM assistance is an interesting metric for evaluating how widely these tools are being used.
- pizlonator 1y agoWhat would it mean to see it in the wild? I think that highly productive people who have incorporated LLMs into their workflows are enjoying a productivity multiplier. I don’t think it’s 2x but it’s greater than 1x, if I had to guess. It’s just one of those things that’s impossible to measure beyond reasonable doubt
- unshavedyak 1y ago> Countless people in comments say this, but other people fail to see evidence of that in the wild. As has been said in response to this point many times in the past: Where's the open source renaissance that should be happening right now? Where are the actual, in-use dependencies and libraries that are being developed by AI? The thing that this comment misses, imo, is that LLMs are not always enabling people who previously couldn't create value to create value. In fact i think they are likely to cause some people who created value previously to create even less value! However that's not mutually exclusive with enabling others to create more value than they did previously. Is it a net gain for society? Currently I'd bet not, by a large margin. However is it a net gain for some individual users of LLMs? I suspect yes. LLMs are a powerful tool for the right job, and as time goes on the "right job" keeps expanding to more territory. The problem is it's a tool that takes a keen eye to analyze and train on. It's not easy to use for reliable output. It's currently a multiplier for those willing to use it on the right jobs and with the right training (reviews, suspicion, etc).
- eitland 1y ago> The thing that this comment misses, imo, is that LLMs are not always enabling people who previously couldn't create value to create value. In fact i think they are likely to cause some people who created value previously to create even less value! Agree. For some time I’ve compared AI to a nail gun: It can make an experienced builder much quicker at certain jobs. But for someone new to the trade, I’m not convinced it makes them faster at all. It might remove some of the drudgery, yes — but it also adds a very real chance of shooting oneself in the foot (or hand).
- badsectoracula 1y agoWell, i haven't used LLMs much for code (i tried it, it was neat but ultimately i found it more interesting to do things myself) and i refuse to rely on any cloud-based solutions, be it AI or not, so i've only been using local LLMs, but even so i've found a few neat uses for it. One of my favorite uses is that i have configured my window manager (Window Maker) that when i press Win+/ it launches xterm with a script that runs a custom C++ utility based on llama.cpp that combines a prompt that asks a quantized version of Mistral Small 3.2 to provide suggestions for grammar and spelling mistakes in text, then uses xclip to put whatever i have selected and filters the program's output through another utility that colorizes the output using some simple regex. Whenever i write any text that i care about having (more) correct grammar and spelling (e.g. documentation - i do not use it for informal text like this one or in chat) i use it to find mistakes as English is not my first language (and it tends to find a lot of them). Since the output is shown in a separate window (xterm) instead of replacing the text i can check if the correction is fine (and the act of actually typing the correction helps me remember some stuff... in theory at least :-P). The [0] shows an example of how it looks. I also wrote a simple Tcl/Tk script that calls some of the above with more generalized queries, one of which is to translate text to English, which i'm mainly using to translate comments on Steam games[1] :-P. It is also helpful whenever i want to try out something quickly, like -e.g.- recently i thought that common email obfuscation techniques in text (like some AT example DOT com) are pointless nowadays with LLMs, so i tried that from a site i found online[2] (pretty much everything that didn't rely on JavaScript was defeated by Mistral Small). As for programming, i used Devstral Small 1.0 once to make a simple raytracer, though i wrote about half of the code by hand since it was making a bunch of mistakes[3]. Also recently i needed to scrape some data from a page - normally i'd do it by hand, but i was feeling bored at the time so i asked Devstral to write a Python script using Beautiful Soup to do it for me and it worked just fine. None of the above are things i'd value for billions though. But at the same time, i wouldn't have any other solution for the grammar and translation stuff (free and under my control at least). [0] https://i.imgur.com/f4OrNI5.png https://i.imgur.com/f4OrNI5.png [1] https://i.imgur.com/jPYYKCd.png https://i.imgur.com/jPYYKCd.png [2] https://i.imgur.com/ytYkyQW.png https://i.imgur.com/ytYkyQW.png [3] https://i.imgur.com/FevOm0o.png https://i.imgur.com/FevOm0o.png
- eitland 1y agoUsing the same arguments people used (use?) against IDEs and I think also against compilers and stuff back in the punch card days. I am not a researcher, but I am a techlead and I've seen it work again and again: IDEs work. And LLMs work. They are force multipliers though, they absolutely work best with people who already know a bit of software engineering.
- whatever1 1y agoAre you Nvidia? If not, then I don’t believe you.
- lm28469 1y agoIf the value produced was so great I think we've be able to measure it by now, or at least see something. If you remove the hype around AI the economy is actually on the way down, productivity wasn't measured to have increased since llm became mainstream either. Lot's of vibes and feelings but 0 measurable impact
- rsynnott 1y agoThe trouble is that peoples' self-evaluation of things that they believe are helping them is generally poor, and there's, at best, weak and conflicting evidence which is _not_ based on polling users. In particular, "producing stuff" is not necessarily "creating value"; some stuff has _negative_ value.
- dpflan 1y agoYes, AI allows for exquisite demos, demos that tantalize the audience into thinking of the infinite potential of the technology, that stunning vision expands and expands until the universe of potential overwhelms the dreamer into a state of terminal fantasy. So it is always a solution looking for a problem. There are cases where the two meet more realistically and a valuable impactful company develops it. The fact it can generate human language that is very compelling for certain context, makes it seem possible of doing so for many, many more contexts.
- js8 1y agoLLM models have evolved to autonomously convince humans that they're useful. They're the ultimate memetic parasites.
- some-guy 1y agoI think LLMs can both be bad for humanity (which I believe) AND useful at certain tasks. The general populace has been convinced that they’re the ultimate authority which is very sad (e.g. “@grok is this true?”)
- rsynnott 1y ago... I mean, it's not evolution. These things have people guiding them. Note the whole 'agreeability' controversy. That one's is a bit like cigarette companies back in the day optimising their products for addictiveness; do you do the right thing, or the thing that makes people buy your product more?
- clueless 1y ago"then fails to consistently help in completing tasks when deployed for daily use." This article seems to be baitware trying to push some outdated perspective. LLMs have only gotten more powerful over the last 3 years (being able to do more things), and so far not much has stopped them from becoming even more powerful (with the help of reasoning, other external methods, etc) in the future. "daily use" is so subjective and this article will be out dated soon as we get closer to an AGI (with LLMs as the base layer and not the main driver)
- slaterbug 1y agoWhat evidence is there that AGI will come “soon”?
- dsr_ 1y agoOr "ever"? (I'm not denying the possiblity. I'm proclaiming a lack of evidence.)
- slaterbug 1y agoI’ve been daydreaming lately about what the fundamental limits of “intelligence” could be, something like the concept of computability but for AI, or even biological brains. Though I will say, surely the existence of the human brain (which by definition is general intelligence), suggests that creating AGI is fundamentally possible?
- dsr_ 1y agoSure, it's possible - as you say, we have an existence proof. We don't know how to do it any other way, though. None of the people who claim that they or somebody else is on the trail has produced any evidence that they are correct.
- lm28469 1y agoThey can "feel it", like people "felt" we'd have commercial space flight "soon" after we put people on the moon, it's all delusion and wishful thinking.
- js8 1y agoNNs are also demoware in the sense they contain extremely condensed and incomprehensible model of the world (or part of). Demo coders would be proud. Edit: I mean their outputs are procedurally generated, like in https://en.m.wikipedia.org/wiki/Demoscene https://en.m.wikipedia.org/wiki/Demoscene
- yvdriess 1y agoI wouldn't say they contain any model of the world. They're a statistical predictive model, which have proven effective at certain tasks. My take is that the demoware part is not inherent to the NN approach, but rather that the tasks it's unreasonably effective produce very cool demo-able tasks for which the audience readily fills in the blanks. Cool demos make it easier to get further resources, so demoware-prone techniques tend to pull more funding, at least for a while.
- tptacek 1y agoIt's wild to me that, of all the things to call LLMs out for, this piece has chosen to include math tutoring. I've been doing Math Academy for a bit over 6 months now, going from (essentially) Algebra II through Calc II (integration by parts, arc lengths, Taylor expansions) and LLMs have been a huge part of what has made that effective: * Clear explanation of concepts that respond to questions and reformulate when things bounce * Step-by-step verification of solutions, spotting exactly where calculations have gone * Instantaneously generating new problem sets to reinforce concepts LLMs are probably not going to live up to all sorts of claims their proponents make. But I don't think you can ever have tried to use an LLM in a math course and reach the conclusion that it's "demoware" for that application. At what point, over 6 months of continuous work, does it stop being a "demo"?
- empath75 1y agoIt seems very hard to maintain the belief that LLMs are useless in the face of the fact that millions of people are using them. It's very much "nobody goes there anymore, it's too crowded"
- tptacek 1y agoI think you'd be crazy to say LLMs are blockchain-style hype when it comes to software development but I don't begrudge anybody who believes they're not currently workable for the kinds of problems they work on; I think reasonable people can disagree about how ready for prime time they are for production software development. But for math tutoring? If you claim LLM math tutoring is demoware, you're very clearly telling on yourself.
- tjr 1y agoI have seen LLMs fabricate bogus calculations; I personally would be hesitant to use an LLM as my one and only source of math learning, but I suppose using it in conjunction with something like Math Academy mitigates that issue? You've clearly had good success here, but any problem areas with the LLM to watch out for?
- 1y ago
- turnsout 1y agoThis is such a weak take to read while I have Claude Code running in the background creating a new database migration for a feature we're building
- dingnuts 1y agoIf Claude Code can do it what do we need you for?
- simonw 1y agoSomeone needs to know what a "database migration" is in order to ask Claude to build one. I think a lot of people are massively underestimating how much knowledge and skill is needed in software engineering beyond typing code into a text editor.
- bgwalter 1y agoYou know that Linus was at one point "typing code into an editor" to review other people's patches because he found it easier to catch mistakes that way. If your field of software engineering is so simple that you can survive on code snippets stolen from other people, great. Please do not generalize that.
- evilduck 1y agoHow do you know it did it correctly?
- turnsout 1y agoGet this: I read the code and understand it
- turnsout 1y agoSomeone has to come up with the idea for app, have a vision for what it needs to be, and continually push it forward without going off-course.
- 1y ago
- tedggh 1y agoI wish demoware or even battle tested software was that easy to sell.
- cronelius 1y agoLLM improvement is a sigmoid, not a parabola. The sooner we understand this, the less money we will lose to deceptive marketing
- book_mike 1y agoLLMs are useful if you use them properly and they are getting better everyday. Arguing against LLMs is like arguing against a shovel. Just use it right.
- soperj 1y agoI haven't noticed them getting any better in the last year.
- simonw 1y agoYou absolutely have not been paying attention then. The difference in quality between September 2025 LLMs (GPT-5, Claude 4/4.5) and September 2024 (we were still on GPT-4o) is huge. For one thing, last year's LLMs were nowhere near winning gold on collegiate math and programming competitions. That's because the "reasoning" thing hadn't kicked off yet - the first model to demonstrate that trick was o1 in ... OK that was September 12th 2024 so it just makes it to a year old now.
- lm28469 1y agoThat's the theory, if the vast majority of people use it wrong the problem is the tool, not the user.
- gipp 1y agoA lot of arguing "against LLMs" is not arguing "shovels aren't useful," it's arguing "maybe shovels aren't actually going to replace all human labor, and sinking so much capital into it we're starting to conceptualize it in terms of 'percent of global GDP' might not be such a great idea."
- rdtsc 1y ago> “Demoware” is a type of software that looks “good” during a demonstration. I like the term. I have been using a similar phrase "looks good in a snippet" when referring to certain styles of programming. Once such instance was when nodejs was becoming popular and everyone was showing how easy concurrent programming can be with a few callbacks in a snippet. However building a large code base with that would eventually turn into a nightmare. Another example is databases which don't fsync after writes by default. They look great in benchmarks (webscale even!) then in production suddenly some of data goes missing. But at least those initial benchmark demos were impressive.
- bgwalter 1y agoThe initial ChatGPT release in 2022 was the product of 7 years of private research that in turn built on decades of public research. Rumors say that Google wasn't far behind at the time, but didn't push releases. Perhaps because they were not that impressed by the applications or did not want "AI" to cannibalize their other products. So it seems very likely that everything has been squeezed out of the decades of research and we have plateaued. Desperate measures like Nvidia buying its own graphics cards through circular investment schemes do not inspire confidence either. Or Microsoft now doing CoPilot product placement ads in teenager YouTube channels. When Google launched, people just used it because it was good. This all fits very well with the demoware angle of the article.
- rsynnott 1y agoFollowing a proud tradition; 4GLs and 5GLs and no-code solutions and so forth were also, essentially, demoware.