27 ms·
Some thoughts on LLMs and software development
- austinallegro 1y ago[flagged]
- krainboltgreene 1y ago> All major technological advances have come with economic bubbles, from canals and railroads to the internet. Is this actually correct? I don't see any evidence for a "airflight bubble" or a "car bubble" or a "loom bubble" at the technologies' invention. Also the "canal bubble" wasn't about the technology, it was about the speculation on a series of big canals but we had been making canals for a long time. More importantly, even if it was correct, there are plenty of bubbles (if not significantly more) around things that didn't have value or tech that didn't matter.
- mmmm2 1y agoThe Intelligent Investor by Graham talks about investors putting so much money into airlines and air freight that it became impossible to make a return. I don't know if you would call that a bubble, maybe just over exuberance.
- sfink 1y agoApologies for arguing from first principles, but for anything that spurs a lot of activity, the only alternative to a bubble is this: people ramp up investment and activity and enthusiasm only as much as the underlying thing can handle, then gradually taper off the increase and gently level off at the equilibrium "carrying capacity" of the new technology. Does that sound like any human, ever, to you? (The only time there isn't a bubble is when the thing just isn't that interesting to people and so there's never a big wave of uptake in the first place.)
- krainboltgreene 1y agoIf you were right then we'd have a lot more than a dozen listed bubbles on wikipedia. That's an absurd framing for a cute quip.
- BlueTemplar 1y agoWhy do you assume that Wikipedia can be trusted to provide an exhaustive answer to this question ?
- Marazan 1y agoThe AI bubble also isn't about the technology.
- antithesizer 1y ago[dead]
- tptacek 1y agoI don't know enough about the early history of the airline industry but there was very definitely a long series of huge bubbles in the railroad industry.
- antithesizer 1y ago[dead]
- daveguy 1y agoI would be very interested in reading about huge bubbles in railroad, airline, or any other industry. Do you happen to have references (genuinely asking; the original article didn't include any references)? Edits-- Found one: https://en.wikipedia.org/wiki/Panic_of_1893 https://en.wikipedia.org/wiki/Panic_of_1893 Another good one:https://en.wikipedia.org/wiki/Public_Utility_Holding_Company_Act_of_1935 https://en.wikipedia.org/wiki/Public_Utility_Holding_Company... (from cake_robot here: https://news.ycombinator.com/item?id=45056621 https://news.ycombinator.com/item?id=45056621) For reference, apple and spotify links to the Derek Thompson podcast in reply below (thank you!): https://podcasts.apple.com/us/podcast/plain-english-with-derek-thompson/id1594471023 https://podcasts.apple.com/us/podcast/plain-english-with-der... https://open.spotify.com/show/3fQkNGzE1mBF1VrxVTY0oo https://open.spotify.com/show/3fQkNGzE1mBF1VrxVTY0oo
- tptacek 1y agoThere was just a long Derek Thompson podcast with a Transcontinental Railroad scholar about this! (That's why I knew about it.) The whole subtext of that podcast was how eerily similar the Transcontinental Railroad was to AI (as an investment/malinvestment/prediction of future trends).
- gjsman-1000 1y agoFor my money, I used this analogy at work: Before AI, we were trying to save money, but through a different technique: Prompting (overseas) humans. After over a decade of trying that, we learned that had... flaws. So round 2: Prompting (smart) robots. The job losses? This is just Offshoring 2.0; complete with everyone getting to re-learn the lessons of Offshoring 1.0.
- draw_down 1y ago[dead]
- the_af 1y ago> Prompting (overseas) humans [...] After over a decade of trying that, we learned that had... flaws. I think this is a US-centric point of view, and seems (though I hope it's not!) slightly condescending to those of us not in the US. Software engineering is more than what happens to US-based businesses and their leadership commanding hundreds or thousands of overseas humans. Offshoring in software is certainly a US concern (and to a lesser extent, other nations suffer it), but is NOT a universal problem of software engineering. Software engineering happens in multiple countries, and while the big money is in the US, that's not all there is to it. Software engineering exists "natively" in countries other than the US, so any problems with it should probably (also) be framed without exclusive reference to the US.
- gjsman-1000 1y agoThe problem isn't that there aren't high quality offshore developers - far from it. Or even high quality AI models. The problems are inherent with outsourcing to a 3rd party and having little oversight. Oversight is, in both cases, way harder than it appears.
- hnbad 1y agoThe problems are actually not the oversight but the motivation: the point of outsourcing is cost-cutting, not finding equally talented developers elsewhere. This is especially evident in manufacturing: China can produce extremely sophisticated technology with excellent quality and well thought-out design but that's not what foreign customers look for when outsourcing to China. They want cheap, so they get cheap. "Oversight" in this case only solves the problem of the quality problems being intentionally hidden initially to get the sale. This might be where the actual parallels with AI models lie: quality of results going down after the initial launch because maintaining the quality is too expensive.
- sebnukem2 1y ago> hallucinations aren’t a bug of LLMs, they are a feature. Indeed they are the feature. All an LLM does is produce hallucinations, it’s just that we find some of them useful. Nice.
- tptacek 1y agoIn that framing, you can look at an agent as simply a filter on those hallucinations.
- th0ma5 1y agoYes yes, with yet to be discovered holes
- Lionga 1y agoIsn't an "agent" not just hallucinations layered on top of other random hallucinations to create new hallucinations?
- tptacek 1y agoNo, that's exactly what an agent isn't. What makes an agent an agent is all the not-LLM code. When an agent generates Golang code, it runs the Go compiler, which is in the agent's architecture an extension of the agent. The Go compiler does not hallucinate.
- Lionga 1y agoThe most common "agent" is an letting an LLM run a while loop (“multi-step agent”) [1] [1] https://huggingface.co/docs/smolagents/conceptual_guides/intro_agents https://huggingface.co/docs/smolagents/conceptual_guides/int...
- tptacek 1y agoThat's not how Claude Code works (or Gemini, Cursor, or Codex).
- rancar2 1y agoMy favorite quote to borrow: “Furthermore I think anyone who says they know what this future will be is talking from an inappropriate orifice.”
- koolba 1y agoReminds me of the classic Yogi Berra, “It's tough to make predictions, especially about the future”.
- th0ma5 1y agoTo me, it is more specific to say that many futurists don't take into account the social and economic network effects of changes taking place. Many just act as if the future will continue on completely unchallenged in this current state. But if you look at someone like Kurzweil, you can see the very narrow and specific focus of a prediction, which has proved to be more informative to me as a high bar of futurism.
- Towaway69 1y agoPredicting the future isn't about being correct tomorrow, rather it’s about selling something to someone today. An insight I picked up along the way…
- atleastoptimal 1y agoThere are only 3 things that I have strong empirical evidence for with respect to LLMs 1. Routinely some task or domain of work that some expert claims that LLM’s will able to do, LLM’s start being able to reliably perform that task within 6 months to a year, if they haven’t already 2. Whenever AI gets better, people move the goalposts regarding what “intelligence” counts as 3. Still, LLM’s reveal that there is an element to intelligence that is not orthogonal to the ability to do well on tests or benchmarks
- hugedickfounder 1y ago[flagged]
- lubujackson 1y agoI like the idea of AI usage comes down to a measurement of "tolerances". With enough specificity, LLMs will 100% return what you want. The goal is to find the happy tolerance between "acceptable" and "I did it myself" via prompts.
- manmal 1y ago> With enough specificity, LLMs will 100% return what you want. By now I’m sure it won’t. Even if you provide the expected code verbatim, LLMs might go on a side quest to “improve” something.
- bko 1y ago> Certainly if we ever ask a hallucination engine for a numeric answer, we should ask it at least three times, so we get some sense of the variation. This works on people as well! Cops do this when interrogating. You tell the same story three times, sometimes backwards. It's hard to keep track of everything if you're lying or you don't recall clearly so you can get a sense of confidence. Also works on interviews, ask them to explain a subject in three different ways to see if they truly understand.
- chistev 1y agoWho remembers that scene on Better Call Saul between Lalo, Saul, and Kim?
- Terr_ 1y ago> This works on people as well! Only within certain conditions or thresholds that we're still figuring out. There are many cases where the more someone recalls and communicates their memory, the more details get corrupted. > Cops do this when interrogating. Sometimes that's not to "get sense of the variation" but to deliberately encourage a contradiction to pounce upon it. Ask me my own birthday enough times in enough ways and formats, and eventually I'll say something incorrect. Care must also be taken to ensure that the questioner doesn't change the details, such as by encouraging (or sometimes forcing) the witness/suspect to imagine things which didn't happen.
- inerte 1y agoTriple modular redundancy. I remember reading that's how Nasa space shuttles calculate things because a processor / memory might have been affected by space radiation https://llis.nasa.gov/lesson/18803 https://llis.nasa.gov/lesson/18803
- anon7725 1y agoTriple redundancy works because you know that under nominal conditions each computer would produce the correct result independently. If 2 out of 3 computers agree, you have high confidence that they are correct and the 3rd one isn’t. With LLMs you have no such guarantee or expectation.
- catigula 1y ago>I’m often asked, “what is the future of programming?” Should people consider entering software development now? Will LLMs eliminate the need for junior engineers? Should senior engineers get out of the profession before it’s too late? My answer to all these questions is “I haven’t the foggiest” I just want to point out that this answer implicitly means that, at the very least, the profession is at least questionably uncertain which isn't a good sign for people with a long future orientation such as students.
- tomku 1y agoIt has never been anywhere close to certain, we just had 20 years of wild, unsustainable growth that encouraged people to cover their eyes and pretend the ride would go on forever. 20 years of telling everyone under the age of 30 that of course they should learn to code and that CS was the new medical or legal degree. 20 years of smugly acting like we are the inevitable future when we are, in fact, subject to the same ups and downs as every other career.
- mvieira38 1y ago(slightly off-topic) Students shouldn't ever have become so laser-focused on a single career path anyway, and even worse than that is how colleges have become glorified trade schools in our minds. Students should focus on studying for their classes, getting their electives and extracurriculars in, getting into clubs... Then depending on which circles they end up in, they shape their career that way. The thought that getting a Computer Science major would guarantee students a spot in the tech industry was always ridiculous, because the industry was just never structured that way, it's always been hacker groups, study groups, open source, etc. bringing out the best minds
- muldvarp 1y agoThe answers are even kind of contradictory. If the answer to "will we ever need software engineers again in the future?" is "we don't know" then the answer to "should I spend time and money to enter software engineering?" should be "no".
- crawshaw 1y agoA bubble is asset prices systematically diverging from reasonable expectations of future cash flows. Bubbles are driven by financial speculation. The claim in the blog post that all technology leads to speculative asset bubbles I find hard to believe. Where was the electricity bubble? The steel bubble? The pre-war aviation bubble? (The aviation bubble appeared decades later due to changes in government regulation.) Is this an AI bubble? I genuinely don't know! There is a lot of real uncertainty about future cash flows. Uncertainty is not the foundation of a bubble. I knew dot-com was a bubble because you could find evidence, even before it popped. (A famous case: a company held equity in a bubble asset, and that company had a market cap below the equity it held, because the bubble did not extend to second-order investments.)
- cake_robot 1y agoJust taking your first example, yes I believe you could characterize a lot of the investments in electrification as speculative bubbles even though there was underlying value that was borne out. https://en.wikipedia.org/wiki/Public_Utility_Holding_Company_Act_of_1935 https://en.wikipedia.org/wiki/Public_Utility_Holding_Company...
- insane_dreamer 1y agoBesides electricity (answered above), there were bubbles related to steel production in its early days intertwined with the railroad bubble that drove huge investments in steel production for railroad development; when the railroad companies crashed it brought down steel as well; ultimately both survived in a more subdued form (as with the dot coms). avation: likewise there was the "Lindberg Boom" https://en.wikipedia.org/wiki/Lindbergh_Boom https://en.wikipedia.org/wiki/Lindbergh_Boom which led to overspeculation and the crash of many early aviation companies
- utyop22 1y agoI want to re-phrase your definition. To me a bubble reflects a market disconnect from fundamentals - wherein prices go up steeply, with no help from the fundamentals (expected growth in base year cash flows and risk). Indeed there is a subtle difference between a bubble and speculation of what could come of a technology. But the two are connected because the effects of technology are reflected by investors in asset prices.
- ares623 1y ago> Other forms of engineering have to take into account the variability of the world. > Maybe LLMs mark the point where we join our engineering peers in a world on non-determinism. Those other forms of engineering have no choice due to the nature of what they are engineering. Software engineers already have a way to introduce determinism into the systems they build! We’re going backwards!
- delusional 1y agoI would rather say it like this: Very good, very hardworking engineers spent years of their lives building the machine that raised us from the non-determinism of messy physical reality. The technology that brought us perfect, replicable, and reliable math from sending electrons through rocks has been deeply underappreciated in the "software revolution". The engineers at TSMC, Intel, Global Foundries, Samsung, and others have done us an amazing service, and we are throwing all that hard work away.
- sodapopcan 1y agoAs pertaining to software development, I agree. I've been hearing accounting (online and from coworkers) of using LLMs to do deterministic stuff. And yet, instead of at least prompting once to "write a script to do X," they just keep prompting "do X" over and over again. Seems incredibly wasteful. It feels like there is this thought of "We are not making progress if we aren't getting the LLM to do everything. Having it write a script we can review and tweak is anti-progress." No one has said that outright, but it's a gut feeling (and it wouldn't surprise me if people have said this out loud).
- tptacek 1y agoThis is the 2025 equivalent of the people who once wrote 2000 word blog posts about how bad it was to use "cat" instead of just shell redirection.
- sodapopcan 1y agoThese are hardly equivalent. One is someone preferring one deterministic way over another. The other is more akin to arguing that it's better to ask someone to manually complete a task for you instead of caching the instructions on your computer. Now if the LLM does caching then you have more of a point, I don't have enough experience there.
- dionian 1y ago> I’ve often heard, with decent reason, an LLM compared to a junior colleague. But I find LLMs are quite happy to say “all tests green”, yet when I run them, there are failures. If that was a junior engineer’s behavior, how long would it be before H.R. was involved? A junior engineer can't write code anywhere nearly as fast. It's apples vs oranges. I can have the LLm rewrite the code 10 times until its correct and its much cheaper than hiring an obsequious jr engineer
- xmprt 1y agoJunior engineers get better and learn from their mistakes. An LLM will happily make then same mistake 10 times if you try asking it to do something similar in 3 months. In the long term, I don't think LLMs actually save you time considering the amount of extra time spent reviewing/verifying its code and fixing tech debt.
- swagasaurus-rex 1y agoIf an AI can’t write the code after two attempts, I’ve never had success trying ten times
- nomilk 1y ago> We should ask the LLM the question more than once For any macOS users, I highly recommend an Alfred workflow so you just press command + space then type 'llm <prompt>' and it opens tabs with the prompt in perplexity, (locally running) deepseek, chatgpt, claude and grok, or whatever other LLMs you want to add. This approach satisfies Fowler's recommendation of cross referencing LLM responses, but is also very efficient and over time gives you a sense of which LLMs perform better for certain tasks.
- _mu 1y agoWhat would such a workflow look like? I have Alfred but mainly just use the clipboard feature. I've tried to get into automation but struggled for inspiration. This one seems good! Are you just opening a browser tab?
- nomilk 1y agoGo to the 'Workflows' tab, make a new one with keyword of your choice (e.g. llm), and map it to open these urls in your default browser: http://localhost:3005/?q={query} https://www.perplexity.ai/?q={query} https://www.perplexity.ai/?q={query} https://x.com/i/grok?text={query} https://x.com/i/grok?text={query} https://chatgpt.com/?q={query}&model=gpt-5 https://chatgpt.com/?q={query}&model=gpt-5 https://claude.ai/new?q={query} https://claude.ai/new?q={query} Modify to your taste. Example: https://github.com/stevecondylios/alfred-workflows/tree/main https://github.com/stevecondylios/alfred-workflows/tree/main (you should be able to download the .alfredworkflow file and double click on it to import it straight into alfred, but creating your own shouldn't take long, maybe 5-10 mins if it's your first workflow)
- mauflows 1y agoI made something similar that uses alacritty / llm / tmux. The referenced script is also in the repo https://github.com/mjmaurer/infra/blob/main/home-manager/modules/alfred/workflows/ai/info.plist https://github.com/mjmaurer/infra/blob/main/home-manager/mod...
- revskill 1y ago[dead]
- skhameneh 1y agoThere are many I've worked with that idolize Martin Fowler and have treated his words as gospel. That is not me and I've found it to be a nuisance, sometimes leading me to be overly critical of the actual content. As for now, I'm not working with such people and can appreciate the article shared without clouded bias. I like this article, I generally agree with it. I think the take is good. However, after spending ridiculous amounts of time with LLMs (prompt engineering, writing tokenizers/samplers, context engineering, and... Yes... Vibe coding) for some periods 10 hour days into weekends, I have come to believe that many are a bit off the mark. This article is refreshing, but I disagree that people talking about the future are talking "from another orifice". I won't dare say I know what the future looks like, but the present very much appears to be an overall upskilling and rework of collaboration. Just like every attempt before, some things are right and some are simply misguided. e.g. Agile for the sake of agile isn't any more efficient than any other process. We are headed in a direction where written code is no longer a time sink. Juniors can onboard faster and more independently with LLMs, while seniors can shift their focus to a higher level in application stacks. LLMs have the ability to lighten cognitive loads and increase productivity, but just like any other productivity enhancing tool doing more isn't necessarily always better. LLMs make it very easy to create and if all you do is create [code], you'll create your own personal mess. When I was using LLMs effectively, I found myself focusing more on higher level goals with code being less of a time sink. In the process I found myself spending more time laying out documentation and context than I did on the actual code itself. I spent some days purely on documentation and health systems to keep all content in check. I know my comment is a bit sparse on specifics, I'm happy to engage and share details for those with questions.
- manmal 1y ago> written code is no longer a time sink It still is, and should be. It’s highly unlikely that you provided all the required info to the agent at first try. The only way to fix that is to read and understand the code thoroughly and suspiciously, and reshaping it until we’re sure it reflects the requirements as we understand them.
- skhameneh 1y ago
- iLoveOncall 1y ago> One of the big problems with these surveys is that they aren’t taking into account how people are using the LLMs. From what I can tell the vast majority of LLM usage is fancy auto-complete, often using co-pilot. This is a completely wrong assumption and negates a bunch of the points of the article...
- epolanski 1y agoHow is it wrong? Do you have any hard number and data to state that tab/auto complete are less popular than agentic coding?
- deleted 1y ago[deleted]
- daveguy 1y agoI would expect it to be almost tautologically true. It's the easiest to implement, and the most deployed. It has to be the most used, even if a lot of people don't know they're using it. It was the original success of LLMs -- code autocomplete. All of the agentic coding have autocomplete under the hood. Autocomplete is almost literally what LLMs do + post-processing.
- daviding 1y agoI get a lot of productivity out of LLMs so far, which for me is a simple good sign. I can get a lot done in a shorter time and it's not just using them as autocomplete. There is this nagging doubt that there's some debt to pay one day when it has too loose a leash, but LLMs aren't alone in that problem. One thing I've done with some success is use a Test Driven Development methodology with Claude Sonnet (or recently GPT-5). Moving forward the feature in discrete steps with initial tests and within the red/green loop. I don't see a lot written or discussed about that approach so far, but then reading Martin's article made me realize that the people most proficient with TDD are not really in the Venn Diagram intersection of those wanting to throw themselves wholeheartedly into using LLMs to agent code. The 'super clippy' autocomplete is not the interesting way to use them, it's with multiple agents and prompt techniques at different abstraction levels - that's where you can really cook with gas. Many TDD experts have great pride in the art of code, communicating like a human and holding the abstractions in their head, so we might not get good guidance from the same set of people who helped us before. I think there's a nice green field of 'how to write software' lessons with these tools coming up, with many caution stories and lessons being learnt right now. edit: heh, just saw this now, there you go - https://news.ycombinator.com/item?id=45055439 https://news.ycombinator.com/item?id=45055439
- tra3 1y agoIt feels like Tdd/llm connection is implied — “and also generate tests”. Thought it’s not cannonical tdd of course. I wonder if it’ll turn the tide towards tech that’s easier to test automatically, like maybe ssr instead of react.
- daviding 1y agoYep, it's great for generating tests and so much of that is boilerplate that it feels great value. As a super lazy developer it's great as the burden of all that mechanical 'stuff' being spat out is nice. Test code being like baggage feels lighter when it's just churned out as part of the process, as in no guilt just to delete it all when what you want to do changes. That in itself is nice. Plus of course MCP things (Playwright etc) for integration things is great. But like you said, it was meant more TDD as 'test first' - so a sort of 'prompt-as-spec' that then produces the test/spec code first, and then go iterate on that. The code design itself is different as influenced by how it is prompted to be testable. So rather than go 'prompt -> code' it's more an in-between stage of prompting the test initially and then evolve, making sure the agent is part of the game of only writing testable code and automating the 'gate' of passes before expanding something. 'prompt -> spec -> code' repeat loop until shipped.
- Scubabear68 1y ago"Hallucinations aren’t a bug of LLMs, they are a feature. Indeed they are the feature". I used to avidly read all his stuff, and I remember 20ish years ago he decided to rename Inversion of Control to Dependency Injection. In doing so, and his accompany blog, he showed he didn't actually understand it at a deep level (and hence his poor renaming). This feels similar. I know what he's trying to say, but he's just wrong. He's trying to say the LLM is hallucinating everything, but Fowler is missing is that Hallucination in LLM terms refers to a very specific negative behavior.
- ares623 1y agoAs far as an LLM is concerned, there is no difference between "negative" hallucination and a positive one. It's all just tokens and embeddings to it. Positive hallucinations are more likely to happen nowadays, thanks to all the effort going into these systems.
- Scubabear68 1y agoThis basically ruins the term “hallucination” and makes it meaningless, when the term actually describes a real phenomenon.
- ares623 1y agoThat's the point. It is meaningless. When it first coined, there were already detractors to the term, that it is an incorrect description of the phenomenon. But it stuck.
- anthem2025 1y agoNo, it’s just an attempt to pretend wrong outputs are some special case when really they aren’t. They aren’t imagine something that doesn’t exist they are just running the same process they do for everything else and it just didn’t work. If you disagree then I would ask what exactly is the “specific behaviour” you’re talking about?
- epolanski 1y ago
- tricky_theclown 1y ago.
- insane_dreamer 1y ago> I’ve often heard, with decent reason, an LLM compared to a junior colleague. But I find LLMs are quite happy to say “all tests green”, yet when I run them, there are failures. If that was a junior engineer’s behavior, how long would it be before H.R. was involved? Reminds me of a recent experience when I asked CC to implement a feature. It wrote some code that struck me as potentially problematic. When I said, "why did you do X? couldn't that be problematic?" it responded with "correct; that approach is not recommended because of Y; I'll fix it". So then why did it do it in the first place? A human dev might have made the same mistake, but it wouldn't have made the mistake knowing that it was making a mistake.
- nicwolff 1y ago> I’ve often heard, with decent reason, an LLM compared to a junior colleague. No, they're like an extremely experienced and knowledgeable senior colleague – who drinks heavily on the job. Overconfident, forgetful, sloppy, easily distracted. But you can hire so many of them, so cheaply, and they don't get mad when you fire them!
- vkou 1y agoExperienced and knowledgeable and they also believe in the technical equivalent of flat-Earthism in many, many non-trivial corners. And if you push back on that insanity, they'll smile and nod and agree with you and in the next sentence, go right back to pushing that nonsense.
- sho_hn 1y agoOne good yardstick, if one has to anthropomorphize, is that LLMs know and believe what's popular. If you ask it to do something that popular developer opinion gets right, it will do fine. Ask it for things that many people get wrong or just do badly, or can be mistakenly likened to a popular thing in a way that produces a wrong result, and it'll often err. The trick is having an awareness of what correct solutions are prevalent in training data and what the bulk of accessible code used for training probably doesn't have many examples of. And this experience is hard to substitute for. Juniors therefore use LLMs in a bumbling fashion and are productive either by sheer luck, or because they're more likely to ask for common things and so stay in a lane with the model. A senior developer who develops a good intuition for when the tool is worth using and when not can be really efficient. Some senior developers however either overestimate or underestimate the tool based on wrong expectations and become really inefficient with them.
- DrewADesign 1y ago“Here’s the code:” looks totally plausible but hallucinated libFakeCall. “libFakeCall doesn’t exist. Use libRealCall instead of libFakeCall.” “You’re absolutely correct. I apologize for blah blah blah blah. Here’s the updated code with libRealCall instead. :[…]” “You just replaced the libFakeCall reference with libRealCall but didn’t update the calls themselves. Re-write it and cite the docs. “ “Sorry about the confusion! I’ve found these calls in the libRealCall docs. Here’s the new code and links to the docs.” “That’s the same code but with links to the libRealCall docs landing page.” “You’re absolutely correct.” It appears that these calls belong to another library with that functionality:” looks totally plausible but hallucinated libFakeCall.
- anthem2025 1y agoIt’s funny how people acknowledge the railroads and the similarity to AI, but then jump back to comparing it to the internet when it comes to drawing conclusions. The internet build out left massive amounts of useful infrastructure. The railroads left us with lots of railroads that fell into disuse and eventually left us with a complete joke of a railway system. Made a few people so rich we started calling them robber barons and talking the gilded age. Are we going to continue to use the 60 billion dollar data centers in Louisiana when the bubble bursts? Is it valuable infrastructure or just a waste of money that gets written off?
- senko 1y ago> Are we going to continue to use the 60 billion dollar data centers in Louisiana when the bubble burst Why not? Tracks in the middle of nowhere are useless. Data centers in the middle of nowhere are useful, as long as energy, cooling and uplink are cost-effective. (Not saying Louisiana is in the middle of nowhere!)
- anthem2025 1y agoI don’t disagree, but was Louisiana chosen because it offers cheap energy/cooling/infra or because the government doesn’t bother to ask questions about the impact to thinks like energy and water usage (which is why we are seeing more and more reports of residents upset about pollution, water usage, and energy costs).
- sudhirb 1y ago> One of the consequences of this is that we should always consider asking the LLM the same question more than once, perhaps with some variation in the wording. Then we can compare answers, indeed perhaps ask the LLM to compare answers for us. The difference in the answers can be as useful as the answers themselves. There was once a coding agent which achieved SOTA performance on SWE Bench Verified by "just" running the agent 5 times on each instance, scoring each attempt and picking the attempt with the highest score: https://aide.dev/blog/sota-bitter-lesson https://aide.dev/blog/sota-bitter-lesson
- jeppester 1y agoIn my company I feel that we getting totally overrun with code that's 90% good, 10% broken and almost exactly what was needed. We are producing more code, but quality is definitely taking a hit now that no-one is able to keep up. So instead of slowly inching towards the result we are getting 90% there in no time, and then spending lots and lots of time on getting to know the code and fixing and fine-tuning everything. Maybe we ARE faster than before, but it wouldn't surprise me if the two approaches are closer than what one might think. What bothers me the most is that I much prefer to build stuff rather than fixing code I'm not intimately familiar with.
- epolanski 1y agoAs Fowler himself states, there's a need to learn to use these tools properly. In any case poor work quality is a failure of tech leadership and culture, it's not AI's fault.
- FromTheFirstIn 1y agoIt’s funny how nothing seems to be AI’s fault.
- johnnienaked 1y agoNo one seems to be able to grasp the possibility that AI is a failure
- raducu 1y ago> No one seems to be able to grasp the possibility that AI is a failure. Do you think by the time GPT-9 comes, we'll say "That's it, AI is a failure, we'll just stop using it!" Or do you speak in metaphorical/bigger picture/"butlerian jihad" terms?
- johnnienaked 1y agoI don't see the use-case now, maybe there will be one by GPT-9
- woah 1y ago> One of the consequences of this is that we should always consider asking the LLM the same question more than once, perhaps with some variation in the wording. Then we can compare answers, indeed perhaps ask the LLM to compare answers for us. The difference in the answers can be as useful as the answers themselves. This is what LLM "reasoning" does. More than "reasoning" in the human sense, it just reduces variance from variations in the prompt and random next token prediction.
- stillsut 1y ago> we should always consider asking the LLM the same question more than once, perhaps with some variation in the wording. Then we can compare answers, Yup, this matches my recommended workflow exactly. Why waste time trying to turn an initially bad answer into a passable one, when you could simply re-generate (possibly with different context) I wrote up an example of this workflow here: https://github.com/sutt/agro/blob/master/docs/case-studies/aba-vidstr-2.md#results https://github.com/sutt/agro/blob/master/docs/case-studies/a...
- oo0shiny 1y ago> My former colleague Rebecca Parsons, has been saying for a long time that hallucinations aren’t a bug of LLMs, they are a feature. Indeed they are the feature. All an LLM does is produce hallucinations, it’s just that we find some of them useful. What a great way of framing it. I've been trying to explain this to people, but this is a succinct version of what I was stumbling to convey.
- jstrieb 1y agoI have been explaining this to friends and family by comparing LLMs to actors. They deliver a performance in-character, and are only factual if it happens to make the performance better. https://jstrieb.github.io/posts/llm-thespians/ https://jstrieb.github.io/posts/llm-thespians/
- lagrange77 1y agoI'll steal that.
- red75prime 1y agoThe analogy goes down the drain when a criterion for good performance is being objectively right. Like with Reinforcement Learning from Verifiable Rewards.
- ACCount37 1y agoRLVR can also encourage hallucinations quite easily. Think of SAT: giving a random answer is right 20% of the time, giving "I don't know" is right 0% of the time. If you only reward for test score, you encourage guesswork. So good RL reward design is as important as ever. That being said, there are methods to train LLMs against hallucinations, and they do improve hallucination-avoidance. But anti-hallucination capabilities are fragile and do not fully generalize. There's no (known) way to train full awareness of its own capabilities into an LLM.
- neonspark 1y ago
- jedimastert 1y agoTalking to LLMs reminds me of my time as a system administrator for my university's physics department. A bunch of incredibly smart, PhD level people...just not in the subject they were talking to me about, most of the time.
- archeantus 1y agoIf all you’ve done with AI is just use it for autocomplete, you’re missing out big time. I built a slick react app using Lovable yesterday and then created a Node BE using Claude Code today. I told Claude Code to look through the FE code to understand the requirements and purpose of the site, and to build a detailed plan (including proposed DB schemas) for a system that could support that functionality. It generated a thousand line file with a robust breakdown of everything that needed to be done and at my command it did it. We went module by module and I made sure that each module Had comprehensive unit test coverage and that the repo built well as we went. After a few hours of back and forth we made 9 modules, 60+ APIs across 10 different tables, and hundreds of unit tests that all are passing. Does that mean that I’m all done and ready to deploy to prod? Unlikely. But it does mean that I got a ton of boilerplate stuff put into place really quickly and that I’m eight hours into a project that would have taken at least a month before. Once the BE was done I had it generate extensive documentation for the agent that would handle the FE integration as a sort of instruction guide - in case we need it. As issues and bugs arise during integration (they will!) the model has everything it needs to keep on track and finish the job it set out to do. What a time to be alive!
- computerex 1y agoI feel exactly the same way. Even this post by Martin Fowler shows he's an aging dinosaur stuck in denial. > I’ve often heard, with decent reason, an LLM compared to a junior colleague. But I find LLMs are quite happy to say “all tests green”, yet when I run them, there are failures. If that was a junior engineer’s behavior, how long would it be before H.R. was involved? I don't know what "LLM's" he's using but I just simply don't get hallucinations like that with cursor or claude code. He ends with this: > LLMs create a huge increase in the attack surface of software systems. Simon Willison described the The Lethal Trifecta for AI agents: an agent that combines access to your private data, exposure to untrusted content, and a way to externally communicate (“exfiltration”). That “untrusted content” can come in all sorts of ways, ask it to read a web page, and an attacker can easily put instructions on the website in 1pt white-on-white font to trick the gullible LLM to obtain that private data. Not sure why he is re-iterating well known prompt injection vulnerability, passing it off as a general weakness of LLM's that applies to all LLM use when that's not the reality.
- rootusrootus 1y agoFor a hot second I thought LLMs were coming for our jobs. Then I realized they were just as likely to end up creating mountains of things for us to fix later. And as things settle down, I find good use cases for Claude Code that augment me but are in no danger of replacing me. It certainly has its moments.
- jama211 1y agoFinally, an opinion on here that’s reasonable and isn’t “AI is perfect” or “AI is useless”.
- mexicocitinluez 1y agoOne of the things that has struck me as odd is just how little self-awareness devs have when talking about "skin in the game" with regard to CEO's hawking AI products. Like, we have just as much to lose as they have to gain. Of course a part of us doesn't want these tools to be as good as some people say they are because it directly affects our future and livelihood. No, they can't do everything. Yes, they can do some things. It's that simple.
- kaosoaksh 1y ago> doesn’t want these tools to be as good as some people say they are No, this is because that would mean AGI. And it’s obviously not that.
- jama211 1y agoAnd they’re similarly not as useless as others say they are, as they draw crosses in the air and hiss at the heathen tech to go away.
- mexicocitinluez 1y agoWhat? I didn't say anything about AGI. Why are you strawmanning this?
- 1y ago
- siffland 1y agoThe issue is a lot of times it takes a senior level programmer to reason with a LLM to get the results needed. What happens when there are no Juniors to replace the Seniors. I guess by then AI will be able to program efficiently enough. So far AI seems to be a great augmentation, but not a replacement.
- dwa3592 1y agoStop comparing LLMs to Humans for fucks sake. Anthropomorphization is fine as long as it serves a purpose, it doesn't in this case.
- paulcole 1y ago[flagged]
- thrown-0825 1y agoDont forget the gas lighting when they make a mistake and try to act like it was you.
- npn 1y agoI will believe his rant when Fowler creates any notable software.
- rhubarbtree 1y ago> Software Engineering is unusual in that it works with deterministic machines. Maybe LLMs mark the point where we join our engineering peers in a world on non-determinism. Recently some people have compared LLMs to compilers and the resulting source code to object code. This is a false analogy, because compilation is (almost always) a semantics preserving transformation. LLMs are given a natural language spec (prompt) that by definition is underspecified. And so they cannot be semantics preserving, as the semantics of their input is ambiguous. The programmer is left with two options: (1) understanding the resulting code, repairing and rewriting it, or (2) ignoring the code and performing validation by testing. Both of these approaches are assistive. At least in its current form, AI can only accelerate a programmer, not replace them. Lovable and similar tools rely on very informal testing, which is why they can be used by non-programmers, but they have very little chance of producing robust software of any complexity. I’ve seen people creating working web apps, but I am confident I could find plenty of strange bugs just by testing edge cases or stressing non functional qualities. The bigger issue is the bugs I can’t find because they’re not bugs a human programmer would create. Option (1) is problematic because LLMs tend not to produce clean code designed to be human readable. A lot of the efforts coders are making is to break down tasks and try to guide the LLMs to produce good code. I have yet to see this work for anything novel and complex. For non trivial systems, reasoning and architecture are required. The hope is that a programmer can write specs well enough that LLMs can “fill in the gaps”. But whether this is a net positive once considering the work involved is still an open question. I’ve yet to see any first hand evidence that there’s a productivity gain here. It’s early days. Option(2) is also difficult because there is a crucial factor missing in AI coding, the “generality of intent” as a human user. This is a problem, because the non-trivial bugs an AI produces are unlikely to be similar to those from a human. Those bugs are usually a failure of reasoning, but LLMs don’t reason in the same sense that humans do, so testing in the same way may not be possible. Your intuitions for where bugs lie are no longer applicable. The likely result is worse code produced more quickly, and that trade-off needs exploring. At the moment I think AI is useful for (a) discussions around design, libraries, debugging, (b) autocomplete, (c) agent analysis of existing code where partial answers are ok and false positives acceptable (eg finding some but not all bugs). Agent coding doesn’t seem ready for production to me, not until we have much better tooling to prevent some of these problems, or AI becomes capable of proper reasoning.
- dragonwriter 1y ago> My former colleague Rebecca Parsons, has been saying for a long time that hallucinations aren’t a bug of LLMs, they are a feature. Indeed they are the feature. All an LLM does is produce hallucinations, it’s just that we find some of them useful. This is an example of my least favorite style of feigned insight: redefining a term into meaninglessness just so you can say something that sounds different while not actually saying anything new. Yes, if you redefine "hallucination" from "produce output containing detailed information despite that information not being grounded in external reality, in a manner distantly analogous to a human reporting sense data produced by a literal hallucination rather than the external inputs that are presumed normally to ground sense data" to "produce output", its true that all LLMs do is "hallucinate", and that "hallucinating" is not a undesirable behavior. But you haven't said anything new about the thing that was called "hallucination" by everyone else, or about the thing--LLM output in general--that you have called "hallucination". Everyone already knew that producing output wasn't undesirable. You've just taken the label conventionally attached to a bad behavior, attached it to a broader category that includes all behavior, and used the power of equivocation to make something that sounds novel without saying anything new.
- scott_w 1y agoFowler is not really redefining "hallucination." He's using a form of irony that emphasises how fundamental "hallucinations" are to the operation of the system. One might also say "you can't get rid of collateral damage from bombs. Indeed, collateral damage is the feature, it's just some of that is what we want to blow up." You're not meant to take it literally.
- xyzzy123 1y agoYou might as well say it's interpolating or extrapolating. That's what people are usually doing too, even when recalling situations that they were personally involved in. I think we call it "hallucinating" when the machine does this in an un-human-like way.
- dapperdrake 1y ago
- TomBers 1y agoOne aspect of using LLM's to code that I have not seen mentioned is the "loss of attachment" to code in my projects. For example, working on my project (https://mudg.fly.dev https://mudg.fly.dev) I wanted to experiment with a new FE graph library. I asked an llm the fantastic (https://tidewave.ai https://tidewave.ai) and it completely implemented a solution, which turned out to be worse than the current solution. If I had spent many hours working on a solution I might have been more "attached" to the solution and kept with it. Perhaps this is a good thing? (https://en.wikipedia.org/wiki/Nonattachment_(philosophy) https://en.wikipedia.org/wiki/Nonattachment_(philosophy))
- raziel2p 1y agoHave you considered that LLMs are also biased against new languages and libraries, so the code quality will be worse compared to something more established regardless of what you personally think/feel?
- kookamamie 1y ago> We are still figuring out how to use LLMs, and it will be some time before we have a decent idea of how to use them well As per typical, Martin is so late to the party that he projects his own lack of understanding of the subject to be the general state of matters. People are, and have been for a while, using LLMs as very effective accelerators or augmentations of themselves. Claude Code, is a great example of something that can easily one-shot most trivial-to-average programming tasks. While some wait for the "bubble to burst", a lot of people are gaining significant benefits from using LLMs to boost their speed in software development. It's not a good idea to muddy the expectations for AI by mixing AGI and LLMs into the same bucket. LLMs are already useful for many purposes, even if they didn't lead to AGI, whatever the definition is.
- ChrisMarshallNY 1y ago> One of the consequences of this is that we should always consider asking the LLM the same question more than once, perhaps with some variation in the wording. Then we can compare answers, indeed perhaps ask the LLM to compare answers for us. The difference in the answers can be as useful as the answers themselves. That’s a useful tip.
- jonnycomputer 1y ago>All an LLM does is produce hallucinations, it’s just that we find some of them useful. Nice callback to "All models are wrong, but some of them are useful."
- mikewarot 1y ago>Should senior engineers get out of the profession before it’s too late? If they are actually using the engineering principles that their job description hints at, they'll probably be fine. Software Engineering is a growing field, the need for actual Engineers (you know, the kind that get licensed in other fields) is unlikely to shrink any time soon.
- siliconc0w 1y agoI started a greenfield project the other day and was excited to see how far AI has gotten for a real internal business use-case - not just a toy project. It was great at setting up the scaffolding and helping with some of the tests but overall it might have been a drain on productivity.. I had to repeatedly correct it and explicitly tell it how to do things (or hunt down examples in the codebase) or it would make up APIs or generate not particularly idiomatic code for our codebase (was writing golang fwiw). There is really a balance between spoon-feeding it the answers vs just abandoning it - the spikiness is really apparent where it can totally nail some things but then utterly fail what should be pretty simple (i.e there was a point where I wanted to refactor two similar functions into a shared library and it just couldn't do it)
- tony2016 1y agoThe developer is supposed to check and verify the code that an LLM creates. Ask for smaller requests in the prompts so you don't get too much code. Unit tests are verified and run by the developer. I don't know what he means by an LLM runs all the tests and gives green. It can fake running tests. I always the tests myself, whether they were created by an LLM or not.
- abcde-bbq 1y agoMore details on tolerance would be nice. Examples: 1: when building games for my kids, the tolerance is high so the LLM can implement a feature in one shot and I will manually test the feature (without bothering to review the code or generating unit tests or using a type-safe language). 2: building a demo at work has lower tolerance for errors, but is still good with LLM with a brief code review. 3: for production code, other processes can be in place to meet the even lower tolerance requirements.