3 ms·
This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can
by plaidfuji 17d ago
This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out.
> the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three:
1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc.
2. those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc.
3. those that can accept or already do by nature the costs of rigorous specification and validation: chip design, drug discovery, and other domains where failure on deployment is an existential concern.
the first two classes are price sensitive and arguably don't need the jump in reasoning quality you see going from cheap to frontier models. most of these firms will be best served by open models running on cheap hardware, perhaps even locally at the site of use. for the first and third classes, the type of fuzzy combinatorial search that has produced headline results in mathematics and security research seems more sensitive to agentic swarm width than reasoning capacity
…
This is just so on point. And for the third class (which I would extend to things like materials research as well), specification and validation are already by FAR the larger costs, so automating search and simulation is really not a massive game changer for the broader business.
- reedlaw 17d agoIn my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law (https://en.wikipedia.org/wiki/Parkinson%27s_Law https://en.wikipedia.org/wiki/Parkinson%27s_Law) is still in effect. You could even expand "Work expands so as to fill the time available for its completion" with "time and tokens available".
- 0c3ca83 17d agoWhy do you belive that agents wouldn't be able to take over product management, and generate prompts for the "software engineer" agents?
- reedlaw 17d agoBecause agents lack human judgment. At the very least there's a need for a human-in-the-loop with agentic processes. Otherwise, it's like running a coding harness with --dangerously-skip-permissions all the time.
- 0c3ca83 17d agoWhy do you think judgement is impossible to automate? What aspects of it do you think make it hard?
- noir_lord 17d agoIt requires general intelligence and we don't even have a good understanding of how our's works or a particularly good way of quantifying it. The counter argument is of course maybe you don't need to understand our kind of intelligence to create a different kind and that could well be true but then how do you determine if a system is intelligent. Unless the new system is intelligent enough to reason with us on our level in a way we can "see" is intelligent it becomes a philosophical argument. We also have a natural inclination towards anthropomorphising systems that mirror us, this is already a problem with LLM's and people overestimating their capabilities or forming actual emotional attachment. Then there are those of us who know more about how they work who in theory should be more immune to that and aren't. I added some stuff to my agent.md to make it sound less human and to communicate more like the machine be it is because I find the faked emotion extremely jarring. It can't be sorry, it's a set of numbers, it sits in the linguistic uncanny valley.
- Analemma_ 17d agoWhy do you think it requires general intelligence? The parent and I aren't being obtuse here: the history of artificial intelligence research is littered with examples of humans confidently declaring that task X requires general intelligence, then getting humiliated by a neural network doing task X better than humans a few years later. See Go, driving, art (you can complain about the quality of AI art, but it's winning competitions with human judges), etc. A priori I'm not sure why you would think being a PM at a FAANG, deciding what color the login button should be, is any different.
- bharatsuthar 17d ago>with "time and tokens available". That's like 1/4 of Codex users burning through their free resets just to build shinier todo apps.
- openfront 17d agoI think the fear is that softwaring engineering becomes a low-skill profession. If the models get good enough you won't need years of experience to be a decent programmer.
- reedlaw 17d agoExpectations around software quality will rise in tandem with gains in productivity so that the same experience curve continues to apply.
- poncho_romero 17d agoI don't think that's been true over the last decade though.
- zffr 17d agoIMO the fear is that software engineering becomes a _high_-skill profession. If the models get good enough to solve all of the low-skill problems, then we only need to keep around the people who are highly skilled. That means a fewer number of software engineers, and a difficult path to becoming someone who is highly skilled.
- entropicdrifter 17d agoIt's already a fairly high skill profession in general. The seemingly-lower-skilled roles tend to be the ones with higher design/creative requirements on the programmers. Agreed about the barrier to entry raising though, we're already seeing that in the glut of CS grads who can't actually get a programming job right now. Personally I suspect that this is a cultural issue more than an economic one, though. Companies need to alter their expectation about entry-level engineers and develop a culture of mentorship/apprenticeship so that advanced analysis, architectural, code review, and AI management skills are all passed down successfully.
- devonkim 17d agoMaybe I'm misunderstanding you but I believe I'd disagree with your first point on the perceived lower skill roles with higher design / creative requirements. I've seen a lot of the software industry being software factories rather than something closer to craftsmanship or even proper ABET-style engineering and AI is indeed better than the manually written slop that was encouraged in low-creativity / low-innovation environments such as QA and a lot of infrastructure engineering like defining CI pipelines. It seems reminiscent of the waves that hit industries like coal mining and manufacturing in the US except at a much faster rate and with different driving and resisting forces in automation. Programmer / engineer compensation in the US market at least has been looking bimodal for at least 12 years now and the first one looks even more devastated in terms of labor than the big tech companies based upon (lack of) job postings and from browsing my connections on LinkedIn.
- jimmaswell 17d agoAs a programming lead on a hobbyist video game project, we're reaping massive rewards being in category #1. I keep the architecture and important details in check while letting frontier models go ham. As alluded, it's a video game, not a life support system, so bugs are low-impact. But even better, the defect rate is actually the lowest it's ever been. Insidious bugs baked in by years of accumulated human error are trivial for Sol or Astra to untangle. This is the best time to be alive so far if you enjoy hobby game dev.
- openfront 17d agoI've found the same, the rate of bugs has dropped pretty dramatically after switching to ai generated code. I think it's partly because ai will write 1000s of lines of unit tests without complaining. I also have a github workflow where claude runs the /code-review command on every PR.
- sthuck 17d agoIn the world of web apps, I find the agent's ability to write good e2e tests to be a real game changer. Turns out with enough rigor you can write pretty stable mostly not flaky e2e tests. And even flaky ones are fixed quickly due to a fuck ton of assertions at every step. Test code looks like a mess, even more than the usual LLM code. takes a while to let it go. Test report looks beautiful though.
- bwestergard 17d agoWhen the test code "looks like a mess", how do you get assurance it is testing the right properties? Like you, I've found that LLMs can improve test coverage by decreasing the amount of developer time spent writing tests. But generally, it takes a lot of manual work to set up the initial testing framework, and even then, a lot of vigilance to ensure that what is actually tested corresponds to the description of the test.
- nonameiguess 16d agoI feel weirdly stuck on both sides of this. I haven't been in a full-time "code writing" roles for years, so my own personal usage of LLMs for anything is roughly nil right now, just because I continue to be able to get done all tasks I wanted to get done in the time I have using tools and techniques I already had available, and have seen little need to adopt new ones. The exceptions have been around things like using the text interface for diagramming tools to at least get the initial scaffolding set up without having to learn the specific quirks of that tool, then spending maybe an hour or two at the end to clean up and make it look nice and ensure it's actually coherent and has no mistakes. This is a task I have to do maybe two or three times a month, so it isn't a huge win, but it's not nothing. But that's digression. The short of it is I've seen enough to convince me these tools may as well be magic and a whole lot of tasks that break down into "produce media content of some sort" that has a well-defined goal and definition of correctness will be permanently sped up by automation. This includes a lot of software writing. At the same time, I shared the skepticism of estimates of economic impact and irrevocably changing the larger world. I'm a lot closer to the business side of the house these days, working with customers and prospective customers to identify use cases, reference architectures, pain points, feature requests, and bring this back to the development teams to attempt using real-world experience like this to inform how we design products. It's not product management as I'm focused more often on the nitty gritty technical details, not high-level user experience or roadmaps. But it gives me a great avenue into seeing what causes organizations to actually buy and/or adopt new software products, and the rate at which they can do that. And frankly, it isn't moving the needle much. They have the same budgets they always had, so they're not buying more, and our business is growing, but no faster or better than it grew before agentic coding became a thing. I always wonder because it seems the glowing success stories on Hacker News come in one of three varieties. It's the solo indie dev, usually targeting mobile app stores, who churns out dozens of roughly "will compile and doesn't immediately crash at runtime" apps in the time it used to take to complete one. It's the hobbyist, making software only they and maybe their immediate friends will ever use. Or it's startups, whose monetization model isn't monetization at all; it's just having something shiny to show investors in order to convince those with loose money to give enough to you personally that you can build up a nest egg whether or not your product ultimately ends up ever having a single paying customer. In my own business, a multi-decade, mature but not hyperscale company selling overwhelmingly self-hosted enterprise open source software, I can see the impacts on output. We have the same major products with the same release cadence. Each point release averages more new features than they used to, but also more regressions. It's overall a mixed bag. Non-technical product management staff is able to contribute code. We have a ton of new internal tools that nobody uses but they're there now. On the customer side, those that hinge decisions on wanting features that didn't exist yet are benefiting from getting those. Those that already had the features they want are losing from the greater rate at which regressions get through. The net business impact seems to be things have definitely changed qualitatively, but in purely financial quantitative terms, things are about the same as they were before. More code being committed to various git forges, but same headcount, same revenue, same margins, and same market cap.
- abletonlive 17d ago[flagged]
- Madmallard 17d agoUhhh what?
- pessimizer 17d ago> 2. those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc. The problem with this angle is that it is still absolutely terrible at doing call center/customer service work, and the profitability story is that the price is going to go up rather than go down. For repetitive physical labor in a controlled environment I'm slightly more bullish, but if you control the environment, you mostly don't need AI. You just use traditional deterministic methods, and send a person in when things get stuck or things are by nature irregular. > 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. Those who can accept failure cheaply can't necessarily detect failure cheaply. A ton of insane attempts will have to be picked through carefully to find the candidates for success, because the lack of a thought process makes AI bad in random, inhuman ways. This is basically a version of 3) that wishes away tests. It will be (and is) certainly helpful to replace interns and aid in rapid prototyping, but not because failure can be accepted, but because those are things that are tightly supervised. According to the world thus far, that is resulting in anything from -15% to +25% productivity gains. I'm not seeing it as a game changer simply because if it was, I'd expect to have seen a lot more useful, original software products by now and I haven't. I've just seen old ones get buggier or rewritten in Rust. I'm only buying 3): when you just want a machine to randomly enumerate through a search space looking for things that make the carefully constructed tests pass. That's a very good thing, though. But as you say, it's not a game changer because you still have to write the tests. but > 5. navier-stokes and statements in pure mathematics like it are the absolute best case scenario for agentic work against rigorous specification. the theorem statement itself is already a rigorous specification. it has undergone decades of auditing by the mathematical community and its rendering in lean is a straightforward translation defined in terms of battle-tested mathematical objects from mathlib. the verifier, the lean theorem prover, has been extensively audited and specifically designed to avoid the types of unsoundness that would make it vulnerable to reward hacks. even lean and theorem provers like it are not invulnerable: soundness bugs have allowed LLMs to launder bogus proofs through the proof kernel before and it is not improbable that more such bugs exist. this is the rosiest setup; the vast majority of human knowledge work does not look like this. i'll comment below on the few areas of knowledge work that do resemble pure mathematics in this respect. This is the real deep point, and one I've been repeating since I heard Navier-Stokes was a fraud. This is exactly where I expected that LLMs would do well, and they are not. It shows that I have a basic misunderstanding of LLMs, and that misunderstanding is causing me to think that they have more potential than they have actually shown. Maybe the nature of the architecture, where it picks out features, intrinsically limits its ability to search a solution space? Maybe the fact that they modally predict what someone might say, and nobody has said a thing as of yet (when many people were knowledgeable enough to have, if it is correct), means that the LLM is not going to say it either? Maybe the fact that it consumes all information and blends it in a structured way, instead of synthesizing an entire space from a relatively very small amount of input like a human does, means that it won't ever accidentally synthesize something that can't be pieced together from things that have already been said? Is its accuracy its flaw, where a human's "mistaken" synthesis might ultimately correct everyone's understanding? Really not beating the charge of being a stochastic parrot. It might just be that we were underestimating stochastic parrots; if a million monkeys on a million typewriters were all getting treats when they satisfied a trainer who wanted to see a new work of Shakespeare; they could look at his published work, and they could watch each other type and when each other got treats; whenever they successfully spelled a word or put words into an intelligible phrase, that was made into a keyboard key for a group of sentence monkeys, and the successes of the sentence monkeys were made into keys for the paragraph monkeys, etc... could you get something that passed for mediocre, drunken Shakespeare in a thousand years? Or maybe even 10?
- visarga 17d ago4. those who do reversible work, where cost of error is tolerable
- Capricorn2481 17d agoI think using AI for customer service is really, really ineffective, and it's plain to anybody that has interacted with it. It's a shame that we've had decades of shitty customer service from companies that have the most responsibility and resources to do it, and people in tech have shrugged their shoulders saying "it's unreasonable to expect people to scale up their customer service! They gotta make money!" Now AI customer service has arrived and things are worse, not better. Now we can't even get in contact with anyone because an agent redirects us to documentation instead of directing us to a human being. The future, when a company steals your money, is making viral posts online trying to get enough public shame going that you will maybe get some of that money back.
- vanuatu 17d agoI've seen it be very effective (moreso than human agents at resolving issues) but its extremely implementation sensitive, so you're more likely to encounter a bad one in the wild than a good one. There are voice agents picking up phones that the vast majority of people don't realize is an agent
- autoexec 17d agoAI is perfect for tasks if you don't care about them being done correctly. Many companies just don't care about providing effective customer service and really don't want to provide it at all.
- kbelder 17d agoI feel that attempts to replace customer service with AI have and will continue to fail. It's not able to do that for both technical reasons (the AI is still to dumb) and business reasons (the AI is usually deliberately scoped to be impotent at helping customers with any important issue). But, when properly used AI can reduce the customer service load. It can handle a lot of simple questions and status updates. There just needs to be a human available when there's any request that escalates past that simple case.