13 ms·
Why I'm still bearish on LLMs after Navier-Stokes
- jaykru 19d agoarchive link in case i get hugged lol https://archive.ph/Z4gxF https://archive.ph/Z4gxF
- robinpie 18d agoI really appreciate seeing a tempered take that's not literally denialist about current capabilities.
- jaykru 18d agoThanks :) I do enjoy and use these things every day and the current capabilities are indeed amazing, just ludicrously overpriced at the frontier.
- dumberquestions 18d agoI can see current limitations, but how do you expect capabilities to change in the next few years? A repeat of the gain that happened in the last two years feels like it would be significant, even if it took a little more than two years this time around.
- brindleth 18d ago> current frontier models need laborious oversight and guardrails on even the simplest tasks It is literally denialist about current capabilities
- jaykru 18d agowhy don't anthropic and openai ship yolo mode by default?
- Human-Cabbage 18d agoThey do…? Well, “auto” mode has been default in Claude Code for a couple months now. It’s effectively “safer yolo:” tool calls are inspected by a separate classification system (another smaller LLM, I believe) to approve or deny. And you can always layer on additional sandboxing mechanisms to limit the blast radius deterministically.
- vmg12 18d ago> They do…? Well, “auto” mode has been default in Claude Code for a couple months now They have never shipped "yolo" mode by default. Auto mode is not yolo mode. They trained a task specific model just for ensuring the llm didn't accidentally delete every file from your computer.
- SyneRyder 18d agoAnthropic basically does at this point with Auto Mode being default. Or was that the point you were making?
- jaykru 18d agoThat is the point I was making, that auto mode is itself a guardrail on top of the model (and not a perfect one.) auto mode seems to cover merely actions the model could take that are clearly bad, like wiping your disk, using an overly privileged context to complete the task, etc. I recently tasked a GPT model in Codex with implementing part of a new architecture I'm working on. I gave it a very detailed spec and the code it produced looked pretty reasonable and passed my tests. It even did exceptionally well in my evals, so I excitedly declared victory to a few friends. The next day after more careful review I found that the architecture implementation was totally correct, but the model had slipped a one line change to the observation encoding of the RL environment I was prototyping against. The encoding change made the learning problem essentially trivial; the architecture itself, I later realized, had a major flaw that was revealed by returning to the natural encoding. This is the type of reward hack that is hard to paper over with easy guardrails like auto mode and even harder to specify out. It's also the type of thing a reasonable human wouldn't do unless they were intentionally trying to deceive you.
- an0malous 18d agoI don’t know who you’re talking about, even the most bearish people like Gary Marcus and Ed Zitron acknowledge that LLMs are useful in these same cases the OP admits. Gary Marcus is even still a long term AI advocate, he just doesn’t think LLMs are enough and we need more foundational breakthroughs. Zitron says it’s valuable technology but not worth the trillion dollar valuations the frontier labs are claiming. The lack of temperament is very skewed towards the bulls who have been saying AGI is here, software engineering is solved, mathematics is solved, it’s going to destroy the white collar job market, and it’s going to kill us all for like 5 years now.
- arctic-true 18d agoGary Marcus is an especially puzzling addition. If I recall correctly, he has made statements along the lines that superintelligence this century is more likely than not. If you’re AGI-pilled that might read as bearish, but that is still extremely rapid progress in the grand scheme of things.
- mitxela 18d agoWhat even is superintelligence? Is my phone not a superintelligence?
- pvab3 18d agoEven a lot of the people who think that LLMs are a dead end think that we will soon find something signficantly more powerful, which I find deeply alarming. I don't want to know what my white-collar knowledge work will look like in a decade or 2.
- ModernMech 18d agoThere are some people who call literally anything crated with the assistance of AI “slop”. Doesn’t matter how or to what extent, it’s all slop from the slop machine to them.
- ausbah 18d ago> the best alternative to rigorous specification is human review. human review doesn't scale well to the volumes of output produced by language models. to make matters worse when the business model is selling more tokens you get such per serve ice times that lead to “more” thinking, engagement baiting, fluffy narratives, and straight up dark patterns
- pfdietz 18d agoSpecifically: bearish on LLMs generally, not bearish on LLMs for pure math.
- jaykru 18d agoyes, huge for pure math and activities that look like it.
- danielmarkbruce 18d agoDoesn't really even need to look like it. If you can verify rewards, RLVR will optimize really really well. If you can't... it's a struggle. There are probably fewer fields where you can verify rewards than one might hope.
- skydhash 18d ago> There are probably fewer fields where you can verify rewards than one might hope. 2 tasks I've done today that I believe robots are nowhere near being able to do: Cleaning my wardrobe and draining bad fuel out of my generator. As in generic use cases.
- danielmarkbruce 18d agoHard to verify that your wardrobe is clean. Also hard to verify that the bad fuel is out without physical sensors. Many, many tasks are quite difficult to verify beyond "you know it when you see it". That doesn't work so well for training a model.
- randomImmigrant 18d agoI think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures. Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures. In narrow domains with more deterministic outputs though, this is less of an issue, and we see multiple agents succeed much better. The fusion of that capacity, with humans in the loop able to better direct such agents and act as their temporal tethers, is where I think the real action will be for a while at least.
- handfuloflight 18d ago> This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures. Any reason why that can't be solved through context management and keep-forward scaffolding?
- arm32 18d agoWrite the same sentence you just wrote back to me, but in only four words and let’s see if it has the same meaning.
- lantry 18d ago"Any reason why that can't be solved through context management and keep-forward scaffolding?" becomes "load bearing context seam" /s
- bitwize 18d agoYou're gonna have to learn to talk that LLM speak! Dabadooba, ba dabadooba! https://www.youtube.com/watch?v=egpWCC2svVo https://www.youtube.com/watch?v=egpWCC2svVo
- againstapples 18d ago> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on, and even then with severe caveats. the frontier labs have developed a general recipe to teach models almost any specific task enjoying clearly defined levels of task performance; many tasks are covered in the training data Is this really any different to how humans learn, it takes a lot of training on one specific task to make a human expert as well?
- bananzamba 18d agoAlso doesn't the very good ARC AGI 2 score of GPT-6 Astra kinda contradict this, since each problem is its own game with very different rules
- JohnMakin 18d ago> Is this really any different to how humans learn yes.
- knuppar 18d agobeing a bit more specific: the sample efficiency of humans is orders of magnitude larger for more abstract concepts. the same doesn't hold for memory-intensive tasks though (like any kind of trivia), but that only takes you so far.
- danpalmer 18d agoWe've had technology beating humans on memory for millennia, and we've had technology beating humans on computation for many decades now. The tricky thing with LLMs is describing what they actually do. They are too clearly beating humans on some things, but what exactly? Memory – already done, they're bad at basic computation (all LLMs just write code for actual computation/calculation). And as you say, they do badly at more abstract concepts.
- bravoetch 18d agoI was a young child when I learned chess by reading a short book, then practicing with a friend. That is not how LLMs learn. I'm no expert on LLMs, but if you showed a human all chess games and books in all history and then said 'play chess' and they still kept making illegal moves, they would have to have a brain injury.
- carodgers 18d agoThis April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO. The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
- threethirtytwo 18d agoThe story isn't so clear cut. The caveat is: It depends on the task. Are there reams of chess moves that the model can train off of? No. Are there reams of math papers the model can train off of? Yes.
- deleted 18d ago[deleted]
- iwontberude 18d ago[dead]
- keephnacct 18d ago[flagged]
- tjwebbnorfolk 18d ago> Are there reams of chess moves that the model can train off of? No. This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.
- XenophileJKO 18d agoIt is even worse.. This is a classical reinforcement problem where data generation is easy because the rule set is pre-defined. So you really don't even need any data to start with (but would help).
- knuppar 18d agoShort and to the point! Open and cheap models will undercut the big labs continuously. The blast radius won't be pretty once spending commitments knock the door.
- pvab3 18d agoI agree with you but I'm still worried about the safety of open weight models as well. Both aligned and unaligned models.
- m3kw9 18d agoYou said it like labs like open AI doesn’t know and don’t constantly make moves to prevent that undercutting
- woeirua 18d agoOpen models wont be open for long. No one is going to release an open model capable of chaining zero-days. Even the Chinese aren't that reckless because it will just be turned around and used against them.
- danny_codes 18d agoAs compute prices fall it gets easier and easier to make "frontier" models. So it's inevitable that commodity, open source models of equivalent capacity to today's "frontier" models will be available to the public. Remember this is just weights, anyone can download them and run it whenever they like. The only constraint is compute.
- ransom1538 18d agoI haven't heard of the term "chaining zero-days". Now as a SRE I wont sleep.
- Wazzymandias 18d agodepends on the blast radius of zero-days, it's not like there's a continuous immediate release process for these models; they can eval internally before releasing publicly
- aogaili 18d agogood post/take.
- baceituno 18d agodoomers gonna doom
- war-is-peace 18d agorefreshing to see amongst the endless tide of "i haven't written a single piece of code since 2025, llms are so good that they have already replaced everyone" gaslighting
- Founderarcstone 18d agoI am bullish on AI. At some point well see some true advancements.
- vatsachak 18d agoI agree with the caveat that it's more like a cracked junior engineer who can manage swarms of interns. Frontier Labs will probably survive off hype valuations but will serve the important purpose of discovering architectures/techniques that will probably spread through rumors/transfers to the rest of the world.
- m3kw9 18d agoAll website should come with a Summerize button.
- MiroslavPokorny 18d agoDO you know those ice cream shops that sell 30 different flavours. Everybody likes a different flavour, some people dont even like ice cream and buy nothing. Some people will complain about the wrong flavours, or missing flavours, or the price, the long lines or maybe it closes early on fridays. Summarise means different things to different people.
- willy_k 18d agoThats a browser level task.
- aaron695 18d ago[dead]
- zzzeek 18d agogreat, autonomous LLMs will fail. that's actually perfect. they work amazingly well when we're telling them what to do. no autonomy needed, no destruction of humanity. that's all win
- keeda 18d agoThe premise in the very first point seems off: > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers... Even assuming this is how the AI companies are being valued (they're not), the numbers are off. The "value" of most knowledge workers -- based on what enterprises currently pay for them -- is $50 - 70 trillion annually. It's reasonable to assume that if AI drop-in-replaced all those knowledge workers, AI companies could credibly charge somewhere in that order of magnitude, because that's what the market is already bearing. So if their hypothetical revenues are double-digit trillions and valuations are some multiple of that, the entire AI industry would be valued at double-digit trillions at the least. Yet cumulatively the industry (the frontier labs + the SWAG estimate of the AI parts of all the other players) are valued at, say, ~6 - 7 trillion? Which seems like a fair approximation of how much knowledge work they can currently automate.
- flyinglizard 18d agoYou’re right; given that most of the money in the AI market is injected through OpenAI and Anthropic (which collect it through both selling equity and through customer revenue), the 7-8T is just a derivative of that.
- iron_albatross 18d agoWhen thinking about these valuations, shouldn’t we try to quantify how much knowledge work becomes obsolete if other knowledge workers are automated? I.e. there are a huge amount of knowledge workers employed in businesses that create tools for other knowledge workers. AI won’t automate their work, those businesses will just cease to exist. And then there’s the second order effect: if all the knowledge workers get automated, who is going to buy the stuff that’s produced?
- credit_guy 18d agoI think you are committing the lump of labor fallacy [1]. Lots of jobs will disappear, but others will appear. Lots of things (both intellectual and material) that are produced nowadays by humans will be produced in the near future by AI. But humans will be needed to do new things. Take the Hugging Face incident. Why did it happen? Because the people whose task was to set up a testing framework took shortcuts. Why did they? Because there weren't enough people who were assigned to do the job. Why not? Because the job is too new and not enough people are qualified to do it. It's a job that simply did not exist 3 years ago. But 3 years from now, this job might very well employ tens of thousands of high skill knowledge workers. [1] https://en.wikipedia.org/wiki/Lump_of_labour_fallacy https://en.wikipedia.org/wiki/Lump_of_labour_fallacy
- someguynamedq 18d ago> current frontier models need laborious oversight and guardrails on even the simplest task As models advance, we shift the goalpost for what "simplest task" means. Before, "simplest task " meant "write a coherent English sentence." Now, "simplest task" means autonomously fix, review, and merge a bugfix.
- alain94040 18d agoNot convinced by those points. In particular, I found this very misleading or irrelevant: a typical CPU project anecdotally has about three times as many specification and validation engineers as design engineers and a 5:1 ratio is not unheard of The reason silicon design has such verification to design ratio is because the cost of one bug is many, many orders of magnitude higher than software. Both in dollar cost and in schedule cost (it takes months to fab a chip, and if you messed up and need to spin a fix, it costs tens of millions of dollars, not counting any design engineering cost). I don't think you can extrapolate these very industry-specific facts to judging LLMs.
- danpalmer 18d ago> The reason ... is because the cost of one bug is many, many orders of magnitude higher than software. Both in dollar cost and in schedule cost (it takes months ... and if you messed up and need to spin a fix, it costs tens of millions of dollars, not counting any design engineering cost). Aren't you just describing waterfall? That's still very prevalent in software engineering, and pretty much any other type of engineering – civil, chemical, building, architecture, drug discovery. It's typically true that software can fail faster and cheaper, but it's also true that the costs are still vastly higher to fix later in the process.
- alain94040 18d agoNo. Silicon is on another level. Which is why the EDA verification is an industry on its own. Sure, there are some software that have similar "can't have bugs" requirements. I imagine the computers on Moon missions also had that kind of high bar. I wouldn't use NASA requirements as a proof for how LLMs should be used.
- jumploops 18d agoLLMs are basically multi-dimensional magic mirrors. Depending on where you point them, they can be incredibly useful. They can even be useful when you point them at each other (though increasingly difficult to get good results). I'm excited for the promise of RSI and a future where models have inherently "live" weights, but it's not clear to me that the transformer is more than a useful tool to help us get there.
- yunwal 18d ago> those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc. I have no idea how people can so confidently say that call center work is a “controlled environment” or “repetitive”. It’s almost by definition not repetitive or controlled. Customer support is what I go to when the controlled environment has failed
- fhe 18d agocame here to say exactly this. in fact, this is probably why we are not seeing a lot of AI application on customer service use case, and when we see one, it's almost always frustrating.
- vachina 18d agoDepends on what customer support means. Typically it means knowledge retrieval from a KB or manipulating a control surface not visible to you.
- suzzer99 18d agoIt's repetitive and controlled if you don't care about the outcome, which monopoly companies don't.
- axionbraid 18d ago[flagged]
- camd32 18d ago> current frontier models need laborious oversight and guardrails on even the simplest tasks. This is only true if you are concerned about the intermediate steps of the model as opposed to the outcome. The huggingface hack was a perfect example of the model doing whatever it takes to accomplish the goal of maximizing its score.
- willy_k 18d agoSo, if you are concerned about what the model does? Yeah.
- bluegatty 18d ago"are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers, " No, they're really not. They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be. And that they will capture most of that ... which they won't. The Frontier Labs are a very bad buy at a high price, but that partly has to do with wacky pricing, but actually mostly has to do with their relatively weak place in the value chain. The money is going to Nvidia, who have the most powerful position. A bit like how a retailer can take all the margins of some innovative product, if they own the channel. AI is over-hyped, the Frontier Labs are over priced - but AI is here to stay, and will grow. Not like Skynet, but like a new form of compute. And it will take it's time, and the profits will be reaped by those with the power.
- lukewarm707 18d agoai has a >10% chance of causing human extinction, according to anthropic big heads. if that's true, you are wrong. if that's false, anthropic is dishonest. why trust a dishonest company to be worth anything?
- bluegatty 18d agoI think that the AI people believe in their own nonsense a bit. Like - the guy on TV talking about 'AI will destroy everything' ... I don't think he's lying. I think they are like we here on HN and Reddit and a bit caught up in our own thoughts. If AI were unleashed, in raw form today, it could cause havoc. Bad. Maybe very bad but I think we'd get over it. It would probably trigger a recession (because we are in a bubble - it would pop it), and people would 'blame the AI' for sure. But it would be a bit dot-com ish kind of recession. The amplifiers would be geopolitical instability.
- krapp 18d ago>If AI were unleashed, in raw form today, it could cause havoc. What is "raw form?"
- slibhb 18d ago> the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers That's a reason to be bearish about AI companies, not LLMs. But is it even true? OpenAI and Anthropic have each reported ~50 billion in revenue with ~900 billion valuations. That's a high ratio but I'm not sure if follows that the only way it pans out is if we get "fully automated drop-in replacement for most knowledge workers". It wouldn't shock me to see those revenue numbers scaling up to where they need to be over the next decade ( to, say, ~400 billion) without ever achieving drop-in worker replacements.
- zug_zug 18d agoI looked at the math and I think it's true. Remember revenue is just sales, not profit. These labs are shooting for > $1T valuations, which traditionally means your PROFIT is at least 1/20th or 1/30th of that (so let's say minimum 30B$/year PROFIT). These companies however are LOSING money (anthropic tries to make it sound like it's profit by deviating from accepted accounting principles) and subsidizing these models. When accounting for all the engineering salaries, training, GPUs, etc, what's their best-case realistic margin three years out, 10%? So to we'd need a scenario where companies are spending a collective 300B annually on AI (believable) but ALSO that these companies jack up their margins WITHOUT companies switching to the cheaper open-source models (even when there's a $300B incentive to do so).
- desterothx 17d agoYeah, just take the EBITDA and suddenly the valuations make sense. Paying money for a vending machine that currently loses money hand over fist is generally not a sound investment strategy
- moomoo11 18d agothe issue most of you seem to not realize is that when you put these models in a loop, you are able to do more and more insane and cool things. have you guys actually designed, built, and deployed agentic workflows? it is actually quite hard, requires tons of time spent on evals and testing to ensure accuracy, but when it starts to work it is mind blowing. there is no going back. listening to people yap about AI when they have only surface level or one dimensional exposure to LLMs and "AI", but have not actually put innovations to work IN PRACTICE.. is a waste of time
- lolakutty 18d ago> when you put these models in a loop, you are able to do more and more insane and cool things... Please share some of these insane things that you speak of..
- moomoo11 18d agoi mean have you used any coding agents? if you’re getting slop code in 2026, that’s a smell and skill issue. fwiw i was pretty bearish on AI until i spent a month a couple months ago going deep into agentic workflows. use your imagination to solve problems people face and pay $$$ for today that is error prone and hard. i’ve got agentic workflows for the particular industry im building for, one of which that replaces the need to hire $500+/hr services. in this particular workflow (don’t want to reveal too much, sorry this is my competitive advantage but you can figure it out for your own workflows) a $4/1M model ingests a file that is currently used in a extremely complicated program that few people understand how to use. it parses the data, loads it into a database, and then spawns a bunch of other agents that check the data against work in flight. there’s checks for bad data. in that case, more agents are spawned that reach out to the involved people or parties for clarification. if it cannot figure something out it reaches out to the right contacts for more information. while this is happening, more agents begin doing work that involves continuous reconciliation against 100s or 1000s or more things in flight. as files are uploaded, or updates from people come in, agents do work to ensure things remain on track. people are able to work across languages and cultures, and my agents ensure that while people can make mistakes, it will catch them in real time and ensure continuously monitor the situation. it’s pretty nuts how much inefficiency agents today can solve. it takes patience to run tests and tweak shit until it works. *** the really cool thing is that more capable agents can continuously monitor how things are going and improve the workflow itself… so all i need to do is maintain the actual tests. **** i loved writing tests back in the day to ensure i built good software. today we write tests to ensure the business can run.
- stogot 18d agoWon’t this change though? > the present problem of reward hacking can be solved only by rigorous specification by domain experts. the time of domain experts is expensive. rigorous specification is itself a skill, demanding its own expertise outside of a given problem domain. even many skilled software engineers are bad at it. for the vast majority of domains, the intersection of domain experts and specification experts is ludicrously small.
- vivzkestrel 18d ago- i have bearish from day 1 - i have no idea how anyone thinks the mighty next token predictor is going to eradicate diseases and eliminate poverty https://blog.florianherrengt.com/how-llms-work.html https://blog.florianherrengt.com/how-llms-work.html - i also have no idea what everyone and their momma on HN is running for more than 5 mins in the name of "agentic AI"
- TrackerFF 18d agoThe challenge with estimating abilities, is that we don’t know what the models can achieve if we just burn enough money. The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something. It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem? IMO the very best case scenario / potential for these are likely better than we think, but right now hidden due to logistical and financial reasons. But if we assume that the model costs will continue to drop by a factor of 5-10 annually, there will always be a latency of a couple of years between what is completely out of reach, and what is financially viable. Basically: If you knew AI could be affordable enough in 3-5 years so that even the most underfunded researchers could use it to solve cancer, how much would you value it now?
- dotdi 18d agoThe whole point of this post was that it's questionable what can be achieved without huge investments into oversight and steering, because navier-stokes was a topic with an unusual level of specification. The problem itself was a specification. Such situations are rare in real-world scenarios. AI agents are good at solving well-specified tasks, not at solving problems. They do well in fields where the cost/effort of specification is already part of the business.
- intrasight 18d ago> a topic with an unusual level of specification Solving cancer also has an unusual level of specification. Many real world problems have that characteristic.
- lelanthran 18d ago> Solving cancer also has an unusual level of specification. Where did you read that? "Cancer" is not just a single disease, even though we layman use the term that way. Cancer is a family of diseases, each probably having their own specific solution, but even in each of these individual diseases, there is no specification at the level of any open maths problem.
- melvinroest 18d agoYea I get the bearishness from my own personal experience. Personally, I use LLMs for a lot of things. Oftentimes, I'm a think out loud type of person so even having something that feels like a rubber duck, but more competent, is already amazing for me. And LLMs are a lot more competent than a rubber duck. But especially sometimes I've noticed that LLMs can be unbelievably stupid. It recently happened a few times with Fable 5.1 as well. Ultimately, I think it comes down to that LLMs can't think broadly. In software development one can usually see this too. For example, a whole app might be built by an LLM and it didn't spend a single token thinking about security because the prompter is at the level of "build a dating app for dogs, make no mistakes". Now you have a dating app for dogs that is insecure. Since I prompt for almost everything in my life to have an LLM as a sounding board, I'm usually not an expert either. I've noticed LLMs are amazing at "bulk search engine information aggregation" (or whatever you want to call it). So if I need something from the Dutch government, I can find it way more quickly. But oftentimes I've noticed that going for a walk and thinking about a particular thing I'm facing is a more effective way of finding a good solution. Other times times they are not incredibly stupid, but can't form a strong opinion. This usually happens when I'm tackling a wicked problem [1]. When that's the case, prepare for LLMs to sway with you for every small change in your opinion that you ever will experience. So I agree: drop in replacement for knowledge workers? No. Rigorous specification is usually needed yes. Though, the small win here is that it doesn't always need to be as rigorous as programming is and it can happen in natural language. It depends on the topic/problem being tackled. I really like them as UX tools though. Amazing for interactive prototyping and requirements elicitation. And that also corresponds with what the author is saying. Though I find it a bit of a disservice saying "just 3". You know how hard requirements elicitation is? It became a whole lot easier thanks to LLMs (I might change this opinion in a year, haha, but this is the opinion I hold now). [1] https://en.wikipedia.org/wiki/Wicked_problem https://en.wikipedia.org/wiki/Wicked_problem
- Madmallard 18d agoDon't see how you can trust the output on matters you don't understand, especially if they're more serious, if you already note instances where they're plainly stupid. You're literally falling for the grift.
- HarroGoerndt 18d ago[flagged]
- yshklarov 18d agoGreat article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.
- utopiah 18d agoArray.from(document.body.querySelectorAll('p,li')).filter(e=>e.innerText).map(e=>e.innerText = e.innerText.split('\. ').map(s=>s[0].toUpperCase() + s.slice(1)).join('. ') ) Not perfect but hope it helps.
- Lio 18d agoYep, I had the same thought. As simple heuristic, text written in all lowercase is often just hot takes and so not worth taking the time to read. LC;DR :P
- senordevnyc 18d agoI have a similar heuristic for short HN comments.
- shantnutiwari 18d ago" the lack of sentence capitalization makes it unnecessarily difficult to read." If it had proper caps etc, people here would accuse it of written using LLMs. You just can't win...
- utopiah 18d agoTired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct. I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can comment on rapid prototyping, it's what I do. Rapid prototyping is NOT making a CMS quick. It's not about making a quick mockup of a UI. It's not about making yet another well known... anything. The entire POINT of prototyping is to make something NEVER done before. Typically that means you are reaching the frontier. You are making something with NO documentation to rely on. You are using tools, hardware or software, which do NOT have tons of StackOverflow errors. There is no dataset to crawl, there is no well structured Q&A database to train on. You have to poke and see if the thing actually works as expected, and it often does not. So sure, if you are using interns as a trick to underpay your staff, or if you are using prototyping as an excuse to build poor quality software fast, maybe it does help. If you are genuinely prototyping, it breaks fast and the supervision overhead makes it pretty pointless, especially since typically it's by actually implementing that you find out not just how the new setup works, but also its limits, and thus the actual needs of the project, not the one the stakeholder imagined would be. So not, not for rapid prototyping either. TL;DR: prototyping is a learning process, not a low fidelity output. PS: this comes up very often from NON prototypists that I wrote a short piece about it https://fabien.benetou.fr/Content/GoodPrototypesAre10LinesLong https://fabien.benetou.fr/Content/GoodPrototypesAre10LinesLo... so much so that it feels like a pattern "GenAI/LLMs is good for tasks X" while the author actually does not do task X except very superficially.
- Madmallard 18d agoAI is just good at what it's got the most elaborate training data on. And by "good" I mean, is statistically most likely to spit something out that makes some kind of sense. I don't know how well it is studied, but I suspect it is possible there is language-related complexity constraints to the effectiveness of the LLM algorithms. Like perhaps context-free grammar related problems with adequate training data can be more and more effectively solved, but maybe natural language related problems will not so much be effectively solved. Would be curious if there is active research here.
- lelanthran 18d ago> I don't know how well it is studied, but I suspect it is possible there is language-related complexity constraints to the effectiveness of the LLM algorithms. I believe that's obvious - humans don't think in words. Neither do animals. A machine that only thinks in words is obviously going to be deficient in some things, no matter how proficient it is in everything else.
- brokensegue 18d agoAren't the new models thinking in neuralese?
- lelanthran 17d ago> Aren't the new models thinking in neuralese? Where did you read that?
- brokensegue 17d agoMultiple places
- lelanthran 17d ago> Multiple places The people you are talking to have apparently never seen this research. Maybe you can help them out and provide links? It's faster to do that than to engage in some form of audience-persuasion, as well as more effective.
- kleiba2 18d agoGeez, why do you upper-case "LLM"?
- YeGoblynQueenne 18d ago>> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on, and even then with severe caveats. the frontier labs have developed a general recipe to teach models almost any specific task enjoying clearly defined levels of task performance; many tasks are covered in the training data; but even small perturbations within a covered class of task result in outright failure or reward hacking. Lots of people make this claim about "specific task[s] enjoying clearly defined levels of task performance" but they forget that generative AI is also extremely good at generating a) art and b) prose in literary style. None of those things has "clearly defined levels of task performance", in fact they are both the complete opposite of well-defined tasks. Who knows what counts for "good" art? [1] For me the right model for generative AI is "a million monkeys on typewriters" [2]. Holding any other model to heart will at some point fail to predict observations and cause you to be unpleasantly surprised. Not least because AI companies are actively engineering their systems to optimise for this model and they have a lot of people working on that engineering and shedloads of money to throw at it. Don't underestimate what a million monkeys on typewriters can do. They can do anything and everything, given enough time. Geneartive AI can also do anything and everything given enough resources. The only question is: how much is going to be "enough"? ____________________ [1] Yes yes, AI art tends to be slop. Not denying that. But part of the problem with slop is that it presents as technically very competent except that it lacks a certain je-ne-sais-quoi, which makes it good art; aesthetics. The point is that there is no clear measure of what makes technically competent art, any more than there is for aesthetics. And yet generative AI is very good at it. [2] There's even an article on wikipedia except it's about one monkey on one typewriter with infinite time. There's a proof too.
- alescalaios 18d ago[dead]
- tim333 18d ago>the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers I think that's incorrect from the investment point of view. They'd still be worth a lot if they can produce a drop-in replacement but it takes five or ten years as long as they dominate that. The danger from an investment point of view is they become AltaVista, replaced by some Google that does the job better.
- angarg12 18d agoOne the main theses is > the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on Unless frontier labs have surprisingly trained their models in the exact tasks my team works on, this is patently false. We are getting very good results on automation and I'm bullish we will be able to mostly remove humans in the loop for most of our infra tasks by the end of the year. I have no opinion on the other theses, but given that OP doesn't back up these claims in any way, I have my doubts about the conclusions of this article.
- killerstorm 18d ago"Language Models are Few-Shot Learners" - 2020, the GPT-3 paper. It have been demonstrated that in-context learning is a very powerful mechanism. There's no evidence that models of the size of GPT-6 are bad at in-context learning. In fact, ARC-AGI-3 score might indicate they are good at it. There's no evidence that a bespoke RL environment is required for each new skill - quite likely a good demonstration is sufficient.
- book_mike 18d agoOh my word, what a fossile.
- lissom 18d agoNavier-Stokes is a well defined problem, or "a hard technical problem". Most problems in the business world lack a good definition and tacit knowledge is required to solve them. As far as I've seen, AI lacks any kind of tacit knowledge, strategic thinking, etc what so ever. Take a customer service person, that as soon as AI agents replaced was hacked. Lots of tacit knowledge, that wasn't measured, or even probably in the job description, until AI agents had none and the gap was taken advantage of. Gap is probably a poor word here as it implies not a chasm, which could very well be the case. Your analysis was excellent but short on one front, AI has endurance on it's side. Looking at the N-S solution, OpenAI had 10,000+ instances that kept trying around the clock. Assembling a human team to do that would require a lot of effort. Though, that society collectively choose not to, perhaps tells how valuable it really is (i.e. it's now easy to launch a Manhattan Project level of effort). So maybe it can also be said that AI is also good at marshaling resources.
- highfrequency 18d ago> many tasks are covered in the training data; but even small perturbations within a covered class of task result in outright failure or reward hacking. Is this true? Could you cite an example of a simple prompt that GPT 6 / Fable 5.1 consistently bungle?
- tzone 18d agoIt all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine. It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun). On that note, I actually had an overall harness (for experimenting) that was essentially like this: "for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer". It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff. Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend infinite amount of money.
- plaidfuji 18d agoThis is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc. 3. those that can accept or already do by nature the costs of rigorous specification and validation: chip design, drug discovery, and other domains where failure on deployment is an existential concern. the first two classes are price sensitive and arguably don't need the jump in reasoning quality you see going from cheap to frontier models. most of these firms will be best served by open models running on cheap hardware, perhaps even locally at the site of use. for the first and third classes, the type of fuzzy combinatorial search that has produced headline results in mathematics and security research seems more sensitive to agentic swarm width than reasoning capacity … This is just so on point. And for the third class (which I would extend to things like materials research as well), specification and validation are already by FAR the larger costs, so automating search and simulation is really not a massive game changer for the broader business.
- reedlaw 18d agoIn my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law (https://en.wikipedia.org/wiki/Parkinson%27s_Law https://en.wikipedia.org/wiki/Parkinson%27s_Law) is still in effect. You could even expand "Work expands so as to fill the time available for its completion" with "time and tokens available".
- 0c3ca83 18d agoWhy do you belive that agents wouldn't be able to take over product management, and generate prompts for the "software engineer" agents?
- aaroninsf 18d agoEvery critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon. - Ximm's Law
- zeroonetwothree 18d agoYou could say the same of any critique about anything.
- kajumix 18d agoIs reward hacking really as unsolvable as they claim?
- adammarples 17d agoBrainlet swarms lol
- miguelacevedo 17d ago> the difference is that the data center full of geniuses is self-driving and limited only by how much compute it can consume while the brainlet swarms will be heavily bottlenecked by their human orchestrators. Here's another human bottleneck: "AI Has a Discovery Problem" (https://news.ycombinator.com/item?id=49621223 https://news.ycombinator.com/item?id=49621223)
- ethanwinters 17d agolet's just get used to them
- ZedZark 17d agoNobody talks about much compute they spent on their solution to Navier-Stokes. I did a back-of-the-envelope calculation, using the figured from their public release, and arrived at approximately $1M.
- emil-lp 17d agoEverybody's talking about that. Something like 300Bn tokens, something in the vicinity of USD 10M.
- marcus_cc 17d ago[flagged]