5 ms·
The idea that LLMs “only know how to predict the next word” is a common but imho wrong cop out. It has been shown that they are good at doing a lot of things th
by low_tech_love 2y ago
The idea that LLMs “only know how to predict the next word” is a common but imho wrong cop out. It has been shown that they are good at doing a lot of things that might not have been immediately obvious, regardless of how their internal process works. 50 years ago, somebody might say “programming is only a way to automate simple repetitive tasks” and that would be obviously wrong.
The real problem in cases like this and other applications, as you and many others have mentioned, is that LLMs are basically correlation machines. They can find very complex, immensely-multivariate correlations in large data sets, and reproduce these correlations very well. But they cannot reason (so far) beyond these correlations in other to find deeper, less obvious causal relationships. They’re simply not trained to do that, yet. But it’ll come…!
- austin-cheney 2y ago> 50 years ago, somebody might say “programming is only a way to automate simple repetitive tasks” and that would be obviously wrong. That is actually extremely correct. The only purpose for software is automation, which is the elimination of labor. Getting that wrong directly influences your quality of product more than any other downstream factor.
- low_tech_love 2y agoIt might be theoretically correct in the same way that it is correct to say that “a building is just a bunch of bricks on top of each other”. But there are thousands of different buildings with different reasons to exist which offer drastically different services and serve different purposes. A video game like Elden Ring is built from the same “automation” pieces as the LS command in my terminal, but it would very disingenuous to say they’re basically the same thing.
- austin-cheney 2y agoThat is an incorrect comparison. A building, all buildings, are dwellings, but they are no more or less the sum of their parts than anything else. It’s not about the construction materials. It’s about the utility.
- teleforce 2y agoNot OP but I think it is correct comparison. Dijkstra's quote succinctly summarized the arguments: "The question of whether a computer can think is no more interesting than the question of whether a submarine can swim."
- CoastalCoder 2y agoWith all due respect to Dijkstra, whether or not something is interesting is subjective. I wonder if his point was that the question didn't need to be answered for whatever they were discussing at the time?
- dTal 2y agoHis point is that question ultimately boils down to a semantic argument over the word "think/swim", which may be interesting to a linguist but is not philosophically meaty in the way the question implies.
- low_tech_love 2y ago> It’s not about the construction materials. It’s about the utility. I think I missed your point. Wasn’t that exactly what I said?
- herval 2y ago> The only purpose for software is automation, which is the elimination of labor. Getting that wrong directly influences your quality of product more than any other downstream factor. I take it you never played a game in your life?
- hinkley 2y agoDevil’s advocate: Fighting games are just better versions of Rockem Sockem Robots. And 4X games could be fancier versions of Settlers of Catan.
- herval 2y agoWhat “labor” does a “fancier version of Catan” eliminate?
- austin-cheney 2y agoDo you mean baseball, football, tag, board games, cross words, Sudoku, Dungeons and Dragons, or something else? There are many games that are not electronic. So, what separates those athletic and paper games from electronic games? Automation.
- herval 2y agoNo, I obviously don’t mean those, buddy
- outofpaper 2y agoVideo games automate whole chunks of play, be that imagination with graphics and sound effects, other players with virtual enemies and allies, let alone all the automation that goes into multi-player games.
- cookie_monsta 2y ago> It has been shown that they are good at doing a lot of things that might not have been immediately obvious Could you point me to some of the places/articles where this is being shown? I'm definitely amongst those who have bought in to the common cop out you are rebutting here
- geoduck14 2y agoNot OP, but I have been poking around with LLMs for a year and I'd like to add to the conversation. In my experience, LLMs are word predictors, and the impacts of that fact are not immediately obvious. LLMs are capable of "explaining" what code does. What it is doing under the hood is pattern matching: I've seen code that looks like X, with an explanation that looks like Y LLMs are capable of formatting text. It has seen English written like X, that is reformatted to look like Y One resounding fact my team has found over and over is that "the things we think are hard for LLMs aren't necessarily hard; the things we think are easy aren't necessarily easy"
- dTal 2y agoWorth noting that pattern matching and term rewriting are a sufficiently general combination that it forms the basis for Mathematica.
- psb217 2y agoTrouble sneaks in when the pattern matching is only correct most of the time. Eg, if some code for regex-based search missed anywhere from 0.1% to 10% of matches, with the miss rate depending on the regex and no obvious way to know which regexes have worse miss rates, the utility of your regex-based search would be limited. LLMs are like this, but their generality makes them useful in spite of this limitation.
- 60654 2y agoI would just add that Mathematica-style term rewriting (e.g. analytical integration, or equation solving) is done with _semantics-preserving_ symbolic solvers, which are hand-made and human-reviewed to guarantee correctness. LLM style pattern matching and rewriting does not preserve semantics, except accidentally due to an overwhelming amount of examples.
- akomtu 2y agoWe rarely use reasoning too. When a lightning strikes, a thunder follows. When the sun shines, trees grow. These are correlations we've learned. Now can you derive them with proper reasoning? In fact, if we dig deep enough, at some point we'll face some principle or axiom that just postulates an observed correlation as a law. What we call reasoning can be the art of finding a chain of small correlations to connect ends of a big correlation. Some sort of quantum-powered DFS algorithm. However a reasoning machine is just a machine. Someone needs to tell it what to reason about.