5 ms·
It's not even a little bit of a joke. Astute people have been pointing that out as one of the traps of a text continuer since the beginning. If you want to ant
by swatcoder 10mo ago
It's not even a little bit of a joke.
Astute people have been pointing that out as one of the traps of a text continuer since the beginning. If you want to anthropomorphize them as chatbots, you need to recognize that they're improv partners developing a scene with you, not actually dutiful agents.
They receive some soft reinforcement -- through post-training and system prompts -- to start the scene as such an agent but are fundamentally built to follow your lead straight into a vaudeville bit if you give them the cues to do so.
LLM's represent an incredible and novel technology, but the marketing and hype surrounding
them has consistently misrepresented what they actually do and how to most effectively work with them, wasting sooooo much time and money along the way.
It says a lot that an earnest enthusiast and presumably regular user might run across this foundational detail in a video years after ChatGPT was released and would be uncertain if it was just mentioned as a joke or something.
- stavros 10mo agoI keep hearing this non sequitur argument a lot. It's like saying "humans just pick the next work to string together into a sentence, they're not actually dutiful agents". The non sequitur is in assuming that somehow the mechanism of operation dictates the output, which isn't necessarily true. It's like saying "humans can't be thinking, their brains are just cells that transmit electric impulses". Maybe it's accidentally true that they can't think, but the premise doesn't necessarily logically lead to truth
- grey-area 10mo agoNo it’s not like saying that, because that is not at all what humans do when they think. This is self-evident when comparing human responses to problems be LLMs and you have been taken in by the marketing of ‘agents’ etc.
- stavros 10mo agoYou've misunderstood what I'm saying. Regardless of whether LLMs think or not, the sentence "LLMs don't think because they predict the next token" is logically as wrong as "fleas can't jump because they have short legs".
- stevenhuang 10mo ago> not at all what humans do when they think. Parent commentator should probably square with the fact we know little about our own cognition, and it's really an open question how is it we think. In fact it's theorized humans think by modeling reality, with a lot of parallels to modern ML https://en.wikipedia.org/wiki/Predictive_coding https://en.wikipedia.org/wiki/Predictive_coding
- stavros 10mo agoThat's the issue, we don't really know enough about how LLMs work to say, and we definitely don't know enough about how humans work.
- grey-area 10mo agoWe absolutely do, we know exactly how LLMs work. They generate plausible text from a corpus. They don't accurately reproduce data/text, don't think, they don't have a world view or a world model, and they sometimes generate plausible yet incorrect data.
- stavros 10mo agoHow do they generate the text? Because to me it sounds like "we know how humans work, they make sounds with their mouths, they don't think, have a model of the world..."
- Arkhaine_kupo 10mo ago> the sentence "LLMs don't think because they predict the next token" is logically as wrong it isn't, depending on the deifinition of "THINK". If you believe that thought is the process for where an agent with a world model, takes in input, analysies the circumstances and predicts an outcome and models their beaviour due to that prediction. Then the sentence of "LLMs dont think because they predict a token" is entirely correct. They cannot have a world model, they could in some way be said to receive a sensory input through the prompt. But they are neither analysing that prompt against its own subjectivity, nor predicting outcomes, coming up with a plan or changing its action/response/behaviour due to it. Any definition of "Think" that requieres agency or a world model (which as far as I know are all of them) would exclude an LLM by definition.
- Antibabelic 10mo ago> The non sequitur is in assuming that somehow the mechanism of operation dictates the output, which isn't necessarily true. Where does the output come from if not the mechanism?
- stavros 10mo agoSo you agree humans can't really think because it's all just electrical impulses?
- Antibabelic 10mo agoHuman "thought" is the way it is because "electrical impulses" (wildly inaccurate description of how the brain works, but I'll let it pass for the sake of the argument) implement it. They are its mechanism. LLMs are not implemented like a human brain, so if they do have anything similar to "thought", it's a qualitatively different thing, since the mechanism is different.
- socialcommenter 10mo agoMature sunflowers reliably point due east, needles on a compass point north. They implement different things using different mechanisms, yet are really the same.
- Antibabelic 10mo agoYou can get the same output from different mechanisms, like in your example. Another would be that it's equally possible to quickly do addition on a modern pocket calculator and an arithmometer, despite them fundamentally being different. However. 1. You can infer the output from the mechanism. (Because it is implemented by it). 2. You can't infer the mechanism from the output. (Because different mechanisms can easily produce the same output). My point here is 1, in response to the parent commenter's "the mechanism of operation dictates the output, which isn't necessarily true". The mechanism of operation (whether of LLMs or sunflowers) absolutely dictates their output, and we can make valid inferences about that output based on how we understand that mechanism operates.
- swatcoder 10mo agoThere's nothing said here that suggests they can't think. That's an entirely different discussion. My comment is specifically written so that you can take it for granted that they think. What's being discussed is that if you do so, you need to consider how they think, because this is indeed dictated by how they operate. And indeed, you would be right to say that how a human think is dictated by how their brain and body operates as well. Thinking, whatever it's taken to be, isn't some binary mode. It's a rich and faceted process that can present and unfold in many different ways. Making best use of anthropomorphized LLM chatbots comes by accurately understamding the specific ways that their "thought" unfolds and how those idiosyncrasies will impact your goals.
- samdoesnothing 10mo agoI never got the impression they were saying that the mechanism of operation dictates the output. It seemed more like they were making a direct observation about the output.
- Ferret7446 10mo agoThe thing is, LLMs are so good on the Turing test scale that people can't help but anthropomorphize them. I find it useful to think of them like really detailed adventure games like Zork where you have to find the right phrasing. "Pick up the thing", "grab the thing", "take the thing", etc.
- immibis 10mo agoAI Dungeon 2 was peak AI.
- internet_points 10mo ago> LLMs are so good on the Turing test scale that people can't help but anthropomorphize them. It's like Turing never noticed how people look at gnarly trees in the dark and think they're human.
- moffkalast 10mo ago> they're improv partners developing a scene with you That's probably one of the best ways to describe the process, it really is exactly that. Monkey see, monkey do.
- Terr_ 10mo ago> they're improv partners developing a scene with you, not actually dutiful agents. Not only that, but what you're actually "chatting to" is a fictional character in the theater document which the author LLM is improvising add-ons for. What you type is being secretly inserted as dialogue from a User character.
- jerf 10mo agoIt seems to me that even if AI technology were to freeze right now, one of the next moderately-sized advances in AI would come from better filtering of the input data. Remove the input data in which humanity teaches the AI to play games like this and the AI would be much less likely to play them. I very carefully say "much less likely" and not "impossible" because with how these work, they'll still pick up subtle signals for these things anyhow. But, frankly, what do we expect from simply shoving Reddit probably more-or-less wholesale into the models? Yes, it has a lot of good data, but it also has rather a lot of behavior I'd like to cut out of my AI. I hope someone out there is playing with using LLMs to vector-classify their input data, identifying things like the "passive-aggressive" portion of the resulting vector spaces, and trying to remove it from the input data entirely.
- undefeated 10mo agoI think part of the problem is that you need a model to classify the data, which needs to be trained on data that wasn't classified (or a dramatically smaller set of human-classified data), so it's effectively impossible to escape this sort of input bias. Tangentially, I'd be far from the first to point out that these LLMs are now polluting their own training data, which makes filtering simulatenously all the more important and impossible.
- mannanj 10mo agoSpoiler: the marketing around themselves has not misrepresented them without reason: its the most effective market and game theory design way to get training for your AIs as a company.