5 ms·
> I feel like LLMs are a fairly boring technology. They are stochastic black boxes. The training is essentially run-of-the-mill statistical inference. There are
by Chabsff 11mo ago
> I feel like LLMs are a fairly boring technology. They are stochastic black boxes. The training is essentially run-of-the-mill statistical inference. There are some more recent innovations on software/hardware-level, but these are not LLM-specific really.
This is pretty ironic, considering the subject matter of that blog post. It's a super-common misconception that's gained very wide popularity due to reactionary (and, imo, rather poor) popular science reporting.
The author parroting that with confidence in a post about Dunner-Krugering gives me a bit of a chuckle.
- miningape 11mo agoI also find it hard to get excited about black boxes - imo there's no real meat to the insights they give, only the shell of a "correct" answer
- yannyu 11mo agoWhat's the misconception? LLMs are probabilistic next-token prediction based on current context, right?
- Chabsff 11mo agoYeah, but that's their interface. That informs surprisingly little about their inner workings. ANNs are arbitrary function approximators. The training process uses statistical methods to identify a set of parameters that approximate the function as best as possible. That doesn't necessarily mean that the end result is equivalent to a very fancy multi-stage linear regression. It's a possible outcome of the process, but it's not the only possible outcome. Looking at a LLMs I/O structure and training process is not enough to conclude much of anything. And that's the misconception.
- yannyu 11mo ago> Yeah, but that's their interface. That informs surprisingly little about their inner workings. I'm not sure I follow. LLMs are probabilistic next-token prediction based on current context, that is a factual, foundational statement about the technology that runs all LLMs today. We can ascribe other things to that, such as reasoning or knowledge or agency, but that doesn't change how they work. Their fundamental architecture is well understood, even if we allow for the idea that maybe there are some emergent behaviors that we haven't described completely. > It's a possible outcome of the process, but it's not the only possible outcome. Again, you can ascribe these other things to it, but to say that these external descriptions of outputs call into question the architecture that runs these LLMs is a strange thing to say. > Looking at a LLMs I/O structure and training process is not enough to conclude much of anything. And that's the misconception. I don't see how that's a misconception. We evaluate all pretty much everything by inputs and outputs. And we use those to infer internal state. Because that's all we're capable of in the real world.
- kmijyiyxfbklao 11mo agoThen why not say "they are just computer programs"? I think the reason people don't say that is because they want to say "I already understand what they are, and I'm not impressed and it's nothing new". But what the comment you are replying to is saying is that the inner workings are the important innovative stuff.
- yannyu 11mo ago> Then why not say "they are just computer programs"? LLMs are probabilistic or non-deterministic computer programs, plenty of people say this. That is not much different than saying "LLMs are probabilistic next-token prediction based on current context". > I think the reason people don't say that is because they want to say "I already understand what they are, and I'm not impressed and it's nothing new". But what the comment you are replying to is saying is that the inner workings are the important innovative stuff. But we already know the inner workings. It's transformers, embeddings, and math at a scale that we couldn't do before 2015. We already had multi-layer perceptrons with backpropagation and recurrent neural networks and markov chains before this, but the hardware to do this kind of contextual next-token prediction simply didn't exist at those times. I understand that it feels like there's a lot going on with these chatbots, but half of the illusion of chatbots isn't even the LLM, it's the context management that is exceptionally mundane compared to the LLM itself. These things are combined with a carefully crafted UX to deliberately convey the impression that you're talking to a human. But in the end, it is just a program and it's just doing context management and token prediction that happens to align (most of the time) with human expectations because it was designed to do so. The two of you seem to be implying there's something spooky or mysterious happening with LLMs that goes beyond our comprehension of them, but I'm not seeing the components of your argument for this.
- ACCount37 11mo ago> But we already know the inner workings. Overconfident and wrong. No one understands how an LLM works. Some people just delude themselves into thinking that they do. Saying "I know how LLMs work because I read a paper about transformer architecture" is about as delusional as saying "I read a paper about transistors, and now I understand how Ryzen 9800X3D works". Maybe more so. It takes actual reverse engineering work to figure out how LLMs can do small bits and tiny slivers of what they do. And here you are - claiming that we actually already know everything there is to know about them.
- LeroyRaz 11mo agoWhat do you mean? what do you think statistical modelling is? I am very confused by your stance. The aim of the function approximation is to maximize the likelihood of the observed data (this is standard statistical modelling), using machine learning (e.g., stochastic gradient decent) on a class of universal function approximators is a standard approach to fitting such a model. What do you think statistical modelling involves?
- parineum 11mo agoI'm not sure what claim your disputing or making with this. What more are LLMs than statistical inference machines? I don't know that I'd assert that's all they are with confidence but all the configurations options I can play with during generation (Top K, Top P, Temperature, etc.) are all ways to _not_ select the most likely next token which leads me to believe that they are, in fact, just statistical inference machines.
- ACCount37 11mo agoWhat more are human brains than piles of wet meat? It's not an argument - it's a dismissal. It's boneheaded refusal to think on the matter in any depth, or consider any of the implications. The main reason to say "LLMs are just next token predictions" is to stop thinking about all the inconvenient things. Things like "how the fuck does training on piles of text make machines that can write new short stories" or "why is a big fat pile of matrix multiplications better at solving unseen math problems than I am".
- zahlman 11mo ago> What more are human brains than piles of wet meat? Calculation isn't what makes us special; that's down to things like consciousness, self-awareness and volition. > The main reason to say "LLMs are just next token predictions" is to stop thinking about all the inconvenient things. Things like... They do it by iteratively predicting the next token. Suppose the calculations to do a more detailed analysis were tractable. Why should we expect the result to be any more insightful? It would not make the computer conscious, self-aware or motivated. For the same reason that conventional programs do not.
- ACCount37 11mo agoDo you have, by chance, a set of benchmarks that could be administered to humans and LLMs both, and used to measure and compare the levels of "consciousness, self-awareness and volition" in them? Because if not, it's worthless philosophical drivel. If it can't be defined, let alone measured, then it might as well not exist. What is measurable and does exist: performance on specific tasks. And the pool of tasks where humans confidently outperform LLMs is both finite and ever diminishing. That doesn't bode well for human intelligence being unique or exceptional in any way.
- LeroyRaz 11mo agoHow is that a misconception? LLMs are just advanced statistical modelling (unsupervised machine learning) with small tweaks (e.g., some fine-tuning for human preference). At the core, they are just statistical modelling. The fact that statistical modelling can produce coherent thoughts is impressive (and basically vindicates materialism) but that doesn't change the fact it is all based on statistical modelling. ...? What is your view?