3 ms·
But isnt the point that it did roll a 17. And no one knows exactly how (in the case of LLMs)? Therefor any description of the conclusion should be thought of as
by nonethewiser 8d ago
But isnt the point that it did roll a 17. And no one knows exactly how (in the case of LLMs)? Therefor any description of the conclusion should be thought of as an anology. Decided, randomly accessed, etc.
- reichstein 8d agoTry "Emitted". That's what it did, with no analogy needed. (But, to be the devil's advocate: the fake can be said about the output of anyone participating here.)
- semi-extrinsic 8d agoIt's actually no different for dice than for LLMs. Explaining accurately the reason for the exact outcome of any given dice roll someone makes would be stupendously hard. It would require lots of instrumentation and math and be poorly transferrable to another surface, another player, etc. But even so people don't say that we don't understand how dice work. Saying that we don't understand how LLMs work is exactly like saying we don't understand how dice, or tires, or golf ball shots work. Or like the old myth that we don't understand how bumblebees fly.
- jacquesm 8d agoThat's precisely the point: you may be able to understand dice statistically and over the course of long rolls of dice you can extract some properties of the dice. But you won't ever understand any particular roll of the dice.
- deleted 8d ago[deleted]
- fc417fc802 8d agoBut importantly for dice we do understand the overarching principles that give rise to this. And dice don't output coherent sentences. Meanwhile in LLM land the analogous "roll of the dice" can result in a coherent response in natural language.
- skydhash 7d agoIf you use a loaded dice, you can be pretty confident about where it will lands. It may not be 100% accurate, but can be quite close to certain. Without training the weight are pure noises. After training, it leans towards coherent sentences and particular statements.
- fc417fc802 7d agoYes, and I believe my point still stands. We thoroughly understand the principle by which a loaded die can be intentionally biased despite not being able to predict the outcome of any given throw due to the system in question being a chaotic one. In contrast, we do not understand LLMs in the same way (nor biological brains). Claiming that anything of that nature is simply biased towards coherent output seems entirely reductive to me - the question is how such coherence arises in the first place. There is no meaning encoded or computation performed by the particular pathway a die travels through the chaotic landscape. Sure an argument can be made that it's "just" a next token predictor thus how is it really any different from a markov model? Yet the output is not even remotely the same.
- skydhash 7d ago> In contrast, we do not understand LLMs in the same way From my point of view, (not a ML researcher), it’s due to the magic of numbers. The same thing happens with computer vision and neural networks. There’s a bunch of magic weights that get created which has no meaning by themselves, but computing them does help with detecting objects. So if you take words, derives them into tokens, use the attention techniques to extract the “coherency” aspect, it’s no wonder you can replicate “coherency”. Add reinforcement learning to that to increase towards certain aspects like correct code syntax and you have heavily loaded the dice again. We have used maths to model chemistry, biology, and physics, as well as economics and sociologic phenomena. Then we use maths (more specifically logic and set theory) to usher in the age of information and computing. Now you want us to act surprised that maths, through ML, can model language. Maybe further down the line, we can have a simpler set of formulas for language coherency, but for now we have to make to with using the whole internet and a bazillion watts of power to guess the weights for the generic ML model.
- theptip 7d agoIt’s different in this way: The best way to model dice is the Physical Stance. You consider rules such as gravity, kinematics, etc. There is no “internal state”, “world model”, “knowledge”. If you prefer, in Friston’s terms, there is no Markov Blanket. The best way to model a human is the Intentional Stance[1]. You mostly need things like beliefs, knowledge, biases, etc to build this model. In Friston’s terms, there is a Markov Blanket, an inside vs outside. Without going into any irrelevant-but-interesting philosophical discussions about consciousness, I believe the intentional stance is most useful for modeling LLMs. Most of the success in predicting, debugging, optimizing these systems is in activities like understanding what they believe, what their intent was, what they observed, what they concluded from those observations. Also note that much simpler creatures benefit from the Intentional Stance; you will be more successful at modeling your dog if you think about what it “wants” rather than trying to run Physics on it. [1]: https://en.wikipedia.org/wiki/Intentional_stance https://en.wikipedia.org/wiki/Intentional_stance - the astute reader will note that I skipped the Design Stance. If we truly understood how NNs actually implement all their cognitive processes then we could perhaps apply this to them; if we actually crafted and designed every parameter of its mind. But we are talking about why dice are different.
- tantalor 7d agoGreat comment. I didn't know about this! Thanks The latest episode of On The Media also uses this framing. > On the Media: How Extinction Entered the AI Debate https://www.wnycstudios.org/podcasts/otm https://www.wnycstudios.org/podcasts/otm
- Cthulhu_ 8d agoIf nobody knows exactly how, then "at random" sounds about right and the results should be treated as such. That is, in this case, it should not be used to influence decisions that can start a war.
- krapp 8d agoPeople will just roll their eyes at you and say "the human mind is nothing but a dice roll too" and call you a slope-headed neanderthal before continuing apace.
- semiquaver 8d agoI agree wholeheartedly about your second sentence, but “we made this artifact and don’t know why the thing it does looks spookily like cognition” and “this artifact makes decisions at random” are obviously distinct categories and pretending otherwise is silly.
- watwut 8d agoWe know why it looks like cognition. Because OpenAI and Antropic put a lot of effort and training to humanize the output and make it sound like a person. Regardless of negative consequences it brings. They have that project of creating tech god which will save the unborn people thousands years in the future ... so people living now dont matter. That is why.
- semiquaver 7d ago“Putting a lot of effort and training” into a dog or an inanimate carbon rod would never result in something that can plausibly substitute for human mental labor and looks likely to eventually surpass us at many tasks, no matter how much you put in. So I don’t think “labs worked hard” is the same thing is “we know scientifically how these things work in any real level of detail”. The ability to build a thing, even if building it is hard, is not the same thing as understanding of what the thing is or how it works, not even a little bit.
- theptip 7d ago