4 ms·
> Just because a move has never been played, doesn't make it illegal, only untried. Well, I think this is nickelpro's point. Namely that the model "does not h
by crosen99 4y ago
> Just because a move has never been played, doesn't make it illegal, only untried.
Well, I think this is nickelpro's point. Namely that the model "does not have a systemic understanding of the rules, only a statistical inference of them."
However, I would still counter nickelpro's point with the idea that what we call systemic understanding is fallible - and most likely statistical in nature. Another response below gives the example of a human understanding the rules of chess but still making incorrect moves.
I might even say that any system truly capable of understanding is inherently fallible. A digital calculator or conventional database is impressive in its accuracy and consistency precisely because the rigor within its programming operates beyond the realm of understanding.
- nickelpro 4y agoI think the argument that human knowledge is itself statistical is far more interesting than the claim that neural nets have accurate models of the systems that feed their training data; but neither claim has much to do with the other.
- crosen99 4y agoI'm not sure why you say these claims are unrelated. The article claims to have found the existence of concepts embedded in the trained parameters of a model. (You called this having a "systemic understanding of the rules.") You are saying that for this to be so, the model would have to perform with 100% accuracy. The general form of this argument is that for a system of intelligence to be endowed with a conceptual understanding of a ruleset, it would have to implement that ruleset with 100% accuracy, right? Well, given that human's don't implement rulesets with 100% accuracy, we'd then need to concede that either 100% consistency is not a requirement for understanding or else human's do not operate with understanding.
- nickelpro 4y agoExactly, the claim that humans do not operate with perfect understanding is far more interesting (and IMHO, likely) than the idea that LLMs are spontaneously producing accurate models of the systems used to generate their data. And again, the claims are unrelated. There's no reason to believe that humans and LLMs operate under similar principles, and a ton of reasons to believe that they absolutely do not.
- crosen99 4y agoThis is what you said originally: "We only need a single counter example to show that Othello-GPT does not have a systemic understanding of the rules, only a statistical inference of them." I don't believe this statement to be true. It is clearly not true that a single counter example rules out systemic understanding of rules for ALL intelligent systems. Human intelligence is offered as the case in point. So, given that the statement is not true for all systems, by what logic do you conclude that the statement is true specifically for an LLM?
- nickelpro 4y ago> It is clearly not true that a single counter example rules out systemic understanding of rules for ALL intelligent systems. I didn't say anything about "intelligence", the OP is not about "intelligence". The OP claim is, "we find interesting evidence that simple sequence prediction can lead to the formation of a world model". A world model, an accurate one, is a very different thing than a statistical model. An inaccurate world model is indistinguishable from a statistical model. An accurate world model, definitionally, does not make mistakes. These models are trivial to produce, classical chess engines are accurate world models. > Human intelligence is offered as the case in point. I think you've hit on a more interesting point about human intelligence, that it is often not an accurate world model, which means it is inaccurate or more likely, statistical in nature. > So, given that the statement is not true for all systems, by what logic do you conclude that the statement is true specifically for an LLM? So what we're left with is that OthelloGPT (and humans, but again, what's going on with humans is irrelevant to the study of an LLM) is in one of two scenarios, it either: A) Has an inaccurate world model, or we might say, an accurate world model of a game that is not Othello B) Is still basically just a curve fitting algorithm, which there's all the evidence in the world to suggest it is It sits upon the authors to distinguish between these two choices, and they never offer any evidence either way. So I side with B.
- karmakaze 4y agoA model that tries these moves is more valuable than one that doesn't. It's easy to filter out by rules if we have them, but harder to get something to look beyond if it's self-restricting.
- phkahler 4y agoOthello is even more interesting here. There are no rules excluding a particular move other than "you must capture at least one of your opponents pieces". In that regard its much simpler than chess, and a valid move is more consequential than in Go. Trying to make an illegal move in Othello would be extremely weird. A reasonable strategy for a computer to play Othello to beat beginners is to literally evaluate the legality of all moves in a predefined order and make the first legal one. The predefined order encodes general positional advantage while the legality check enforces validity. In short, I strongly agree that making even 1 illegal move in Othello indicates a lack of understanding the rules.