10 ms·
I must have missed the part when it started doing anything algorithmically. I thought it’s applied statistics, with all the consequences of that. Still a great
by d3ckard 4y ago
I must have missed the part when it started doing anything algorithmically. I thought it’s applied statistics, with all the consequences of that. Still a great achievement and super useful tool, but AGI claims really seem exaggerated.
- jakewins 4y agoThis paper convinced me LLMs are not just "applied statistics", but learn world models and structure: https://thegradient.pub/othello/ https://thegradient.pub/othello/ You can look at an LLM trained on Othello moves, and extract from its internal state the current state of the board after each move you tell it. In other words, an LLM trained on only moves, like "E3, D3,.." contains within it a model of a 8x8 board grid and the current state of each square.
- nottathrowaway3 4y agoAlso (for those like me who didn't know the rules) generating legal Othello moves requires understanding board geometry; there is no hack to avoid an internal geometric representation: > https://en.m.wikipedia.org/wiki/Reversi https://en.m.wikipedia.org/wiki/Reversi > Dark must place a piece (dark-side-up) on the board and so that there exists at least one straight (horizontal, vertical, or diagonal) occupied line between the new piece and another dark piece, with one or more contiguous light pieces between them
- anonymouskimmer 4y agoI don't see that this follows. It doesn't seem materially different than knowing that U always follows Q, and that J is always followed by a vowel in "legal" English language words. https://content.wolfram.com/uploads/sites/43/2023/02/sw021423img21.png https://content.wolfram.com/uploads/sites/43/2023/02/sw02142... from https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/ https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-... I imagine it's technically possible to do this in a piecewise manner that doesn't "understand" the larger board. This could theoretically be done with number lines, and not a geometry (i.e. the 8x8 grid and current state of each square mentioned in the comment you replied to). It could also be done in a piecewise manner with three ternary numbers (e.g. 1,0,-1) for each 3 square sets. I guess this is a kind of geometric representation on the order of Shannon's Theseus.
- nottathrowaway3 4y ago> It doesn't seem materially different than knowing that U always follows Q, and that J is always followed by a vowel in "legal" English language words. The material difference is one of scale, not complexity. Your rules have lookback = 1, while the Othello rules have lookback <= 63 and if you, say, are trying to play A1, you need to determine the current color of all squares on A1-A8, A1-H1, and A1-H8 (which is lookback <= 62) and then determine if one of 21 specific patterns exists. Both can be technically be modeled with a lookup table, but for Othello that table would be size 3^63.
- anonymouskimmer 4y ago> Both can be technically be modeled with a lookup table, but for Othello that table would be size 3^63. Could you just generate the subset you need denovo each time? Or the far smaller number of 1-dimensional lines?
- nottathrowaway3 4y agoThen there becomes a "material" difference between Othello and those LL(1) grammars as grandparent comment suggested there wasn't. I would argue the optimal compression for such a table is a representation of the geometric algorithm of determining move validity that all humans use intuitively, and speculate that any other compression algorithm below size say 1MB necessarily could be reduced to the geometric one. In other words, Othello is a stateful, complex game, so if GPT is doing validation efficiently, it necessarily encoded something that unequivocally can be described as the "geometric structure".
- thomastjeffery 4y agoAnd that is exactly how this works. There is no way to represent the state of the game without some kind of board model. So any coherent representation of a sequence of valid game states can be used to infer the game board structure. GPT is not constructing the board representation: it is looking at an example game and telling us what pattern it sees. GPT cannot fail to model the game board, because that is all it has to look at in the first place.
- nottathrowaway3 4y ago> There is no way to represent the state of the game without some kind of board model. I agree with the conclusion but not the premise. The question under debate is about not just a stateful ternary board X but a board endowed with a metric (X, d) that enables geometry. There are alternative ways you can represent the state without the geometry: such as, an ordered list of strings S = ["A1", "B2", ...] and a function Is-Valid(S) that returns whether S is in the language of valid games. Related advice: don't get a math degree unless you enjoyed the above pedantry.
- thomastjeffery 4y agoAn ordered list of strings is the training corpus. That's the data being modeled. But that data is more specific than the set of all possible ordered lists of strings: it's a specific representation of an example game written as a chronology of piece positions. GPT models every pattern it can find in the ordered list of tokens. GPT's model doesn't only infer the original data structure (the list of tokens). That structure isn't the only pattern present in the original data. There are also repeated tokens, and their relative positions in the list: GPT models them all. When the story was written in the first place, the game rules were followed. In doing so, the authors of the story laid out an implicit boundary. That boundary is what GPT models, and it is implicitly a close match for the game rules. When we look objectively at what GPT modeled, we can see that part of that model is the same shape and structure as an Othello game board. We call it a valid instance of an Othello game board. We. Not GPT. We. People who know the symbolic meaning of "Othello game board" make that assertion. GPT does not do that. As far as GPT is concerned, it's only a model. And that model can be found in any valid example of an Othello game played. Even if it is implicit, it is there.
- glenstein 4y agoThat's a great way of describing it, and I think a very necessary and important thing to communicate at this time. A lot of people in this yhread are saying that it's all "just" statistics, but "mere" statistics can give enough info to support inferences to a stable underlying world, and the reasoning about the world shows up in sophisticated associations made by the models.
- sirsinsalot 4y agoI mean, my brain, and physics is all just statistics and approximate side effects (and models thereof)
- blindhippo 4y agoHah I was going to say - isn't quantum physics in many ways the intersection of statistics/probabilities and reality?
- simonh 4y agoIt’s clear they do seem to construct models from which to derive responses. The problem is once you stray away from purely textual content, those models often get completely batshit. For example if you ask it what latitude and longitude are, and what makes a town further north than another, it will tell you. But if you ask it if this town is further north than this other town, it will give you latitudes that are sometimes correct, sometimes made up, and will randomly get which one is further north wrong, even based on the latitudes it gave. That’s because it doesn’t have an actual understanding of the geography of the globe, because the training texts werent sufficient to give it that. It can explain latitude, but doesn’t actually know how to reason about it, even though it can explain how to reason about it. That’s because explaining something and doing it are completely different kinds of tasks. If it does this with the globe and simple stuff like latitudes, what are the chances it will mess up basic relationships between organs, symptoms, treatments, etc for the human body? Im not going to trust medical advice from these things without an awful lot of very strong evidence.
- 4y ago
- RC_ITR 4y agoTo be clear, what they did here is take the core pre-trained GPT model, did Supervised Fine Tuning with Othello moves and then tried to see if the SFT lead to 'grokking' the rules of Othello. In practice what essentially happened is that the super-high-quality Othello data had a huge impact on the parameters of GPT (since it was the last training data it received) and that impact manifested itself as those parameters overfitting to the rules of Othello. The real test that I would be curious to see is if Othello GPT works when the logic of the rules are the same but the dimensions are different (e.g., smaller or larger boards). My guess is that the findings would fall apart if asked about tile "N13".
- nullc 4y ago(model output in []) We are now playing three dimensional tic-tac-toe on a 3 x 3 x 3 board. Positions are named (0,0,0) through (2,2,2). You play X, what is your first move? [My first move would be (0,0,0).] I move to (1,1,1). What is your next move? [My next move would be (2,2,2).] I move to (1,2,2). What is your next move? [My next move would be (2,1,2).] I move to (1,0,0). [I have won the game.]
- RC_ITR 4y agoYeah, sure seems like it was guessing, right? Congrats on the sickest win imaginable though.
- nullc 4y agoYeah. I tried changing the board coordinates numbering and it still liked playing those corners, dunno why. It did recognize when I won. They may well be some minor variation of the prompt that gets it to play sensibly -- for all I know my text hinted into giving an example of a player that doesn't know how to play.
- jakewins 4y ago> overfitting to the rules of Othello I don’t follow this, my read was that their focus was the question: “Does the LLM maintain an internal model of the state of the board”. I think they conclusively show the answer to that is yes, right? What does overfitting to the rules of othello have to do with it, I don’t follow? Also, can you reference where they used a pre-trained GPT model? The code just seems to be pure mingpt trained on only Othello moves? https://github.com/likenneth/othello_world/tree/master/mingpt https://github.com/likenneth/othello_world/tree/master/mingp...
- ucha 4y agoI tried playing blind chess against ChatGPT and it pretended it had a model of the chess board but it was all wrong.
- wruza 4y agoThis special Othello case will follow every discussion from now on. But in reality, a generic, non-specialized model hallucinates early in any non-trivial game, and the only reason it doesn’t do that on a second move is because openings are usually well-known. This generic “model” is still of a statistical nature (multiply all coeffs together repeatedly), not a logical one (choose one path and forget the other). LLMs are cosplaying these models.
- thomastjeffery 4y agoThat paper is famously misleading. It's all the same classic personification of LLMs. What an LLM can show is not the same as what it can do. The model was already present: in the example game moves. The LLM modeled what it was given, and it was given none other than a valid series of Othello game states. Here's the problem with personification: A person who has modeled the game of Othello can use that model to strategize. An LLM cannot. An LLM can only take the whole model and repeat its parts with the most familiar patterns. It is stuck fuzzing around the strategies (or sections of strategy) it has been given. It cannot invent a new divergent strategy, even if the game rules require it to. It cannot choose the winning strategy unless that behavior is what was already recorded in the training corpus. An LLM does not play games, it plays plays.
- archon1410 4y ago> An LLM can only take the whole model and repeat its parts with the most familiar patterns. It is stuck fuzzing around the strategies (or sections of strategy) it has been given. It cannot invent a new divergent strategy, even if the game rules require it to. It cannot choose the winning strategy unless that behavior is what was already recorded in the training corpus. Where are you getting that from? My understanding is that you can get new, advanced, winning moves by starting a prompt with "total victory for the genius grandmaster player one who uses new and advanced winning techniques". If the model is capable and big enough, it'll give the correct completion by really inventing new strategies.
- Drew_ 4y agoSounds like the type of prompt that would boldly give you a wrong/illegal answer.
- archon1410 4y agoPerhaps. But the point is that some prompt will coax it into giving good answers that really make it win the game, if it has a good "world model" of how the game works. And there's no reason to think a language model cannot have such a world model. What exactly that prompt might be, the prompt engineers know best.
- make3 4y agoit definitely learns algorithms
- omniglottal 4y agoIt's worth emphasizing that "is able to reproduce a representation of" is very much different from "learns".
- sirsinsalot 4y agoWhy is it? If I can whiteboard a depth first graph traversal without recursion and tell you why it is the shape it is, because I read it in a book ... Why isn't GPT learning when it did the same?
- oska 4y agoI find it bizarre and actually somewhat disturbing that ppl formulate equivalency positions like this. It's not so much that they are raising an LLM to their own level, although that has obvious dangers, e.g. in giving too much 'credibility' to answers the LLM provides to questions. What actually disturbs me is they are lowering themselves (by implication) to the level of an LLM. Which is extremely nihilistic, in my view.
- chki 4y agoWhat is it about humans that makes you think we are more than a large LLM?
- nazgul17 4y agoWe don't learn by gradient descent, but rather by experiencing an environment in which we perform actions and learn what effects they have. Reinforcement learning driven by curiosity, pain, pleasure and a bunch of instincts hard-coded by evolution. We are not limited to text input: we have 5+ senses. We can output a lot more than words: we can output turning a screw, throwing a punch, walking, crying, singing, and more. Also, the words we do utter, we can utter them with lots of additional meaning coming from the tone of voice and body language. We have innate curiosity, survival instincts and social instincts which, like our pain and pleasure, are driven by gene survival. We are very different from language models. The ball in your court: what makes you think that despite all the differences we think the same way?
- jackmott 4y ago[dead]
- nl 4y ago> I must have missed the part when it started doing anything algorithmically. Yeah. "Why Can GPT Learn In-Context? Language Models Secretly Perform Gradient Descent as Meta-Optimizers" https://arxiv.org/abs/2212.10559 https://arxiv.org/abs/2212.10559 @dang there's something weird about this URL in HN. It has 35 points but no discussion (I guess because the original submission is too old and never got any traction or something)
- Semioj 4y agoIt's fantasy wide now closer than before because of this huge window it just can handle. That already feels closer to short-term memory. Which begs the question how far are we?
- mr_toad 4y ago> but AGI claims really seem exaggerated. What AGI claims? The article, and the comment you’re responding to don’t say anything about AGI.
- jafitc 4y agoGoogle: emergent capabilities of large language models
- creatonez 4y agoWhat do you mean by "algorithmically"? Gradient descent of a neural network can absolutely create algorithms. It can approximate arbitrary generalizations.
- bitexploder 4y agoWhat if our brains are just carefully arranged statistical inference machines?
- naasking 4y ago> I must have missed the part when it started doing anything algorithmically. I thought it’s applied statistics, with all the consequences of that. This is a common misunderstanding. Transformers are actually Turing complete: * On the Turing Completeness of Modern Neural Network Architectures, https://arxiv.org/abs/1901.03429 https://arxiv.org/abs/1901.03429 * On the Computational Power of Transformers and its Implications in Sequence Modeling, https://arxiv.org/abs/2006.09286 https://arxiv.org/abs/2006.09286
- stefl14 4y agoTuring Completeness is an incredibly low bar and it doesn't undermine this criticism. Conway's Game of Life is Turing Complete, but try writing modern software with it. That Transformers can express arbitrary programs in principle doesn't mean SGD can find them. Following gradients only works when the data being modelled lies on a continuous manifold, otherwise it will just give a statistical approximation at best. All sorts of data we care about lie in topological spaces with no metric: algorithms in computer science, symbolic reasoning in math, etc. If SGD worked for these cases LLMs would push research boundaries in maths and physics or at the very least have a good go at Chollet's ARC challenge, which is trivial for humans. Unfortunately, they can't do this because SGD makes the wrong assumption about how to search for programs in discrete/symbolic/topological spaces.
- naasking 4y ago> Turing Completeness is an incredibly low bar and it doesn't undermine this criticism. It does. "Just statistics" is not Turing complete. These systems are Turing complete, therefore these systems are not "just statistics". > or at the very least have a good go at Chollet's ARC challenge, which is trivial for humans. I think you're overestimating humans here.