11 ms·
From word models to world models
- cs702 3y agoAfter a quick/superficial read, my understanding is that the authors: (a) induce an LLM to take natural language inputs and generate statements in a probabilistic programming language that formally models concepts, objects, actions, etc. in a symbolic world model, drawing from a large body of research on symbolic AI that goes back to pre-deep-learning days; and (b) perform inference using the generated formal statements, i.e., compute probability distributions over the space of possible world states that are consistent with and conditioned on the natural-language input to the LLM. If this approach works at a larger scale, it represents a possible solution for grounding LLMs so they stop making stuff up -- an important unsolved problem. The public repo is at https://github.com/gabegrand/world-models https://github.com/gabegrand/world-models but the code necessary for replicating results has not been published yet. The volume of interesting new research being done on LLMs continues to amaze me. We sure live in interesting times! --- PS. If any of the authors are around, please feel free to point out any errors in my understanding.
- skepticATX 3y agoI have not yet read the paper, but based on this description it seems like it provides grounding in the context of the training data, which is kind of the rub with current LLMs to begin with, right? We don't have a set of high quality training data that is completely unbiased and factual.
- cs702 3y agoI'd describe it as grounding the model with a formally specified symbolic world model.
- agnosticmantis 3y ago> … which is kind of the rub with current LLMs to begin with, right? No, the bigger problem with current LLMs is that even with high quality factual training data, they often generate seemingly plausible nonsense (e.g. cite nonexistent websites/papers as their sources.) This is by design imo; they’re trained to generate ‘likely’ text, and they do that extremely well. There’s no guarantee for faithful retrieval from a corpus.
- novaRom 3y agoImportant addition to your partially right statement: "they’re trained to generate ‘likely’ text" is they are trained to produce most probable next word so that the current context look as "similar" to training data as possible. Where "similar" is not "equal".
- andsoitis 3y agoHumans’ experience and understanding of the world around them isn’t limited to a symbolic representation. It remains to be seen whether you can truly be an effective intelligence with understanding of the world if all you have are symbols that you have to manipulate.
- mjburgess 3y agoIt's a surprise to see a paper actually try to solve the problem of modelling thought via language. Nevertheless, it begins with far too many hedges: > By scaling to even larger datasets and neural networks, LLMs appeared to learn not only the structure of language, but capacities for some kinds of thinking There's two hypotheses for how LLMs generate apparently "thought-expressing" outputs: Hyp1 -- it's sampling from similar text which is distributed so-as-to-express a thought by some agent; Hyp2 -- it has the capacity to form that thought. It is absolutely trivial to show Hyp2 is false: > Current LLMs can produce impressive results on a set of linguistic inputs and then fail completely on others that make trivial alterations to the same underlying domain. Indeed: because there're no relevant prior cases to sample from in that case. > These issues make it difficult to evaluate whether LLMs have acquired cognitive capacities such as social reasoning and theory of mind It doesnt. It's trivial: the disproof lies one sentence above. Its just that many don't like the answer. Such capacities survive trivial permutations -- LLMs do not. So Hypothesis-2 is clearly false.
- rytill 3y agoI don't think you really disproved anything. You're just saying another hypothesis. Often, LLMs produce impressive results on domains that aren't in the training set.
- sgt101 3y ago>LLMs produce impressive results on domains that aren't in the training set. How do we know? Who knows what they're trained on?
- famouswaffles 3y ago>It is absolutely trivial to show Hyp2 is false No it's not > Current LLMs can produce impressive results on a set of linguistic inputs and then fail completely on others that make trivial alterations to the same underlying domain. >Indeed: because there're no relevant prior cases to sample from in that case. That's not what that tells us. Humans have weird failure modes that look absurd outside the context of evolutionary biology (some still look absurd) and that don't speak to any lack or presence of intelligence or complex thought. Not sure why it's so hard to grasp that LLMs are bound to have odd failure modes regardless of the above. and trivial here is relative. In my experience, "trivial" often turns out to be trivial in the way a person may not pay close attention to and be similarly tricked. For instance, GPT-4 might solve a classic puzzle correctly then fail the same puzzle subtlety changed. I've found more often than not, simply changing names of variables in the puzzle to something completely different can get it to solve the changed puzzle. It takes memory shortcuts but can be pulled out of that. LLMs have failure modes that look like human failure modes too.
- antiquark 3y agoI doubt that word models can lead to world models. To quote Yann LeCun: "The vast majority of our knowledge, skills, and thoughts are not verbalizable. That's one reason machines will never acquire common sense solely by reading text." https://twitter.com/ylecun/status/1368235803147649028 https://twitter.com/ylecun/status/1368235803147649028
- FrustratedMonky 3y ago>"solely by reading text". Of course, that does leave the door Open, that when these models are put in a physical real body, a robot, and have to interact with the world, then maybe they can gain that "common sense". This doesn't mean a silicon based AI can't become conscious of skills that are hard to verbalize. Just that they don't yet have all the same inputs that we have. And when they do, and they have internal thoughts, they will have the same difficulty verbalizing them that we do.
- moffkalast 3y agoThat just seems like an unfounded hot take. Of course we can explain most of our knowledge, skills, and thoughts in words, that's how we don't lose everything when the next generation comes around lol. It's the core reason we're different from animals. Now sure you can't describe qualia, but that's basically a subjective artefact of how we sense the world and (to add another unfounded hot take) likely not critical to have an understanding of it on a physical level.
- pphysch 3y ago> Of course we can explain most of our knowledge, skills, and thoughts in words, that's how we don't lose everything when the next generation comes around lol. I would wager if you put a newborn human to be raised in the absence of any physical human contact, but somehow taught them to read/write, and gave them access to a universal corpus (text only, no audio/video), or heck, even internet access with `curl`, and lastly dropped them into the "real world" at age 25, they would be utterly incapable of performing, say, a basic service job at a restaurant. Words help us symbolize and reason about our sense experiences, but they are not a substitute for them.
- mjburgess 3y agoThe level of understanding of the problem that this paper expresses is extraordianry in my reading of this field --- it's a genuinely amazing synthesis. > How could the common-sense background knowledge needed for dynamic world model synthesis be represented, even in principle? Modern game engines may provide important clues. This has often been my starting point in modelling the difference between a model-of-pixels vs. a world model. Any given video game session can be "replayed" by a model of its pixels: but you cannot play the game with such a model. It does not represent the causal laws of the game. Even if you had all possible games you could not resolve between player-caused and world-caused frames. > A key question is how to model this capability. How do minds craft bespoke world models on the fly, drawing in just enough of our knowledge about the world to answer the questions of interest? This requires a body: the relevant information missing is causal, and the body resolves P(A|B) and P(A|B->A) by making bodily actions interpreted as necessarily causal. In the case of video games, since we hold the controller, we resolve P(EnemyDead|EnemyHit) vs. P(EnemyDead| (ButtonPress ->) EnemyHit -> EnemyDead)
- mercurialsolo 3y agoHumans come in all shapes and forms of sensory as well as cognitive abilities. Our true ability to be human comes from objectives (derived from biological and socially bound complex systems) that drive us, feedback loops (ability to morph / affect the goals) and continuous sensory capabilities. Reasoning is just prediction with memory towards an objective. Once large models have these perpetual operating sensory loops with objective functions, the ability to distinguish model powered intelligence and human like intelligence tends to drop.
- gibsonf1 3y agoUnfortunately, this effort fully misses the boat. Human cognition is about concepts, not language, and that's where one must start to understand it. Language simply serializes our conceptual thinking in multiple language formats, the key is what's being serialized and how that actually works in conceptual awareness.
- buzzy_hacker 3y agoMaybe they can’t be so fully separated. https://en.m.wikipedia.org/wiki/Linguistic_relativity https://en.m.wikipedia.org/wiki/Linguistic_relativity
- gibsonf1 3y agoI think the key point is that serialized words symbolize concepts and other logic such that if you can't retrieve that concept into your awareness, you will not understand the word. Learning and forming the concepts comes prior to attaching common word symbols to them based on the region you live in. So if you start with words, you never get anywhere, hence the complete lack of any intelligence in the LLM approach.
- taliesinb 3y agoExactly. Thought is prior to language, and much confusion happens when you conflate the them. In particular the surface syntax of language tells you next to nothing about the "syntax" of thought, which is hypergraphical, not tree-structured.
- canjobear 3y agoRead more carefully. Their "language of thought" is not a natural language, it's a variant of lambda calculus with probabilistic semantics.
- gibsonf1 3y agoRight, derived from word pattern statistics. The CYC project tried first order predicate calculus with complete failure. This is not how we think or how conceptual awareness works. The key give away is what they don't talk about, Concepts.
- antisthenes 3y agoWorld modeling is impossible without sensory input. You need constant modeling of touch/smell/vision/temperature, etc. These senses give us an actual understanding of the physical world and drive our behavior in a way that pure language will never be able to.
- stevenhuang 3y agoA facsimile of sufficient equivalence to the world models we derive from our 5 senses may be approached through derivation of descriptive language only. "sufficient equivalence" is important because sure it may not _really_ know the color of red or the qualia of being, but if for all intents and purposes the LLM's internal model provides predictive power and answers correctly as if it does have a world model, then what is the difference?
- esafak 3y agoThat's not how physics works. We understand the world by interacting with it. How do you know your internal model is right until it is tested in reality?
- stevenhuang 3y agoSeems you're unaware the amount of world knowledge that already exist in written form. Think of all the top journals, textbooks, etc. People have understood the world by interacting with it, detailed their hypothesis, conducted experiments, recalled their learning and written down conclusions. It's not at all obvious to say a useful world model cannot be derived strictly from all this written information.
- thewataccount 3y agoYeah but we can serialize the world to numbers and already have. I asked GPT3.5turbo "Pretend you are a character called Samatha and you're in your house. You go up to the thermostat and select a comfortable temperature and explained your reasoning" > Next, I take into account my personal preferences and comfort levels. Everyone has their own ideal temperature range, and it's essential to find the sweet spot that makes me feel most comfortable. For me, it's usually between 22 to 24 degrees Celsius (72 to 75 degrees Fahrenheit). This range allows me to feel neither too cold nor too warm, striking the perfect balance. It also goes on about how the humidity could effect the desired temperature, etc. It doesn't need the ability to feel temperature (which could also be a single floating number using kelvin), but it can already describe a "comfortable temperature" and what factors would effect it. Side note: It doesn't "know" anything, it can only make a "best guess" which is now fairly reliable enough to be useful. It doesn't need the ability to test things to learn, we did it already for it, and it's using that to predict the results. You could make a recursive system to allow it to test data if you'd like though.
- ilaksh 3y agoSo they are using GPT-4 to write Lisp? Or some probabilistic language that looks like Lisp. They keep saying LLMs but only GPT-4 can do it at that level. Although actually some of the examples were pretty basic so I guess it really depends on the level of complexity. I feel like this could be really useful in cases where you want some kind of auditable and machine interpretable rationale for doing something. Such as self driving cars or military applications. Or maybe some robots. It could make it feasible to add a layer of hard rules in a way.
- dimatura 3y agoThis is really interesting. The title is referencing the "Language of Thought" hypothesis from early cognitive psychology, that posited thought consisted of symbol manipulation akin to computer programs. The same idea was behind was also what is often referred to GOFAI. But the idea has largely fallen out of fashion in both psychology and AI. There's a twist here in the "probabilistic" part, and of course the surprising success of LLMs makes this a more compelling idea than it would've been only a couple of years ago. And there's also an acknowledgement of the need for some kind of sensorimotor grounding as well. Pretty cool!
- wilonth 3y agoWas excited for a moment, thought it was related to this https://worldmodels.github.io/ https://worldmodels.github.io/. World models are meant to be for simulating environments. If this was something like testing if a game agent with llm can form thoughts as it play through some game it would be very interesting. Maybe someone on HN can do this?
- sgt101 3y agoI : hhmmppp a paper from Tenenbaum's group, let's read. Paper : Hi! I am 94 pages long. I : omg...