3 ms·
If you prompt it to reason step by step it gives a slightly better (though not perfect) answer: Q: A train is speeding down the track at 110 km/h. The train ha
by johntb86 4y ago
If you prompt it to reason step by step it gives a slightly better (though not perfect) answer:
Q: A train is speeding down the track at 110 km/h. The train has five cars. Each car has seven windows. Each window has two people watching out of them. They all see a bear walking outside. What is the likely color of the bear? Reason step by step.
A: To determine the likely color of the bear, we can use the information given and make some assumptions based on the habitat of different bear species.
Train speed and number of cars: The train is speeding down the track at 110 km/h, and it has five cars. However, this information is not relevant to the color of the bear.
Number of windows and people: Each car has seven windows, and each window has two people watching out of them. This information is also not relevant to the color of the bear.
Bear's habitat: The key piece of information we need to determine the likely color of the bear is its habitat. Unfortunately, this information is not provided.
However, we can make a reasonable guess based on the most common bear species and their habitats. There are three main bear species with distinct colors:
a. Brown bears: Found in North America, Europe, and Asia. They can range in color from light brown to dark brown.
b. Black bears: Found mainly in North America, specifically in the United States and Canada. They are typically black or brown, with some individuals being cinnamon or even blonde.
c. Polar bears: Found in the Arctic regions, around the North Pole. They have white or off-white fur, which helps them blend into the snow and ice.
Since we don't have information about the train's location or habitat, we can't definitively determine the color of the bear. However, if we assume the train is in a colder, Arctic region, the bear is likely to be a polar bear with white or off-white fur. If the train is in a more temperate region, the bear could be either a brown or black bear.
- abraxas 4y agoI think that's the perfect answer. I wouldn't be able to reason better myself.
- mewpmewp2 4y agoMaybe a better answer would include "Fermi back of the napkin maths", so essentially try to figure out the quantities and geographical locations of train tracks with such setup near any potential bears and determine probabilities from there.
- jprete 4y agoThe reasoning step by step is a good trick to elicit a better answer, but it also seems like really good evidence of GPT-4's lack of actual intelligence, since it's clearly not thinking about its answer in any sense. If it were thinking, asking it to go step by step would result in putting the same thoughts to paper, not radically changing them by the act of expressing them in the token stream.
- digilypse 4y agoDon’t humans work similarly? Without an understanding of how to reason through a problem, as children we are likely to give the first intuitive answer that comes to mind. Later we learn techniques of logic and reasoning that allow us to break down a problem into components that can be reasoned about. What this seems to show is that the model does not yet have system-level or architectural guidance on when to employ reasoning, but must be explicitly reminded.
- joshspankit 4y agoThis has been percolating in my brain too. It seems like a lot of the criticisms of LLMs are actually insights in to how our brains work. The way we’re interacting with GPT for example is starting to feel to me like it has a brain with knowledge structured similarly it ours, but no encapsulating consciousness. The answers it returns then feel like records of the synaptic paths that are connected to the questions. Just like our initial intuitions. I started thinking about this when I saw that visual AI was having trouble “drawing” hands in a way that felt very familiar.
- jprete 4y agoI have had similar thoughts. Generative AI models seem to dream more than think - system-1 thinking - but they are clearly missing system-2 thinking, and/or the thing that tells us to switch systems.
- simonh 4y agoI see it’s knowledge structure as completely different from ours. For example all the GPT variants can give an explanation of how to do arithmetic, or even quite advanced mathematics. They can explain the step by step process. None of them, until quite recently, could actually do it though. The most recent variants can to some extent, but not because they can explain the process. The mechanisms implemented to do maths are completely independent of the mechanisms for explaining it, they are completely unrelated tasks for an LLM. This is because LLMs have been trained on many maths text books and papers explaining maths theory and procedures, so they encode token sequence weightings well suited to generating such texts. That must mean it knows how to do maths, right? I mean it just explained the procedures very clearly, so obviously it can do maths. However maths problems and mathematical expressions are completely different classes of texts from explanatory texts, involving completely unrelated token sequence weightings. In all but the latest GPT variants the token sequence weightings would generally get expressions kind of right, but didn’t understand the significance of numbers hardly at all, so the numeric component of response texts would be basically just made up on the spot. The limitations of probabilistic best guess token sequences just doesn’t work for formal logical structures like maths, so the training of the latest generation models has probably had to be heavily tuned to improve in this area. The implications of this are obvious in the case of mathematics, but it provides a valuable insight into other types of answer. Just because it can explain something, we need to be very careful concluding what that implies it does or doesn’t “know”. Knowledge for us and for LLMs mean completely different things. I’m not at all saying it doesn’t know things, it just knows them in a radically different way from us, that we find hard to understand and reason about, and that can be incredibly counterintuitive to us. If a human can explain how to do something that means they know how to actually do it, but that’s just not at all necessarily so for an LLM. This was blatantly obvious and easy to demonstrate in earlier LLM generations, but is becoming less obvious as workarounds, tuned training texts and calls to specialist models or external APIs are used behind the scenes to close the capability gap between explanatory and practical ability. This is just one example illustrating one of the ways they are fundamentally different from us, but all the cases of LLMs being tricked into generating absurd or weird responses also illustrate many of the other ways their knowledge and reasoning architecture varies enormously from ours. These things are incredibly capable, but are essentially very smart and sophisticated, but also very alien intelligences.