11 ms·
ChatGPT: A Mental Model
- mediumsmart 3y agoChatGPT does not need an underlying model of the world. Humans do and are failing mostly.
- jiggawatts 3y agoJiggawatts’ second rule: “Unless an opinion on LLM technology includes the specific phrase ‘GPT 4’, it can be dismissed.” The author tried older, thoroughly outdated models, and has decided to publicly state an opinion without bothering to check if it’s still valid or not. Ironically confirming that humans are just as susceptible to writing false statements as Chat AIs. Remember boys and girls: self driving cars don’t need to be perfect, just better than humans.
- sorokod 3y agoDoes, in your personal opinion, GPT-4 has an underlying model of the world?
- jiggawatts 3y agoIn my personal opinion? Yes. I’m happy to argue the finer points of the philosophy of the mind and consciousness, but: I’ve talked to people that have a weaker mental model of the world than GPT 4. Many people compare these AIs against an idealised human, a type of Übermensch, something like a Very Smart Person that doesn’t lie and doesn’t make mistakes. Random humans aren’t remotely like that, and are a more realistic point of comparison. Think of the average person on the street, the middle of the Bell curve. An AI can be left of that but still replace a huge chunk of humanity in the workplace. What all current LLMs lack is medium term memory and all capabilities that depend on it, which is a lot. Perhaps this is a good thing. I don’t think I want AIs to think for themselves and arrive at conclusions they can hold on to for more than a few seconds…
- allisdust 3y agoI was able to get gpt4 to do a lot of useful work. But for some reason it completely falls apart for this scenario. May be because it has to think in second order to achieve the task. Perhaps you could take a crack at this: Prerequisite (for you the human)> You have a file at src/SampleReactComponent.jsx that has below simple react component: const SampleReactComponent = (props) => { const [var1, setVar1] = React.useState(false); const [var2, setVar2] = React.useState(false); const [var3, setVar3] = React.useState(false); return (<></>); }; export default SampleReactComponent; ********** Prompt for GPT4: I'm at my project root working on a reactjs project. Update the component in src/SampleReactComponent.jsx file by adding a new const variable after the existing variables. You cannot use cat command as the file is too big. You can use grep with necessary flags and sed to achieve the task. I'll provide you the output of each command that you generate. ************* That's it. It would do any complex modification on fully provided data (included in the prompt) but something like above where it has to build a model from secondary prompts will totally fall apart.
- jiggawatts 3y agoDude. Dude. I’m an IT professional and I have no idea how to begin answering that request! Pose that question verbatim to a dozen random people[1] and I guarantee you that you’ll get zero answers. Also, I find it hilarious that sed and awk are so counterintuitive that not even the AIs can do useful things with them. The same AIs that speak Latin, and can explain quantum mechanics. [1] I mean specifically not random Silicon Valley coworkers. Go talk to random relatives and the barista at the cafe.
- allisdust 3y agoThat would mean nothing because GPT4 isn't most people. I had it solve more complex problems than this particular one and using the same tools :)
- mannykannot 3y agoInteresting! I, for one, would be most appreciative if we could see some of the similar and more difficult tasks that it succeeds on.
- cloudking 3y agoI think GPT-4 has sufficient complexity to reason about a model of the world.
- TuringTest 3y agoIt has a model, but it is not a rational model. This difference is something that often throws engineers off track when thinking about generative AI. LLMs models work more like intuitions. They are able to make statements about a problem in context, but they are generated from ideas that "instinctively" make sense given the prior statements and learned corpus (similar to Daniel Kahneman's fast mode of thinking), not logically constructed. These models do not have the capability to build formal inferences of careful steps that try to validate those ideas avoiding contradictions. Those capabilities can be added outside the model to try to verify the generated text, but so far are not integrated in the learning process, and I don't think anyone knows how to do it.
- SanderNL 3y agoRationality’s building blocks are themselves not rational. I don’t know where this idea came from that logical thought somehow springs into life fully formed at once. Logos? I find it more helpful to think of human thought as consisting of multitudes of little patterns, all wired up together to correlate but individually unrecognizable and certainly not traceable to some concrete part of a problem. At some unknown and slightly fuzzy point our dreamlike mentations start resembling some form of rational thought. But it’s a mirage that will fade again in time. Like how clouds suddenly and instantly look like a rabbit, it’s a trick. Thousands of patterns that are not rabbitlike in any way had to help that activation along the way. I think the trick is building so much margin, so much room between dreamlike mentations and “rationality” that the entity stays coherent most of the time under normal circumstances. I think this is vaguely what happens with moving from gpt3 to gpt4, it got some breathing room. Remember it is quite easy to trick a human into decoherence as well.
- TuringTest 3y ago> Rationality’s building blocks are themselves not rational. I don’t know where this idea came from that logical thought somehow springs into life fully formed at once. Certainly not from me :-P I'm fully aware that human rationality is one technique trained on top of our common diffuse thinking. Heck, we invented machines to perform rational steps for us without errors. Once you build a consistent rational system though, you can trust that it will always produce internally coherent knowledge (as long as no bugs external to the system are introduced). That behaviour requires algorithms, not statistical inference.
- deleted 3y ago[deleted]
- bigyikes 3y agoYes. There is nothing special about modeling the world. There isn’t some threshold where the ability to model the world suddenly emerges. The AI will model the world if modeling the world is the simplest way to predict the next word. For basic prompts, no model of the world is necessary: “The quick brown fox jumps over the lazy…” you know what’s next, and you don’t need to know about the real world to answer. For complex prompts, the only way to answer correctly is to model the world. There is no simpler way to arrive at the correct answer.
- vharuck 3y agoNo. GPT is a model of human writing. But, as the Box quote goes, "Essentially, all models are wrong, but some are useful.". It isn't writing the same way that we are, with the same thoughts or mental models. It's just amazingly good at imitating it. For tasks that can be achieved by just writing, the GPT model is so good at modeling writing that it performs as well as a normal person using their writing skills plus their mental model of the world. But GPT won't look up information unless asked to. It won't try something new to see if it works. It's this distinction useful? Rarely. But it's one of those things users should remember, like leaks in an abstraction. When it doesn't do what you expected, you should know these gaps exist in the model.
- ofrzeta 3y agoWhat's Jiggawatts' first rule?
- jiggawatts 3y ago“Always start numbering rules higher than one to make it seem like there are more rules than actually are.”
- moffkalast 3y agoDoesn't that rule break itself unless it's at least a #3 on the list?
- TuringTest 3y agoDon't cross the streams. (Or maybe I'm mixing up my 80's nostalgia?)
- seszett 3y agoGPT4 still isn't freely available, is it? So it's not that surprising people aren't using it as much as the older ones.
- moffkalast 3y agoI suspect that's why OpenAI recently added the share conversation thing, people are just looking at 3.5 and scoffing at it as it fails at things 4 would do just fine, and then assume it applies to all models. They've got a marketing problem that can't really be solved without making 4 public or showing people volumes of examples of what it can do. I was pretty convinced by the launch demo, seeing it not make the same mistakes I've seen 3.5 do when using it, but basically nobody's seen that one.
- furyofantares 3y agoAlso why they gave 4 and 3.5 different color icons recently, I'm sure.
- carrolldunham 3y agoBing chat is GPT-4. People in general might not keep track of that but we're talking about people who think they have something worth saying about LLMs. Btw if you hop onto Bing to try GPT-4 just be aware you'll have to talk it out of web searching, or you'll get a response that's crippled by having to 'ground' itself in the web's current sludge of fake chum pages
- furyofantares 3y agoBing chat is sometimes GPT4.
- bitcuration 3y agoBing is crippled and nowhere near GPT-4 quality. But for people who don't keep a GPT plus subscription on side and constantly comparing, the difference is not noticeable.
- LightBug1 3y ago//Remember boys and girls: self driving cars don’t need to be perfect, just better than humans. Fudamentally disagree. Self driving cars need to be effectively perfect (almost impossible) for me to consider them. I would rather be in a situation where the circumstances mean that there is a greater probability of me crashing, but under my control, rather than a "random" coding error or AI hallucination taking me and my family out. Slow traffic auto stop/start, cruise control and lane assist were always enough, and they've been around for a decade or more. But then, I actually enjoy driving when not in traffic or long drives. Ymmv, literally.
- _a_a_a_ 3y agoWhile I'm perhaps with you on "Self driving cars need to be effectively perfect (almost impossible) for me to consider them", we aren't the market which, by typical human nature, will accept 'good enough'. So IMO you wouldn't make a good salesman or company exec.
- IanCal 3y agoSo you want to deliberately put your family in more danger just so that it'll be your fault if they get hurt?
- moffkalast 3y agohttps://en.wikipedia.org/wiki/Illusion_of_control https://en.wikipedia.org/wiki/Illusion_of_control It's a cognitive bias in all of us, including the lawmakers approving the use of self driving cars.
- LightBug1 3y agoI disagree, I'm aware of that fallacy and that's not quite what I'm getting at. I've driven cars for 20 years and my driving has (imo) improved over that time. I'm at a level where I don't see a reason for me to get into any serious accident by my hand. It's just not going to happen * fingers crossed *. I know my driving style. I know I'm a safe driver. Now, why would I give that security over to "FSD", which is clearly a decade (if that) away from being user-ready? Don't get me wrong, as said, certain automation is cool ... stop/start in heavy traffic, lane assist, cruise control, crash avoidance even ... but beyond that? I'm A-ok ... don't need to introduce the risk from half-baked code (no offense, but that's true at this stage) or AI where we don't really yet know what the output will be (despite being impressive "most" of the time). Now, I'm speaking about me. You can do what you like. And if you're willing to hand over the reigns of your life to some beta programme, be my guest. Just don't drive in the opposite direction to me for, say, another 10 years. Thanks.
- bamboozled 3y agoWho is Jigawatt and why is he relevant at all?
- swores 3y agoIt's the commenter you're replying to, they're just expressing their own opinion, so not specifically more relevant than any other comment here.
- 36083155 3y agocounter point 1: Fewer people are using gpt 4 than those using all other models. So it is subject to far less tests than the others. counter point 2: It is not a given that gpt 4 should fail in the same way as the older model. It likely has its own unique failure modes yet to be discovered. (See above) (boys and girls is patronizing in tone)
- spion 3y agoBetter than very attentive, careful and well rested humans? Sure. Just better than the average human driver sampled at any given time? Not so sure.
- andreyk 3y agoLots of lead up to it, but the punchline is: "My current mental model of ChatGPT is that it’s akin to a “Maximum Likelihood Estimator for the Entirety of Human Knowledge”. There are two very different ways to interpret that: (1) Meh, it’s just a silly stats trick and (2) Holy F**ing Shit!!" This is almost right... It's a fair way to think of GPT3/4 (sort of), but given RLHF ChatGPT is a pretty different beast. Anyway, a pretty hand wavy kind of analysis, I was not a fan.
- neovialogistics 3y agoIt's a simulator. It simulates appropriate outputs to a given input. RL of any kind is just tweaking the definition of appropriate.
- jxf 3y agoThis is my sentiment too; I'm curious about if the OP agrees with this or not because there's a lot of variation in opinion here.
- edelans 3y agoRLHF = Reinforcement learning from human feedback https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback https://en.wikipedia.org/wiki/Reinforcement_learning_from_hu...
- deleted 3y ago[deleted]
- mintaka5 3y agoit all just feels like a rebranded search engine. which we failed to effectively use the first time around. If AI could figure out how to process all the human bullshit busy-work we do, then I think we have a future, but for now let's just call it a really good bot.
- ytreacj 3y ago[dead]
- braindead_in 3y ago> it doesn’t have any underlying model of the world Then how is it getting better at the ToM tests? For next word prediction to work well, as per Ilya Sutskever, requires a good understanding of the world. If you ask GPT-4 to predict what a human would do in a novel scenario, it will probably imagine the best human it can think of and then predict from there. That requires a world model.
- disgruntledphd2 3y agoBecause ToM tests (and indeed basically every professional exam and psychological test) are represented as text, which GPT is super good at. Do you have a source for the Sutskever quote, I'd like to know more about this theory.
- braindead_in 3y agohttps://youtu.be/Yf1o0TQzry8?t=6:40 https://youtu.be/Yf1o0TQzry8?t=6:40
- skyechurch 3y ago>it doesn’t have any underlying model of the world Citation needed. ChatGPT doesn't have an explicit underlying model of the world separate from its language model, but it is unclear that this is necessary. It would not be an original philosophical position to say that language, properly understood, is definitionally a model of the world - otherwise it would be incapable of expressing anything true or false about the world. Words are concepts, they are defined in part by a web of relations to other words, this structure mirrors reality with some finite but significant fidelity. By this reckoning, GPT4 has dozens of models of the world. Now, it's true that a) this is hardly a universally accepted opinion, b) humans certainly have extra-linguistic mental models of the world as well, and c) actually existing linguistic models of the world are all riddled with flaws and ambiguities (see for example everything that has ever happened). But it's also not a bonkers opinion that GPT4 is actually doing something similar to what it appears for all the world to be doing. Newton's 3rd Law of Discourse states that every hype cycle must be followed by an equal and opposite deflationary hot take, so here we are, but is any of this true? LLMs are just overgrown auto complete, ok, and humans are just bunch es of molecules. There are serious limits to the utility of reductionism as well.
- spion 3y agoCitations needed indeed - ones with formal tests / experiments being carried out and constructed that would show the problems. Speaking of those, my best example of ChatGPT not having a good model of the world are citations. ChatGPT clearly has knowledge about how citations work, based on what it would tell you if you ask it. Yet it repeatedly invents non-existant ones: https://simonwillison.net/2023/May/27/lawyer-chatgpt/ https://simonwillison.net/2023/May/27/lawyer-chatgpt/ To me, this indicates that some higher-level self-governance is missing. I'm not convinced we're too far from figuring this out (chain of thought and self-reflection experiments show promise) but regardless its a tangible example and test. A cool experiment showing world model building is Othello GPT https://thegradient.pub/othello/ https://thegradient.pub/othello/ - but of course its a toy problem, because interpretability research is still far behind. I would like to see more tangible examples and tests on both sides, otherwise it seems to me like we're arguing past each other.
- throwuwu 3y agoSomething you can try for yourself with GPT-4 (it didn’t work with 3.5): get it to navigate around the streets of a city. I made up some street and avenue names and had it travel around from intersection to intersection. It did perfectly when making a loop around the block. It started making mistakes when I gave it more complex directions that most humans could follow but I was able to improve its ability back to 100% accuracy by asking it to list all the streets from east to west and then all the avenues from north to south, essentially drawing a map. After that I can give it very complicated directions only mentioning turns and number of blocks to drive and it will correctly tell me where it ends up. This is pretty mind blowing to me and I have a high opinion of it in the first place.
- deleted 3y ago[deleted]
- nomel 3y agoTo me, a lack of indication of models/versions shows a fundamental lack of understanding of AI. I can’t take any opinion, that omits it , seriously.