5 ms·
I tried the example about the table not fitting into the car using GPT4. Not only it answered it correctly, but it actively (but politely) argued with me when I
by ay 4y ago
I tried the example about the table not fitting into the car using GPT4. Not only it answered it correctly, but it actively (but politely) argued with me when I tried to imply a different target for “it”:
https://mobile.twitter.com/ayourtch/status/1638872045281787904 https://mobile.twitter.com/ayourtch/status/16388720452817879...
(Edit: it did of course end up in a nonsense, so one can argue the point of not having the model of the world still stands; there is an “edge of bullshit”, but it seems to be moving further with each update)
At this point the arguments about about LLMs need to be tagged, which exact version they are talking about. (Probably most are talking about GPT3.5)
My (small) experience playing with both indicates there is a significant gap between the two.
- sparsevector 4y agoI also tried this out and the balloon example as well. User: Barbara, Mary, and Harry had red balloons. Ted had a blue balloon. Sam also had a red balloon. How many children have red balloons? ChatGPT (GPT-4): Barbara, Mary, Harry, and Sam each had a red balloon. That means there are 4 children with red balloons. User: What about Ted? ChatGPT (GPT-4): Ted had a blue balloon, not a red one. So, he is not counted among the children with red balloons. Therefore, the total number of children with red balloons remains 4. User: Why do you think I asked my original question? ChatGPT (GPT-4): It seems that you asked your original question to test my ability to understand and process information provided in a text. You wanted to see if I could correctly identify and count the number of children with red balloons based on the details given. User: Can you explain what mistakes a simpler model may have made on this question? ChatGPT (GPT-4): A simpler language model may have made the following mistakes when answering your original question: ... I like trying these sort of puzzles and "gotcha" prompts and I've found GPT-3.5 and lower versions are indeed not hard to confuse, but it's much harder to find simple examples that confuse GPT-4. Often times GPT-4 will even be able to explain the puzzle and why it's interesting. I think this is part of why GPT-4 feels qualitatively different. I understand (roughly) how transformers work, but GPT-4 does seem to have some emergent ability to reason, at least in the sense of making short chains of logical inferences, and this does seem to inform text completions not just on these sort of puzzles. That said I do think there's real danger of overconfidence and over reliance.
- imjonse 4y agoI thought the gotcha was that nowhere was it stated which of them are actually children.
- kzrdude 4y agoHow it just takes these implicit assumptions in stride makes me worry that in the future these chatbots will be fine with assumptions and me - human - has a too square brain to go on without definitions. :)
- usaar333 4y agoAgreed GPT4 is better, but there's still simple world modeling issues: > I have a chair in front of the TV and one in front of the computer. I put a watch under the chair in front of the TV. I swap the chairs. Where is the watch? > After swapping the chairs, the watch is now under the chair that is in front of the computer. Interestingly, GPT3 got it correct.
- kenjackson 4y agoI just tried it and it says it’s under the chair in front of the computer. I used the Bing model.
- letmevoteplease 4y agoI asked GPT 4 (Bing) your question, but added "(Please give a plausible explanation for why this question could be answered either way.)" > This question could be answered either way depending on how you interpret the phrase “under the chair”. If you mean under the chair as in attached to the bottom of the chair, then the watch would move with the chair when you swap them. In that case, the watch would be in front of the computer. However, if you mean under the chair as in on the floor below the chair, then the watch would stay on the same spot when you swap the chairs. In that case, the watch would be in front of the TV. So, it depends on what you mean by “under”.
- deleted 4y ago
- yorwba 4y agoThe answering style matches the usual presentation of a Winograd schema, which is a bit of a problem when you actually want to test its general reasoning ability, because it could be that someone already posted your go-to test question and the answer on the internet and the model was trained on that. Maybe we need to start obfuscating this stuff import base64;import hashlib;f=lambda s:bytes(k^v for k,v in zip(hashlib.shake_256(b'wOmGLuIVp70').digest(len(s)),s));print(f(base64.decodebytes(b'Iq6VQOBRifQwwwO6gluzzWEnGIICFKKFwM1oMWmBsTrIhMj5AseeNmNUwtEZkthcz8m8v8qKmVIx7nEjPOsqOUimKaTIJ8OKk2STdo/SRZGLAOsBSmgGaNTYEgT3KaayJWmGVf7K/UN06VyosEHfyFZlsS+PHDS6B3bN94qrzdnOA9f12FwWuaTPNJhLGcXFX7r5H8mtWyt9uWq6n5AItEcRXId04ssR8jfvNray2fwFIh5qPHTdQvZ9ogKLJ4Y+nAro7ecRSXXgskAj5EBmo2YobRkfE26er/Tj9DZHNx81N64ujWvN8jiS7aNcYs/oaEyN0oqnZvia9qocrv6CfBr+wGGSG1oxk5mbhAgkQhfuyR6c8MVKNFKp8HFo6SR7auju8vLnYjcObcII88vRbbua/jQmakiWwmS68Y1e1Gqmqg==')).decode()) to avoid burning test questions. (GPT-3.5 gave the answer I expected on the 4th try, didn't test with GPT-4.)
- thefreeman 4y agoThe biggest takeaway from the article for me was GPT 4 has 100 trillion parameters, 500x more then GPT 3. that type of exponential scaling is going to hit an upper bound really quickly. So when you say it “moves further with each update” you should realize that the improvements are coming from throwing massively more scale at the problem and not some underlying improvement in the technique, and also balance the amount of improvement with the scale itself. obviously gpt-4 is nowhere near 500x better than gpt-3. let’s say it’s 20% better (very generous imo). can they realistically 500x the model again? and if so, is that going to be worth an additional 4% gain to the original model quality? numbers are completely made up and math is probably wrong but i think i’m hopefully making my point, that diminishing returns will quickly become a blocker with this type of scaling.
- carbocation 4y agoI would hesitate to assume that is true. The 100 trillion param number has not been confirmed by anyone who could know if it’s true.
- cinntaile 4y agoWe shouldn't forget that at the same time effort is being spent to achieve the same results with a lot fewer parameters.
- famouswaffles 4y agoGpt-4 doesn't have 100 trillion parameters. It's not much larger than 3. It just has a lot more data according to Altman.
- qualudeheart 4y agoAltman himself has said it doesn’t have 100T parameters. It has more data in line with the Chinchilla laws.
- blueorange8 4y agoI don't believe they spent 7 months just "making gpt-4 safer" - I think they spent a very long time doing human reinforcement learning to make it better and now hope to speed that process up the next time using gpt-4 itself as the reinforcement.
- HarHarVeryFunny 4y agoI tried the car-table example too (via Bing), using the exact same wording, and got a totally different confused reply, so it must think the two parsings of the sentence are closely probable (and isn't deterred by the semantics). I tried it with Bard too, which also thought the table was too small to fit in the car, and doubled down on this when pressed by explaining how the table might be too narrow or deep or heavy(!) to fit in the car. I was a bit surprised to see GPT-4 get this wrong, even it it's only doing so some of the time (sampling temperature randomness?).
- HarHarVeryFunny 4y agoJust tried again with GPT-4 (exactly as before), and reply this time was simply "The table didn’t fit in the car because it was too big. Do you have any other questions?". Bing/GPT bolded the words "too big". But as I said before, the fact that it sometimes gets it wrong must mean it doesn't see one parsing as much to be favored over the other, which is surprising given how competent it generally is.