3 ms·
I tried doing the same thing with the Radeon R7 250 (an old GPU that was weak when it was introduced) and got essentially the same answer, so it doesn't seem li
by chc 4y ago
I tried doing the same thing with the Radeon R7 250 (an old GPU that was weak when it was introduced) and got essentially the same answer, so it doesn't seem like it's just regurgitating praise.
What impressed me just as much about these answers, though, was that it correctly identified that the comparison was silly. I don't believe it has seen that exact comparison before, so it seems like it has to have a granular enough categorization of tokens to know that people would say a comparison between a GPU and a computer is silly, and that those are a GPU and computer, and that one is old and the other is new.
- Jensson 4y agoGPU's typically compute faster than CPU's, so maybe that is the conversation it identifies it as? Meaning, when people ask about power difference between CPU and GPU they say they aren't directly comparable, but the GPU has much more processing power. That would mean it would identify an old GPU as faster than a modern CPU, even though it isn't, since it would slot that into the same conversation. The hard part with finding out the source for the logic is that it can map items to similar items in many ways, and there are trillions of conversations out there that it could map it to, so it can do quite a lot of heuristics that will work fairly well for naive questions even though it can't apply any direct logic. If it could link to similar conversations that it bases its reasoning on then that would probably greatly help understand where its logic would fail. So before it can do that we probably wont be able to eliminate these failure modes, because humans aren't random enough to generate the edge cases unless they understand what logic the model used. For example, comparing an old computer to a modern GPU might seem random to you, but to ChatGPT it just sees something like "This looks like a CPU vs GPU comparison, aha so I map it to: {CPU} isn't really comparable to {GPU}, but {GPU} has much more processing power". If you could see something like that then it would be obvious where it would fail, but without that it looks like magic.
- Jensson 4y ago> but to ChatGPT it just sees something like "This looks like a CPU vs GPU comparison, aha so I map it to: {CPU} isn't really comparable to {GPU}, but {GPU} has much more processing power" Jup, just tested with this: "what is more powerful Athlon PRO 3125GE or Geforce 6800". It said that they aren't comparable, but that the GeForce is probably faster: "In summary, while the GeForce 6800 is likely to be more powerful than the integrated graphics in the Athlon PRO 3125GE, the two cannot be directly compared as they have different functions and are designed for different purposes." Which is wrong, the modern integrated GPU is much faster, and they are directly comparable. So that is how it answered your question, and since I maanged to figure out what conversation I could find this failure mode where it gives the completely wrong answer. Once you understand that ChatGPT just slots what you ask for into some conversation based on fuzzy matching then it is really easy to understand what sort of answers it can give and can't give.
- notahacker 4y ago> What impressed me just as much about these answers, though, was that it correctly identified that the comparison was silly. That's been one of ChatGPT's most consistently impressive features, but it's clearly a function of a lot of training to encourage strings that reject comparisons when they're in different categories. Sometimes it does so based on simple heuristics like "FamousCharacter is fictional", sometimes it demonstrates a surprisingly nuanced grasp of historical periods or events so it will insist that x and y lived in different centuries or tell you that during a conflict p attempted to invade q, not the other way round. Even when it makes logical errors (no sooner has it told you that if Finland did not exist, there would be no Winter War, it suggests that the absence of Finland would not have affected anything about the outcome...) it's pretty good overall. But the trouble with lots of training to reject silly user inputs, is when it gets it wrong it can end up insisting it's been a good Bing dealing with a bad user...