3 ms·
"GPT-3 randomizes answers in order to avoid repetition that would give the appearance of canned script. That’s a reasonable strategy for fake social conversatio
by FrenchDevRemote 5y ago
"GPT-3 randomizes answers in order to avoid repetition that would give the appearance of canned script. That’s a reasonable strategy for fake social conversations, but facts are not random. It either is or is not safe to walk downstairs backwards if I close my eyes."
I stopped there, completely inaccurate article, there is parameters like temperature that you need to take care of. You can set it up to give extremely similar answers all the time.
They have humans mostly to remove offensive or dangerous content. Humans are not what's "making it work"
- nullc 5y agoCame here to make essentially the same comment as you? Why should we care about the opinions on GPT3 from people who aren't interested (or able?) to understand even the most simple ideas about how it works. These sort of models take the context and output so far and predict a probability distribution over the next character. The next character is then sampled from the probability. In written text there is essentially never a single correct next character-- it's always some probability. This has nothing to do with trying to fake the inconsistent answers humans give. Always choosing the most likely character drives GPT3 into local minima that give fairly broken/nonsense results.
- RandomLensman 5y agoUltimately, you likely need to convince people who don't care about how it works/who are only interested in that it does or doesn't work. Right now, time might not have come for use cases that need such buy-in, but if and when it happens, need to be prepared for it.
- disgruntledphd2 5y agoCan I see your contributions to statistical theory and data analysis please?
- nullc 5y agoWhat bearing does my publication history have on the hot-take by someone commenting outside of their (sub)field that clearly don't understand the basic operation of the mechanism they're commenting on? The author of that text is simply mistaken about the basic operation of the system, thinking that the sampling is added to imitate human behavior. It isn't. You can see the same structure in things as diverse as wavenet-- a feedforward cnn rather than a transformer-- and for the same reason, if you feed back only the top result you rapidly fall into a local minima of the network that gives garbage output. Another more statistical way of looking at it is that the training process produces (or, rather, approaches) the target distribution of outputs even without any lookahead, but it can't do that if it selects the most likely symbol every time because in the real distribution (if we could evaluate it) there are necessarily some outputs which are likely but have prefixes which are unlikely relative to other prefixes of the same length. If you never sample unlikely prefixes you can't reach likely longer statements that start with them. To give a silly example: "Colorless green ideas sleep furiously" is a likely English string relative to its length which GPT3 should have no problem producing (and, in fact, it produces it fine for me). But the prefix "Colorless green" without the rest is just nonsense-- extremely unlikely compared to many other strings of that length. [Not the best example, however, because the prevalence of that specific nonsense statement is so great that GPT3 is actually prone to complete it as the most likely continuation even after just the word colorless at the beginning of a quote. :P but I think it still captures the idea.] If you derandomized GPT* by using a fixed random seed for a CSPRNG to make the sampling decisions every time, the results would be just as good as the current results and it would give a consistent answer every time. For applications other than data compression doing so would be no gain, and would take away the useful feature of being able to re-try for a different answer when you do have some external way of rejecting inferior results. In theory GPT without sampling could still give good results if it used a search to look ahead, but it appears that even extraordinary amounts of computation for look-ahead still is a long way from reaching the correct distribution, presumably because the exponential fan out is so fast that even 'huge' amounts of lookahead are still only testing a tiny fraction of the space.
- gwern 5y ago> In theory GPT without sampling could still give good results if it used a search to look ahead, but it appears that even extraordinary amounts of computation for look-ahead still is a long way from reaching the correct distribution, presumably because the exponential fan out is so fast that even 'huge' amounts of lookahead are still only testing a tiny fraction of the space. This is a reasonable guess, but unfortunately, it turns out to be more fundamental than just 'search is expensive'; there's something pathological in the model tree of completions that leads to degeneration, and definitely does not lead to the best possible completion, the more compute you use (nodes explore). If you use something like beam search, the more you search, the more likely you are to get trapped in one of the notorious repetition traps where it prints 'the the the'. This is a long-standing problem in NMT, and IIRC, I recall reading a paper where they went to the trouble of doing full Viterbi-style brute force to get the exact optimal result (according to the model) to avoid incomplete search possibly screwing things up, and the result were all still garbage. There are a couple theories why, but nothing I consider definitive. You can boost models a lot with limited search (best-of can make a big difference) and with other tricks which ought to be sorta equivalent (self-distillation and InstructGPT's RL finetuning), but it's never clear what tricks will work in advance. So, we'll see! I have a hunch it may just be, like adversarial examples, another blessing-of-scale in waiting.
- mardifoufs 5y agoYeah, this blog is usually very interesting but this is definitely not a good article. A bit disappointing