3 ms·
I agree with the author on GPT-2. But GPT-3, which became available shortly after this was published, is quite a bit more powerful and there are many commercial
by picodguyo 6y ago
I agree with the author on GPT-2. But GPT-3, which became available shortly after this was published, is quite a bit more powerful and there are many commercial applications being built on it now.
- leereeves 6y agoEven GPT-3 knows nothing about the real world; it's merely trained to repeat the words that most often followed the prompt in its training data. That's obviously not useful for news...if a fact is in the training data, it's not news. It's not useful for "hard facts about how your pension fund is performing" unless you want to know how it performed a long time ago. But I agree there are some applications it is useful for, like education.
- FeepingCreature 6y ago> Even GPT-3 knows nothing about the real world; it's merely trained to repeat the words that most often followed the prompt in its training data. I don't know why that would imply that it knows nothing about the real world, unless the data corpus it is trained on likewise bears no relation to reality...
- nmfisher 6y ago> unless the data corpus it is trained on likewise bears no relation to reality It’s trained on Reddit, so I wouldn’t rule that out.
- probably_wrong 6y agoI do not see how GPT-3 could solve the basic architectural problem that the parent comment quotes, namely, that "driven as it is by information that is ultimately about language use, rather than directly about the real world, it roams untethered to the truth". As an experiment I used a GPT-3-powered website [1] to see what GPT-3 has to say about bears, and the first answer was: > "Weird that every day, there are so many cute/funny/entertaining bears to enjoy online but hardly any on the ground." When asked about beards, the first answer has no relation with beards at all: > "If a person doesn’t constantly outwit, outplay, outlast, others, the strong eat the weak." And then there's that time when GPT-3 told someone to kill themselves [2]. While funny and (mostly) grammatically correct, these "thoughts" are nonsense and no amount of extra parameters is going to solve the disconnection between GPT-3 and reality. I imagine you could condition GPT-3 to generate text for a specific piece of data in such a way that guarantees the correctness of its output, but at that point you might as well throw GPT-3 away and write a rule-based system. [1] https://thoughts.sushant-kumar.com/bears https://thoughts.sushant-kumar.com/bears [2] https://www.nabla.com/blog/gpt-3/ https://www.nabla.com/blog/gpt-3/
- picodguyo 6y agoIs language use not inherently shaped by the real world? The site you tried is a tweet generator, not a question answering site. I prompted GPT-3 with "Bears and beards are different because" and got... "Bears and beards are different because they are not the same thing. Bears are animals. Beards are facial hair. Bears are dangerous. Beards are not. Bears live in the woods. Beards live on your face. Bears eat people. Beards do not." But my original point was mainly that this field is moving fast and the the old school NLG companies (I created one back in the day!) are toast.
- hannasanarion 6y agoThere is more to "the real world" than the definition of words, which is the most that you can expect a language model to learn. Yes it is true that bears are animals. No it is not true that, as GPT-3 said, "There aren't any on the ground"
- visarga 6y agoLet's just wait a couple of years to see if GPT-3 was any good in applications. Doesn't matter what we think, what matters is if it is viable. It's younger sibling DALL-E is capable of language grounded in images, I expect the next version to be multi-modal as well. On another line of research there's effort to tame the horse (GPT) by attaching a secondary neural net. This can monitor language, topic, style and bias and ensure increased accuracy in tasks by auto-learning good prompts. It would make development of applications much easier because the base model which was super expensive to train can be reused many times while the secondary net is small and fast to train. Other efforts are related to including a search engine on an inner loop, to make the language model able to query large collections. Also, there's an open effort to create a huge text corpus, so far 800GB (The Pile). It improves on the GPT-3 training corpus on some categories that were lacking. I think it's safe to say the article is way off the current research level.