3 ms·
It's important to keep in mind when reading anything about GPT-3 that the stated objective in training it wasn't to produce any kind of AGI, but specifically to
by ChefboyOG 6y ago
It's important to keep in mind when reading anything about GPT-3 that the stated objective in training it wasn't to produce any kind of AGI, but specifically to see if a model could perform context-specific language tasks without first being fine-tuned to the specific context. For example, GPT-2 (GPT-3's predecessor) was originally used in AI Dungeon, but it first had to be fine-tuned on a large corpus of choose-your-own-adventure texts. GPT-3 doesn't need to be fine-tuned—it just works.
From the original paper:
"Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model."
Of course, it makes sense that people critique GPT-3 in terms of its actual progress towards human-like intelligence, since every news publication writes about it like it's Skynet and OpenAI's stated goal is AGI. I do think, however, that we miss all the parts of GPT-3 that are exciting and innovative when we view it through the binary of "Is it really understanding as a human does?" (Not that you were reducing it to that, just speaking about the conversation around GPT-3 more broadly.)
- reader_mode 6y agoI agree, and I'm impressed with the result quality - I expect that this approach is viable for creating something like a verbalisation layer for an AGI process.
- alisonkisk 6y agoHow? By design, GPT has no model of the meaning of what it's saying. One of GPT's major tells is that it is so context-free that short passages will contain contradictory facts.
- derefr 6y agoI believe GPT has been demonstrated to be usable for machine translation of long passages without "losing the thread." As long as the input has structure, and GPT just has to "follow along" the structure of the input when generating the output, then it the output is structured intelligibly. Also, I've seen GPT used in recursive fill-in-the-blanks approaches, where it does equally well, since the "skeleton" of its answer is already there. I imagine the task of "translating" between natural language, and a set of expert-system assertion predicates (or vice-versa) would work very well for GPT, and be "scalable." I'd love someone to go ahead and try this experiment, actually. It could potentially be meta-learnable in just a few prompts.