3 ms·
are these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?
by kewp 4y ago
are these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?
- colechristensen 4y agoYes, they are being trained, to simplify, to complete sentences. You can then use the resulting model to do lots of things. How you train a model and the inference jobs it can do don't necessarily have to be the same.
- Enginerrrd 4y agoIn essence, yes I think, but... isn't that essentially not much different than what I'm doing in making this comment?
- sebzim4500 4y agoThat's how they are trained initially, but the resulting model isn't all that useful (was SOTA two years ago but this field moves fast). A lot of the utility comes from the later finetuning. You can see this using the examples from the article, every mistake they identify with GPT-3 (which is the unfinetuned version) is answered correctly by chatGPT, which has gone through an extensive finetuning process called RLHF.
- astrange 4y agoThat's how the text decoder works, but the model gets to define "most likely" and an RLHF model uses this to make the text decoder produce useful answers instead.