4 ms·
The generated text is a little too on the nose and sure enough: > In each case, we generated more than one response and selected the predicted text that follow
by sparsely 7y ago
The generated text is a little too on the nose and sure enough:
> In each case, we generated more than one response and selected the predicted text that follows each section in the article.
Good article even so.
- hyperbovine 7y agoAgree, I was blown away by the quality of the generated text until realizing that it was simply too good to be true. The fact that they picked and chose makes for a way less impressive demo. Good piece though.
- gwern 7y agoBut on the other hand, 'picking and choosing' can be automated in many ways (eg A/B testing), and in a very real sense, that's what the last research OA released on GPT-2 is about: https://openai.com/blog/fine-tuning-gpt-2/ https://openai.com/blog/fine-tuning-gpt-2/ (Because the human preferences are used to train a critic which then 'picks and chooses' by ranking output by quality to further train the text generator.)
- gwern 7y agoI have been experimenting with the preference-learning code for writing poetry, and it does seem to work: I've already noticed improvements in the generated poetry in increasing the amounts of rhymes, avoiding undesirable text genres like Greek/Latin/Spanish/English-prose, and decreasing the amount of repetition. (One of the most irritating and longstanding failure modes in neural text generation is that it often degenerates into repeating the same phrase many times, so it's pleasing to see that penalizing this in ratings does seem to then cause the generator to learn to avoid it.) Unfortunately, model-free RL training using PPO is extremely slow. Fortunately, I can't see any reason that it has to be PPO or model-free RL at all, since the reward model ought to be fully differentiable and allow calculation of true gradients / model-based RL, so OOM speedups ought to be possible quite easily.
- hammock 7y agoThe cheat you point out IS the state of the art for predictive text, though. When I'm texting on my Google phone, or emailing from Gmail, it gives me 3 possible options to choose from.