4 ms·
Skimmed a bit and found some snippets, from which I can't take this paper seriously as it dismisses unsupervised learning / language models over large datasets.
by nfiedel 7y ago
Skimmed a bit and found some snippets, from which I can't take this paper seriously as it dismisses unsupervised learning / language models over large datasets. Yes, sec 4.3.4 briefly discusses recent work in this area, but only briefly and dismisses it by cherry-picking the least positive result of many.
"Only if we have a sufficiently large collection of input-output tuples, in which the
outputs have been appropriately tagged, can we use the data to train a machine so that
it is able, given new inputs sufficiently similar to those in the training data, to predict
corresponding outputs"
This ignores of recent work with large language models that do generalize, zero-shot, to novel tasks.
"supervised learning with core technology end-to-end sequence-to-sequence deep
networks using LSTM (section 4.2.5) with several extensions and variations, including use of GANs"
This reads like something generated from a LM (e.g. GPT-2):
* Where is any mention of attention or Transformer?
* GANs? Have any recent works used GANs successfully for text? There are a few, e.g. CycleGAN, but not widespread afiact.
- YeGoblynQueenne 7y ago>> This ignores of recent work with large language models that do generalize, zero-shot, to novel tasks. Which work is that?
- nfiedel 7y agoOpenAI trained a large (1.5B parameter) Transformer model called GPT-2 on a diverse set of pages from the web. From their paper, GPT-2 "achieves state of the art results on 7 out of 8 tested language modeling datasets in a zero-shot setting" Blog entry with link to the paper: https://openai.com/blog/better-language-models/ https://openai.com/blog/better-language-models/
- YeGoblynQueenne 7y agoThank you for the link. I'm not sure I'm convinced by OpenAI's claim that their model performs zero-shot learning. It depends on what exactly do they mean by zero-shot learning. My understanding, from reading the linked article (again; I remember it from when it was first published) is that, although their GPT-2 model was not trained on task-specific datasets there was no attempt to ensure that testing instances for the various tasks they used to evaluate its zero-shot accuracy were not included in the training set. The training set was a large corpus of 40 gigs of internet text. The test set for e.g. the Winograd Schema challenge was a set o 140 Winograd schemas (i.e. short sentences followed by a shorter question), so it's very likely that the training set had comprehensive coverage of the testing set, for this task anyway. I don't know about the other tasks.