3 ms·
What you describe is exactly what practitioners in the field have been doing for years. I think that's why the parent is a bit puzzled at the publication, as it
by emcq 10y ago
What you describe is exactly what practitioners in the field have been doing for years. I think that's why the parent is a bit puzzled at the publication, as it's difficult to understand what's novel.
- joshuamorton 10y agoIndeed, although I'm not as familiar with transfers this far (that is, fine-tuning models is really common with image recognition tasks, but often you take a model trained to recognize objects in images, and train it to recognize different objects in images, but you're still recognizing objects in images). This feels a bit different, in that yesterday I would have had no strong intuition that "a char-rnn can detect text sentiment better than sota". Looking now, I can rationalize that idea. I get why it might make sense, but it was non-obvious. Do you disagree with any of that? (and if you do, I'd love to see this distance of transfer in literature, its always cool to read up on these things)
- backpropaganda 10y agoPractitioners have been doing this for vision, and transfer learning is not very popular in language. This work is the first that I have heard of that uses transfer learning for language.
- Cybiote 10y agoYes, I agree with you but with a caveat. Semi-supervised learning is well known but has, I'll argue, recently fallen out of fashion in favor of throwing gallons more of labeled data at a really big neural net, crossing your fingers and hoping for the best. Usually, the neural net is either a really big conv-net with a novel architecture or a biLSTM with some elaboration on attention (which is actually closer to memory/state). Most of the time, in neural net land, what people are doing with the fine tuning part is taking a model trained on looaads of supervised data, chopping off the head and using those features to train on smaller data. This OpenAI method is different in that it used patterns it learned on its own, instead of the recently more common technique of features extracted from a heavily label trained model to reduce the supervised learning burden in a nearby domain. Arguably yes, this is an ancient technique but it has mostly been forgotten when it became clear that many problems are surmountable with a large enough helping of GPUs and a small moon's worth of data. OpenAI's is a good idea because it makes you say 'yeah that's obvious, pretrain a simple char rnn on loads of free text and oh wait, why has no one tried this before!?' What is interesting here is that such a straight forward method compares so well to glittering methods that laboriously advanced the state of the art. What I also found surprising was that there was a 'neuron' that was tracking something very close to sentiment. Why? A bit of thinking and I came to a simple idea. One way of looking at the LSTM in the practical setting (as opposed to a theoretically Turing Equivalent thing) is as a really big finite state rube goldberg machine. In learning to predict the next character, it makes sense that one set or part of a set of states it can enter/track is extremely correlated with what we humans call sentiment in review text. In summary, the trained model can be thought of as a computable theory of amazon reviews that also works really well on IMDB reviews (and probably short but probably not sarcastic text reviews in general).
- dnautics 10y agothanks for clarifying this- it isn't transfer learning at all, more like the techniques like, digging through LSTMs post hoc to find the neuron responsible for opening and closing quotation marks (insert karpathy youtube vid here), except for a more "high level" feature - in this case, sentiment.