2 ms·
X. Zhang and Y. LeCun had an article [1] about using this technique for regularization very recently. We aren't talking specifically about overfitting in this c
by kastnerkyle 11y ago
X. Zhang and Y. LeCun had an article [1] about using this technique for regularization very recently. We aren't talking specifically about overfitting in this case, but rather a lack of diversity / size in the training data, however it seems like this kind of thing might help for a variety of tasks.
One more idea would be to have a (non-differentiable / REINFORCE?) penalty based on sentence likelihood using doc2vec or skip-thoughts to avoid the "blah and a blah on a blah" type errors that seem to be common in captioning.
One more would be to use TV / YouTube captions, but that data is extremely noisy - even more the COCO captions, unfortunately!
[1] http://arxiv.org/abs/1511.03719 http://arxiv.org/abs/1511.03719
- emcq 11y agoI get excited whenever we can improve performance with manipulations to the training sets such as data transformation tricks (inducing variances like rotations, translations, whitening [1] etc), labeling tricks (such as those like LeCun or [2]), or including information/learnings/regularizations from external corpus like doc2vec. It feels like getting something for free :) Of course there is a limit to how much signal you can extract from a noisy dataset, but the amount of time and human energy invested into creating and improving datasets can be quite large relative to finding another cool trick that can improve performance. However, I wonder which will come first to make these systems "robust" for the average joe's real world uses for these perceptual systems; a large, well labeled dataset or more transformations and semi-supervised learning approaches? [1] http://www.cs.stanford.edu/~acoates/papers/coatesleeng_aistats_2011.pdf http://www.cs.stanford.edu/~acoates/papers/coatesleeng_aista... [2] http://cseweb.ucsd.edu/~elkan/posonly.pdf http://cseweb.ucsd.edu/~elkan/posonly.pdf