4 ms·
Hey Mike - I did some work on an industry active learning system a few years ago. The high level finding was that transfer & nonparametric methods were huge win
by nmca 6y ago
Hey Mike - I did some work on an industry active learning system a few years ago. The high level finding was that transfer & nonparametric methods were huge wins, but online uncertainty and "real active learning" hardly worked at all and were super complicated (at least in caffe1 anyway lol).
Can you point to any big breakthroughs that have helped in recent years? Linear Hypermodels (https://arxiv.org/abs/2006.07464 https://arxiv.org/abs/2006.07464) seem promising, but that original experience has left me with some healthy skepticism.
- razcle 6y agoHi NMCA, I'm Raza one of the founders of Humanloop. I totally agree that transfer learning is one of the best strategies for data efficiency and it's pretty common to see people start from large pre-trained models like BERT. Active learning then provides an additional benefit, especially when labels are expensive. For example we've worked with teams where lawyers have to annotate. In terms of breakthroughs in recent years, some things I'd point to would BALD (https://arxiv.org/abs/1112.5745 https://arxiv.org/abs/1112.5745) and its applications in deep learning as well. There has also been progress in coreset methods (https://openreview.net/forum?id=H1aIuk-RW https://openreview.net/forum?id=H1aIuk-RW). I think that you're right that it used to be much to hard to get active learning to work. Part of what we're trying to do is make it easy enough that its worth the benefits.
- razcle 6y agoI'd also point out that people always focus on just the labelling savings from active learning but there are other benefits in practice too: 1) faster feedback on model performance during the annotation process and 2) Better engagement from the annotators as they can see the benefit of their work.
- andy99 6y agoI'd add that there is a deep connection between active learning and understanding the "domain of expertise" of a model, for example what inputs are ambiguous or low confidence, and which are out of distribution. E.g. BALD is a form of out of distribution detection - a point with high disagreement it not only useful to add to the training pool, it is a point for which the current model has no business making a prediction.
- peadarohaodha 6y agoAdding to what Raza as said - to your point on "real active learning" hardly working. I would be interested to hear what approaches you took? We've found that the quality of the uncertainty estimate for your model is quite important for active learning to work well in practise. So applying good approximations for the model uncertainty for modern sized transformer models (like BERT) is an important consideration
- nmca 6y agoIrrc we were using a Bayesian classification model on top of of fixed pretrained features from transfer, something along the lines of refitting a GP every time the number of classes changed. This was images as opposed to text, and after an epoch classification was ~ok but during training (eg the active bit) we didn't see much benefit.