5 ms·
While this is true there is also a lot of interesting research going on that makes model training and adaptation to new tasks much more efficient. For example i
by Datenstrom 5y ago
While this is true there is also a lot of interesting research going on that makes model training and adaptation to new tasks much more efficient. For example in meta-learning methods like Model Agnostic Meta-Learning (MAML)[1] you learn a model that learns to initialize weights for another model such that it converges very quickly for a new task.
For datasets benchmarking new task adaptation there are Meta-Dataset (MD)[2] for Meta-Learning, and the Visual Task Adaptation Benchmark (VTAB)[3] for Representation Learning. Recently to compare the two approaches VTAB+MD[4] was created. Of course, there is also model quantization[5], mixed precision, sparse networks[6], brain floats[7].
Well I've dropped a lot there, but there is a lot more outside the mainstream. I've been building something on a tight budget for a few years so this has been important for success there. We are are very focused on being data and compute efficient.
[1]: https://proceedings.mlr.press/v70/finn17a/finn17a.pdf https://proceedings.mlr.press/v70/finn17a/finn17a.pdf
[2]: https://arxiv.org/abs/1903.03096 https://arxiv.org/abs/1903.03096
[3]: https://arxiv.org/abs/1910.04867 https://arxiv.org/abs/1910.04867
[4]: https://openreview.net/pdf?id=Q0hm0_G1mpH https://openreview.net/pdf?id=Q0hm0_G1mpH
[5]: https://arxiv.org/abs/2105.08819 https://arxiv.org/abs/2105.08819
[6]: https://arxiv.org/abs/2112.13896 https://arxiv.org/abs/2112.13896
[7]: https://en.wikipedia.org/wiki/Bfloat16_floating-point_format https://en.wikipedia.org/wiki/Bfloat16_floating-point_format