3 ms·
This period of model scaling at all cost is going to be a major black eye on the industry in a couple years. We already know that language models are few shot l
by valine 2y ago
This period of model scaling at all cost is going to be a major black eye on the industry in a couple years. We already know that language models are few shot learners at inference time, and yet OpenAI seems to be happy throwing petaflops of compute training models the slow way.
The question is how can you use in-context learning to optimize the model weights. It’s a fun math problem and it certainly won’t take a billion dollar super computer to solve it.
- johnsutor 2y agoSeems like this is already being answered: https://arxiv.org/abs/2407.10930 https://arxiv.org/abs/2407.10930 https://arxiv.org/abs/2006.04439 https://arxiv.org/abs/2006.04439
- valine 2y agoNot really the first paper is just fine-tuning on synthetic data. The second paper doesn’t optimize the model weights.