3 ms·
Has anyone empirically assessed the claims of the Bitter Lesson? The article may sound convincing, but ultimately it's just a few anecdotes. It seems to have a
by t_mann 2y ago
Has anyone empirically assessed the claims of the Bitter Lesson? The article may sound convincing, but ultimately it's just a few anecdotes. It seems to have a lot of 'cultural' impact in AI research, so it would be good to have some structured data-based analysis before we dismiss entire research directions.
- slowtrek 2y agoWe got better and better models when we threw more and more compute? I gotta work on my snarkiness. Seriously, that's pretty good empirical evidence. The smaller models we get are all some kind of distillation or student model of a larger model, so they can never claim they are not the result of large compute.
- reportgunner 2y agoBetter and better how ? Isn't this also based on an overcomplicated web of anecdotes ?
- neverokay 2y agoNo.
- t_mann 2y agoAn empirical study could start by quantifying model quality and available compute and plot them over time, see how they correlate, investigate whether there might be confounding factors,... in essence, putting numbers behind the qualitative statements. I see no reason for snarkiness here.
- coderenegade 2y agoThe Bitter Lesson is about general methods that scale to hard problems once you unlock a minimum threshold for compute, so neural nets definitely qualify. Traditional ML methods exist for all of the things we use neural nets for, but none of them are as effective, for a plethora of reasons, but one of the biggest reasons is how much training data they can handle. If you have to invert an NxN matrix, for example, where N is the size of your training set, you aren't getting very far. But a neural net scales to datasets containing billions of samples, and can be adapted to multiple domains that previously had their own special techniques. The bottleneck was being able to train them, and letting go of restrictions like provable optimality. Once we could train them, we quickly discovered that scaling to larger datasets produced models that dominated everything else.