5 ms·
On the contrary, this suggests that the bitter lesson is alive and kicking. The bitter lesson doesn't say "compute is all you need", it says "only those methods
by grbsh 2y ago
On the contrary, this suggests that the bitter lesson is alive and kicking. The bitter lesson doesn't say "compute is all you need", it says "only those methods which allow you to make better use of hardware as hardware itself scales are relevant".
This chain of thought / reflection method allows you to make better use of the hardware as the hardware itself scales. If a given transformer is N billion parameters, and to solve a harder problem we estimate we need 10N billion parameters, one way to do it is to build a GPU cluster 10x larger.
This method shows that there might be another way: instead train the N billion model differently so that we can use 10x of it at inference time. Say hardware gets 2x better in 2 years -- then this method will be 20x better than now!
- WanderPanda 2y agoI'd be shocked if we don't see diminishing returns in the inference compute scaling laws. We already didn't deserve how clean and predictive the pre-training scaling laws were, no way the universe grants us another boon of that magnitude