3 ms·
> > The first part is true of all statistical models: To improve performance by a factor of k, at least k^2 more data points must be used to train the model. T
by BeatLeJuce 5y ago
> > The first part is true of all statistical models: To improve performance by a factor of k, at least k^2 more data points must be used to train the model. The second part of the computational cost comes explicitly from overparameterization. Once accounted for, this yields a total computational cost for improvement of at least k4.
Those claims are entirely new to me, and I've been a researcher in the field for almost 10 years. Where do they come from/what theorems are they based on? It's unfortunate this article doesn't have any citations.
- 6gvONxR4sf7o 5y agoFor the former, you can find reasonably general forms of it in mathematical statistics books. For the latter, I'd love to know too.
- mitmatt 5y agoThe first statement may be referring to the Cramer-Rao bound: https://en.wikipedia.org/wiki/Cram%C3%A9r%E2%80%93Rao_bound https://en.wikipedia.org/wiki/Cram%C3%A9r%E2%80%93Rao_bound But it only applies to estimation (like how well a population parameter can be estimated) in certain regimes, and not e.g. expected risk (like how well one can do at prediction), so I’m not sure how it would apply here.