4 ms·
> they do not quite care about parameter efficiency. Google Research is pretty big, I used to think like you did but I think it's mostly b/c DeepMind just hogs
by vladf 6y ago
> they do not quite care about parameter efficiency.
Google Research is pretty big, I used to think like you did but I think it's mostly b/c DeepMind just hogs all the spotlight.
Check out PRESS [0] for example.
[0]: https://research.google/pubs/pub46141 https://research.google/pubs/pub46141
- thesz 6y agoThank you! I skimmed over the abstract and will read the paper later, it seems interesting. But you gave me another point to support my view: PRESS uses stochastic gradient, not second-order method like IRLS.
- vladf 6y agoI agree, I was really only proposing PRESS for the "parameter efficiency" part of your comment. It'd be interesting to see some modern takes on IRLS. I think generally this goes against the grain of the Cheap Gradient Principle which is why you see less of it (edit: eh, I think Fisher scoring can be cast in this light). For instance, on modern modelling problems with non-linearities and change points, it's a lot less easy to do something like IRLS in an end-to-end system, but interesting as a research direction.
- thesz 6y agoI also agree with "Cheap Gradient Principle". I see it as a case of "width of two horse backs from Ancient Rome determine Shuttle buster width" (which is untrue but cool as a reference). The very SGD thing was developed because it was the only way to train something like neural network with small memory and, more importantly, in reasonable time. Multiplying of training time by N (number of parameter) meant having good result in a year, not in a day. And today we have large batch training with complex synchronization systems to speed up training even more. Which bring us closer to the whole-dataset training and, I guess, second-order optimization as well.