3 ms·
Actually there is no need for gradient decent at all if you have a linear regression problem, because you can find the min/max exactly by inverting your weight-
by nil-sec 6y ago
Actually there is no need for gradient decent at all if you have a linear regression problem, because you can find the min/max exactly by inverting your weight-matrix (pseudo-inverse if degenerate). Additionally, second order optimisation, i.e. using the Hessian in addition to the gradient is a well studied problem in the literature. My understanding is that it is in general not worth it because calculating the Hessian is very expensive (see e.g. https://arxiv.org/abs/2002.09018 https://arxiv.org/abs/2002.09018).
- James_Henry 6y agoIt seems to me that gradient descent would be, often, a better solution than the moore-penrose inverse. Is it not more efficient in most cases?
- oivey 6y agoOne important example is large, sparse systems/matrices. The pseudo-inverse requires computing an inversion that can be very expensive due to the size of the matrix. That matrix generally will also no longer be sparse, so it will be expensive to store and work with.
- zwaps 6y agoI always thought you‘d use Gram Schmitz for linear regression https://en.m.wikipedia.org/wiki/Gram–Schmidt_process https://en.m.wikipedia.org/wiki/Gram–Schmidt_process
- deleted 6y ago[deleted]