4 ms·
But what about "Broken Neural Scaling Laws" (https://arxiv.org/abs/2210.14891 https://arxiv.org/abs/2210.14891)?
by evc123 3y ago
But what about "Broken Neural Scaling Laws" (https://arxiv.org/abs/2210.14891 https://arxiv.org/abs/2210.14891)?
- loonginthetooth 3y agoI think my ignorance is showing here, but that paper's tldr to me seems to be: neural network performance is not a monotonic function of network width. Like the conclusion and the problem statement seem to be trivially equivalent. They admit that this law is only useful if you already know where the 'breaks' are: "If an additional break of sufficient sharpness happens at a scale that is sufficiently larger than the maximum (along the x-axis) of the points used for fitting, there does not (currently) exist a way to extrapolate the scaling behavior after that additional break."