3 ms·
This paper partially finds disagreeing evidence: https://arxiv.org/abs/2403.17887 https://arxiv.org/abs/2403.17887
by underlines 3y ago
This paper partially finds disagreeing evidence: https://arxiv.org/abs/2403.17887 https://arxiv.org/abs/2403.17887
- Y_Y 3y agoGood reference. I actually work on this stuff day-to-day which is why I feel qualified to comment on it, though mostly on images rather than natural language. I'll say in my defense that work like this is why I put a little disclaimer. It's well-known that plenty of popular models quantize/prune/sparsify well for some tasks. As the authors propose "current pretraining methods are not properly leveraging the parameters in the deeper layers of the network", this is what I was referring to as the networks not being "at capacity".