3 ms·
Why is it behaving sparsely? There are only dense operations, right?
by nynx 4y ago
Why is it behaving sparsely? There are only dense operations, right?
- w1nk 4y agoI also have this question, yes it should be. The forward pass should require accessing all the weights AFAIK.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- HarHarVeryFunny 4y agoFrom what I've read there's no evidence it's "behaving sparsely".. That was just offered as a suggestion why it might not be loading all the weights, but makes no sense in terms of the model. It's going to be using all the weights. Another suggestion is that not all of the word/token embedding table might be used, which would be a function of the input used to test, but that would be easy enough to disprove as there would then be different memory usage for different inputs. It seems possible the reported memory usage is lower than reality if that's how mmap/top work. In any case, a good use of mmap it seems, especially since for a multi-layer model layer weights will be used sequentially so paged load-on-demand will work relatively well even in a low memory situation.