3 ms·
Anything that’s not an actual computational kernel (python interpreter, Pytorch dispatcher, etc.) See the flame graph here in the overhead section: https://hor
by chillee 4y ago
Anything that’s not an actual computational kernel (python interpreter, Pytorch dispatcher, etc.)
See the flame graph here in the overhead section: https://horace.io/brrr_intro.html https://horace.io/brrr_intro.html
- mochomocha 4y agoRight, all this overhead is very pytorch specific and doesn't have to be present in all cases. If you resolve everything at compile-time for a limited class of cases (which is what vectorflow does), it's possible to have a solution where none of this overhead exists at runtime (no dispatch, no indirections, no memory copies nor allocations...), neither during training nor serving. I think pytorch has a genius design and makes excellent tradeoffs. But it's not the perfect solution to every problem. Genericity vs specialization: hard to win on all fronts.
- chillee 4y agoYeah, my main point is that I don't think the issue is memory movement. PyTorch/Tensorflow do care a lot about memory movement, as memory movement doesn't stop being an issue with larger networks.