4 ms·
In "small" setups, you're actually more likely to be overhead-bound if anything (especially on CPU).
by chillee 4y ago
In "small" setups, you're actually more likely to be overhead-bound if anything (especially on CPU).
- mochomocha 4y agoWould you mind defining "overhead"?
- spullara 4y agoLoading the model, preprocessing, post processing, i.e. things not even done in the NN.
- mochomocha 4y agoIf you're loading the model for every prediction, you're doing something really wrong. Pre-processing and post-processing are indeed usually trumping the core computation itself, but as you mentioned it's not done in the NN.
- spullara 4y agoCould be something that is just one once per process or the like. But yeah, if you are doing multiple inferences it would be silly to load the model each time.
- chillee 4y agoAnything that’s not an actual computational kernel (python interpreter, Pytorch dispatcher, etc.) See the flame graph here in the overhead section: https://horace.io/brrr_intro.html https://horace.io/brrr_intro.html
- mochomocha 4y agoRight, all this overhead is very pytorch specific and doesn't have to be present in all cases. If you resolve everything at compile-time for a limited class of cases (which is what vectorflow does), it's possible to have a solution where none of this overhead exists at runtime (no dispatch, no indirections, no memory copies nor allocations...), neither during training nor serving. I think pytorch has a genius design and makes excellent tradeoffs. But it's not the perfect solution to every problem. Genericity vs specialization: hard to win on all fronts.
- chillee 4y agoYeah, my main point is that I don't think the issue is memory movement. PyTorch/Tensorflow do care a lot about memory movement, as memory movement doesn't stop being an issue with larger networks.
- dragontamer 4y agoIMO, one of the big sources of overhead is the PCIe bus. Transferring data to-and-from the PCIe / GPU takes time. If the CPU is faster and can perform the calculation before the PCIe is done transferring, then there's no point even touching the GPU. PCIe is on the scale of ~5000 nanoseconds (20,000 CPU cycles assuming 4GHz). A PCIe write, then read could be ~40,000 CPU cycles or so, and it turns out that a lot of things can be done in that time before the GPU was even _NOTIFIED_ that there was work to do. CPUs can contact other CPU-cores within 50 to 500 nanoseconds or so, depending on the distance. A 64-core CPU could then have 64-cores * 18000 clocks == ~1-million CPU-clock cycles before the GPU gets any message at all.
- westurner 4y agoThere's no mention of GPUs, TPUs, or indeed QPUs in the memory hierarchy described by Wikipedia! Locality of reference ("Data locality") > Spatial and temporal locality usage : https://en.wikipedia.org/wiki/Locality_of_reference https://en.wikipedia.org/wiki/Locality_of_reference Memory hierarchy https://en.wikipedia.org/wiki/Memory_hierarchy https://en.wikipedia.org/wiki/Memory_hierarchy : > Most modern CPUs are so fast that for most program workloads, the bottleneck is the locality of reference of memory accesses and the efficiency of the caching and memory transfer between different levels of the hierarchy [citation needed]. As a result, the CPU spends much of its time idling, waiting for memory I/O to complete. This is sometimes called the space cost, as a larger memory object is more likely to overflow a small/fast level and require use of a larger/slower level. The resulting load on memory use is known as pressure (respectively register pressure, cache pressure, and (main) memory pressure). Terms for data being missing from a higher level and needing to be fetched from a lower level are, respectively: register spilling (due to register pressure: register to cache), cache miss (cache to main memory), and (hard) page fault (main memory to disk). Is it that PCIe is necessarily implied by the debuggable pipeline specified by the von Neumann architecture? https://en.wikipedia.org/wiki/Von_Neumann_architecture https://en.wikipedia.org/wiki/Von_Neumann_architecture Otherwise, computation within RAM avoids interconnect saturation. "Neuromorphic" computing, stateful RAM with operators mapped to particle interactions: Memristor > Derivative devices > memtransistor https://en.wikipedia.org/wiki/Memristor#Derivative_devices https://en.wikipedia.org/wiki/Memristor#Derivative_devices Quantum reservoir computing: https://en.wikipedia.org/wiki/Reservoir_computing#Quantum_reservoir_computing https://en.wikipedia.org/wiki/Reservoir_computing#Quantum_re... But these still need a faster and wider (and qubit) bus than PCIe, too: https://en.wikipedia.org/wiki/PCI_Express https://en.wikipedia.org/wiki/PCI_Express
- dekhn 4y agoanything that prevents my code from running at one or more instruction per cycle.