4 ms·
Yes but coordinating the computationally expensive parts usually needs to be done across multiple CPU threads. Typically one should feed each GPU from a differe
by jhj 3y ago
Yes but coordinating the computationally expensive parts usually needs to be done across multiple CPU threads. Typically one should feed each GPU from a different CPU thread, even in native C/C++ land, as in many cases the kernels being run on the GPU may not be entirely heavy compute but instead on the order of the kernel launch overhead from the host (1-4 us); not every kernel run would be orders of magnitude longer runtime than the overheads involved. Data loaders and other non-computationally expensive parts (I/O latency or throughput bound things) need to be done as well, which make sense to multi-thread as well.