2 ms·
Terrible post, really. Did they need 5 pages to say that a silly dot product micro-bench that is PCIE-bound loses to a CPU implementation? Why are they even com
by juunpp 2y ago
Terrible post, really. Did they need 5 pages to say that a silly dot product micro-bench that is PCIE-bound loses to a CPU implementation? Why are they even comparing computation vs computation + memory transfer?
- refulgentis 2y agoBecause going to the GPU adds the overhead of memory and this is a simple way to demonstrate that. Did you know that already? Congrats, you're in the top 10% of software engineers
- juunpp 2y agoYou still don't need the unwieldy blog post to explain that. The SIMD section is unnecessary. Comparing the naive dot product on CPU vs GPU (+round-trip transfer) would have sufficed.