10 ms·
So that's why Xeon Phi has been so successful! /s Real talk, though. Why do you think larabee is the "right" hardware?
by twtw 8y ago
So that's why Xeon Phi has been so successful!
/s
Real talk, though. Why do you think larabee is the "right" hardware?
- hajile 8y agoI'd assume because of the extra serial processing power of a CPU. x86 isn't the right choice though. As the chips get smaller, all the x86 cruft starts to become a larger portion of each core. RISCV with 40-ish instructions plus a SIMD extension would probably be ideal. They've already proven that those designs use less space per core than ARM.
- vl 8y agoThe truth is that for many types/sizes of models AVS2 or AVS512 in Xeons is as fast and GPUs.
- frankchn 8y agoI am interested in seeing benchmarks regarding this. I believe this is true for models with truly huge embeddings, but otherwise if your models are even a little compute dense then GPUs are faster.
- Cybiote 8y agoMy experience is that you need the additional caveat of a streaming fairly homogeneous and highly parallelizable approach or the memory transfer, branching and communication overheads will eat away nearly all the GPU gains. I've also noticed that a lot of the time, people are comparing GPU to naive implementations (and sometimes, implementations written in dynamic languages) instead of to highly tuned BLAS or MKL implementations. For a large swathe of problem types, using the fastest math library will reduce the CPU/GPU gap to less than an order of magnitude or less.
- twtw 8y ago> comparing GPU to naive implementations (and sometimes, implementations written in dynamic languages) instead of to highly tuned BLAS or MKL implementations In your experience, are these comparisons using naive GPU implementations as well? If so, that still seems like a valuable comparison.
- shaklee3 8y agoNo, they're not. Straight front the horse's mouth: https://software.intel.com/en-us/mkl/features/benchmarks https://software.intel.com/en-us/mkl/features/benchmarks In no case will you see it get close to 10TFLOPS. GPUs easily do this, and can approach 100TFLOPS with tensor cores.
- Cybiote 8y agoI don't dispute this. If you can remove memory transfer overhead and meet the requirements I mentioned, then GPUs will be much better. But there are many problems that do not fit that (SIMD per warp) regime. In short, GPUs are no panacea and have their own trade-offs and bottleneck sensitivities, just like everything else.
- shaklee3 8y agoI agree that they are no panacea, but we also have an example of the Intel Xeon phi, which did have high speed hbm memory. That too did not compete with GPUs. I think the majority of it has to do with the GPU chip itself being extremely simple, with no Branch prediction, no branch difference, very simple caching behavior, etc. They are able to allocate a much larger portion of the die to just throw fma units at it.
- vl 8y agoWell, I'm speaking from personal experience, for many (but not all) text-classification IO-intensive models with large embedding tables Xeon on TF compiled with AVX2 or AVX514 and FMA is as fast as GPU. Obviously these models need to be on the lower spectrum of computational complexity.
- shaklee3 8y agoThis is not true. I've benchmarked it, and you can see yourself by looking at the MKL benchmarks versus cuBLAS. cuBLAS is always significantly faster.