3 ms·
No tools available, or at least, none that are easy to use. It's the classic parallelism problem: tools are either too low level (e.g. CUDA/OpenCL for the GPU s
by 14113 10y ago
No tools available, or at least, none that are easy to use. It's the classic parallelism problem: tools are either too low level (e.g. CUDA/OpenCL for the GPU space) for day to day developers to leverage effectively, or too high level to provide the performance in the other 10% of the codebase, wiping out your parallelism benefits. In this space, the low level alternative is manually writing SIMD x86 instructions, or using a similarly low level C library to do the same. The alternative is switching to something like haskell and using one of their high level array manipulation DSLs. In the first case, you have to explicitly manage the parallelism, and there _will_ be bugs there, in the second case you probably lose more performance in the other 90% of your application by switching to (say) haskell than you gain in parallelism.
In terms of automation, SIMD is hard to automatically implement (i.e. through the compiler) with a lot of traditional programming languages (e.g. C), and hard to add to dynamic languages, as you end up adding extra code paths/jit passes for each new type of simd/parallelism hardware construct available.