4 ms·
> unfortunately, the use of SIMD can make the code less portable and less maintainable Unfortunately general purpose processors aren’t getting any faster, but
by BooneJS 8y ago
> unfortunately, the use of SIMD can make the code less portable and less maintainable
Unfortunately general purpose processors aren’t getting any faster, but they are getting more specialized. At some point in the future package managers for OSS May have to do LLVM compilation on the machine just to specifically target the bits and baubles of the specific architecture it’s running on.
- gameswithgo 8y agoAOT compilation on install! the best of AOT and JIT worlds! Android does this already.
- pjmlp 8y agoYour information is bit outdated regarding Android. Android did AOT on install between 5 and 7 versions. With Android 7, they introduced a mixed model, of an hand written interpreter in Assembly, JIT compiler with PGO, followed by AOT compilation with PGO data when the device is idle. As extra information, AOT on install for mobile phones, was introduced on Windows Phone 8, and has been available in mainframes and some Pascal USCD systems since the early 80's.
- pjmlp 8y agoAs neighboring comment states, a common feature in Android, and in all managed languages with JIT/AOT toolchains.
- dragontamer 8y agoGood GPU (SIMD) code is written in a very, very different manner than good CPU code. Case in point: NVidia has 32 x 32-bit shaders per L1 cache / Shared Memory (a Symmetric Multiprocessor, or SM). Each SM can run 32-warps at a time, for an effective 1024 "shader threads" per L1 cache. In contrast: CPUs have 1-core, with 2-way SMT. That's 2-threads per L1 cache. IBM has a Power9 CPU with 8-way SMT for 8-threads on one core, but that's as far as you get with traditional CPUs. Effectively: GPUs are memory-constrained. Not memory-bandwidth constrained btw (GPUs have a TON of bandwidth), but literally memory constrained. Each shader only has ~500 bytes, probably less, that it can access efficiently (more if you have "uniform" data that can be shared between shaders). And maybe 2MB total per shader before you run out of RAM entirely. --------- GPU code is innately parallel, so any memory allocation you do is multiplied by a thousand fold, or more. Vega64 needs at least 16384 work-items before it has occupancy 1 per vALU (64 compute units, 4 work queues per compute unit, 64 vALUs per queue)... and needs x10 work-items for max occupancy (a total of 163840 work-items in flight). If you have max occupancy 10 and allocate 16-bytes per work-item, and you just allocated 2.6MB of data on the GPU. I'm not kidding. You run out of space very quickly on GPUs. While GPUs struggle with this multiplicative problem per work-item, CPUs have a ton of memory. That's 32kB of L1 cache typically available per thread (64kB per core, but typically you have two threads per core thanks to Hyperthreading). That's the core of GPU vs CPU programming IMO. Dealing with the grossly constrained memory conditions of the GPU environment.
- jcranmer 8y ago> In contrast: CPUs have 1-core, with 2-way SMT. That's 2-threads per L1 cache. Comparing GPUs and CPUs are tricky. What you want to compare is not the number of logical threads that can be independently scheduled [1], but the number of simultaneous FMA units you can access. If you have an AVX2 processor, that's 2 FMA units each processing 8 lanes of 32-bit values, or a total of 16 lanes per L1 cache. AVX-512 has 32 lanes due to the wider vectors. [1] The independent scheduling of a GPU thread is also a bit of a lie. GPU threads are closer to maskable lanes of a wide vector, rather like the mask vectors of AVX-512.
- dragontamer 8y ago> Comparing GPUs and CPUs are tricky I agree it is tricky, but my point in the above post is not to compare processing girth, but instead to compare memory-allocation. If you are aiming for occupancy 10 on Vega64, every 32-bit DWORD you allocate will allocate a total of 655,360 bytes. Yes, this code, a simple 4-byte allocation in OpenCL: __private uint32_t foobar; This code just allocated 655,360 bytes on Vega64 (assuming worst-case Occupancy 10 per queue). While this following OpenCL code: __private uint64_t wow; Just allocated 1.25 MB. By itself. The single line of code. The total amount of memory used adds up surprisingly quickly on GPUs. This is IMO the biggest issue with GPU programming, the absurdly small amount of memory you're expected to work with per shader.
- flukus 8y ago> At some point in the future package managers for OSS May have to do LLVM compilation on the machine just to specifically target the bits and baubles of the specific architecture it’s running on. Future? This sounds basically like what gentoo has been doing for 20 years.