7 ms·
C-for-Metal: High Performance SIMD Programming on Intel GPUs
- 37ef_ced3 6y agoDomain-specific compilers that generate explicit SIMD code from a high-level specification are even nicer. These can fully exploit the capabilities of the instruction set (e.g., fast permutes, masking, reduced precision floats, large register file, etc.) for a particular domain For example, generating AVX-512 code for convnet inference: https://NN-512.com https://NN-512.com NN-512 does four simultaneous 8x8 Winograd tiles (forward and backward) in the large AVX-512 register file, accelerates strided convolutions by interleaving Fourier transforms (again, with knowledge of the large register file), makes heavy use of the two input VPERMI2PS permutation instructions, generates simplified code with precomputed masks around tensor edges, uses irregular/arbitrary tiling patterns, etc. It generates code like this: https://nn-512.com/example/11 https://nn-512.com/example/11 This kind of compiler can be written for any important domain
- the_optimist 6y agoThis is great, but it doesn't address GPUs. If you built it for GPUs, from what I understand, that outcome would basically look like tensorflow, or maybe tensorflow XLA. Is that right?
- 37ef_ced3 6y agoMy point is that a less general compiler can yield better SIMD code for a particular domain, and be easier to use for a particular domain. And I gave a concrete illustration (NN-512) to support that claim Consider NN-512 (less general) versus Halide (more general). Imagine how hard it would be to make Halide generate the programs that NN-512 generates. It would be a very challenging problem
- the_optimist 6y agoUnderstood: NN-512 is a local optimum in an optimization of hardware and problem structure.
- skavi 6y agoWhy are Intel GPUs designed in such a way that typical GPU languages don’t fully exploit it? Is the new Xe architecture still SIMD?
- pauljurczak 6y agoThe Intel Gen GPU architecture, which includes the newest incarnation Gen12 aka Xe, is a SIMD architecture as opposed to Nvidia and AMD SIMT architecture. The reasons are historical, i.e. CPU centric design: x86 CPU, Larrabee, Xeon-Phi, etc.
- dragontamer 6y agoOpenCL works on Intel GPUs, while CUDA doesn't because CUDA is an NVidia technology. > Is the new Xe architecture still SIMD? SIMD is... pretty much all GPUs do. There's a few scalar bits here and there to speed up if-statements and the like, but the entire point of a GPU is to build a machine for SIMD.
- oivey 6y agoOpenCL is basically dead at this point, too. The de facto standard is CUDA and there aren’t currently any real challengers. Maybe eventually AMD’s ROCm or Intel’s oneAPI will get traction.
- pjmlp 6y agoFor them to get traction, they need to invest in debugger tooling that allows the productivity as on CPUs, and to help language communities other than C and C++ to target GPGPUs. NVidia started doing both around CUDA 3.0, whereas Khronos, AMD and Intel only started paying attention that not everyone wanted to do printf() style debugging with a C dialect until it was too late to get people's attention back.
- MaxBarraclough 6y agoAMD had a good Visual Studio plugin for OpenCL, complete with debugging support, although I believe it's since been discontinued.
- raphlinus 6y agoAnother interesting reference from a few years ago: http://www.joshbarczak.com/blog/?p=1028 http://www.joshbarczak.com/blog/?p=1028 Also read the followups (1120 and 1197), as they go into considerably more detail about the SPMD programming model and some use cases. The author is now at Intel working on ray tracing.
- astrange 6y agoIntel has a previous SPMD compiler here: https://ispc.github.io https://ispc.github.io Although the author seemed to have fled Intel soon after releasing it, and apparently spent the whole development process terrified that corporate politics would make him cancel it.
- einpoklum 6y ago> The SIMT execution model is commonly used for general GPU development. CUDA and OpenCL developers write scalar code that is implicitly parallelized by compiler and hardware. On Intel GPUs, however, this abstraction has profound performance implications as the underlying ISA is SIMD and important hardware capabilities cannot be fully utilized What? That makes no sense. GPU processor cores are basically just SIMD with a different color hat. The SASS assebly simply has _only_ SIMD instructions - and with the full instrunction set being SIMD'ized, it can drop the mention of "this is SIMD" and just pretend individual lanes are instruction-locked threads . So, an OpenCL compiler would do very similar parallelization on a GPU and on an Intel CPU. (It's obviously not exactly the same since the instruction sets do differ, and the widths are not the same, and Intel CPUs has different widths which could all act at the same time etc.) So, the hardware capabilities can be utilized just fine.
- my123 6y agoModern NVIDIA GPUs (since Volta) drop that pretence at the ISA level, each thread has its own instruction pointer there. Your GPU ISA is scalar, not vector on modern NVIDIA machines.
- the_optimist 6y agoCompiling from high-level lang to GPU is a huge problem, and we greatly appreciate efforts to solve it. If I understand correctly, this (CM) allows for C-style fine-level control over a GPU device as though it were a CPU. However, it does not appear to address data transit (critical for performance). Compilation and operator fusing to minimize transit is possibly more important. See Graphcore Poplar, Tensorflow XLA, Arrayfire, Pytorch Glow, etc. Further, this obviously only applies to Intel GPUs, so investing time in utilizing low-level control is possibly a hardware dead-end. Dream world for programmers is one where data transit and hardware architecture are taken into account without living inside a proprietary DSL Conversely, it is obviously against hardware manufacturers' interests to create this. Is MLIR / LLVM going to solve this? This list has been interesting to consider: https://github.com/merrymercy/awesome-tensor-compilers https://github.com/merrymercy/awesome-tensor-compilers
- banachtarski 6y agoI'm not a hardware engineer, but I am a GPU-focused graphics engineer. > C-style fine-level control over a GPU device as though it were a CPU. Personally, I think this is a fool's errand, and this has nothing to do with my desire for job security or anything. When I look at how code in the ML world is written for a GPU for example, it's really easy to see why it's so slow. The CPU and GPU architectures are fundamentally different. Different pipelining architecture, scalar instead of vector, 32/64-wide instruction dispatches, etc. HLSL/GLSL and other such shader languages are perfectly "high level" with other needed intrinsics needed to perform relevant warp level barriers, wave broadcasts/ballots/queries, use LDS storage, execute device level barriers, etc. This isn't to say that high level shader language improvements aren't welcome, but that trying to emulate a CPU is an unfortunate goal.
- mpweiher 6y agoWhat kinds of improvements would you like to see?
- banachtarski 6y ago
- moonbug 6y agoain't no one gonna use that.
- fulafel 6y agoSkimming through the paper, it seems they don't referece or review other recent GPU languages, just OpenCL and CUDA. Seems curious as it's an active area.
- zvr 6y agoThe compiler is open source, available at https://github.com/intel/cm-compiler https://github.com/intel/cm-compiler More documentation at https://01.org/c-for-metal-development-package https://01.org/c-for-metal-development-package