5 ms·
Seems to me Intel is placing their bet on the compiler doing the lion's share of the work here -- and where hand optimization is needed, another bet that having
by gonewest 10y ago
Seems to me Intel is placing their bet on the compiler doing the lion's share of the work here -- and where hand optimization is needed, another bet that having a consistent instruction set across their big Xeon cores and these MIC devices will make things easier for developers?
- dogma1138 10y agoIntel has been betting on the compoler doing low level optimizations for about forever now, their compilers and software development tools are not a small part of their business. At this point Intel is betting that anyone who's capable of doing large scale low level optimization will be designing their own hardware including the CPU so they'll better focus on high performance computing for the masses. Interoperability with x86 big cores is also what intel wants because it means that the software can run on anything, even GPU based HPC efforts want X86 compatibility this is why AMD and Intel dragged NVIDIA to court a couple years back.
- Twirrim 10y agoIt doesn't seem like a particularly bad approach. I would guess that most developers don't have the skills (or time, for that matter) to be able to efficiently write code at the hand optimisation level. Abstracting this away to the compiler is a safe bet, and developers will soon learn ways to at least generally optimise the code for the compiler, much like people have already learned to optimise code for javac, v8, etc.
- gpderetta 10y agoActually the whole masking and scatter-gather in AVX512 is to simplify the job of the compiler by pretty much allowing any loop to be vectorized relatively trivially. It doesn't really require any new compiler breakthrough, as it is pretty much what all the GPU compilers (i.e. cuda) have been doing for a while already.
- greglindahl 10y agoNot to mention scatter-gather and masking in Cray's vector processors. You do still need dependency understanding in the compiler, but that's pretty easy now compared to the late 1970s.
- gnufx 10y agoI don't know about that, but Intel are putting some effort into relevant libraries (typically free software, other than MKL, I'm pleased to say). An example is the small matrix multiplication library libxsmm, which is written up for Supercomputing 16 as referenced from the repo on github. ("Simple loops"...)
- dbcurtis 10y agoSo, how was that 20 year nap you just woke up from? That's pretty much the way everyone does it now, and has been for some time. The central lesson of RISC is not "fewer, simpler instructions is good", it is "let the compiler do what it does well, and let the hardware do what it does well." The increase in available silicon has been moving that boundary for many years. Ever since we finally had enough silicon to implement out-of-order execution with synchronous exceptions, for example, register allocation is no longer a problem to be solved in the compiler. Classic RISC only made sense when available bandwidth to main memory was at rough parity with on-die memory bandwidth. Those days were brief. But the idea of intelligently dividing optimization tasks between the CPU and the compiler is a timeless idea. The optimal implementation, however, is a function of available technology on both sides.
- innocenat 10y agoHaving consistent instruction set with CPU doesn't help much. Learning new instruction doesn't takes that long as more recent instructions set all shares same ideas and don't really have quirks like in the old days. When you go hand-optimizing, you want to squeeze the last bit of performance. And that (currently) is very CPU-specific. Because all CPUs have different cache/memory system, instruction latency, and pipeline structure. When I was doing low-level work (for video processing), a lot of time I has to strike a balance across various microarchicture. Some more extreme people (including the code by Intel ICC) target each CPU model specifically. Because the MIC has vastly different architecture, the resulting hand-optimized code will be vastly different. You still need to learn both of them anyway.