5 ms·
I read a post yesterday on a generalized notion of compositionality [0]. It was neat and extolled the virtues of modularity and compositionality and being able
by imh 8y ago
I read a post yesterday on a generalized notion of compositionality [0]. It was neat and extolled the virtues of modularity and compositionality and being able to reason about system by reasoning about its parts.
If I'm understanding OP, this means that to use the AVX-512 instructions well, a compiler that has to think about instruction speed as a function of what other instructions are around it. It might be faster to write operation X with these instructions than without, but only if you don't also write operation Y with them, because then the CPU would get too hot.
That sounds so much harder! Hot damn! I know CPUs are complicated and 1 instruction = 1 cycle is wrong in many ways, but this just sounds especially difficult.
[0] https://news.ycombinator.com/item?id=17923075 https://news.ycombinator.com/item?id=17923075
- spitfire 8y agoEssentially a compiler would have to add energy pressure alongside its existing scheduling for memory atency/pressure. Doable, but someone will have to take the first jump.
- jcranmer 8y agoNot really. The power throttling and voltage gating that goes on takes a long time--at least microseconds, up to a few milliseconds. The scheduling concerns that compilers deal with are worried about around tens to hundreds of clock cycles, a factor of well over a thousand.
- BeeOnRope 8y agoSure, but it's not really a scheduling decision. I think the GP is correct in as much compiler now have to make the hard choice of whether to use any AVX at all, and it's a global trade-off: even though using a few 64-byte moves might be locally optimal, you now need a higher license hence slower CPU and you can only evaluate if that trade-off makes sense in the scope of the larger program: how much such speedups do you get and does it compensate for the lower frequency?
- BeeOnRope 8y agoCurious, does any compiler implement any kind of general algorithm for "memory pressure"? For register allocation (hence pressure), they do I think - but the memory layout, at least in lower level languages, is mostly fixed by the source so I didn't think there was much flexibility there.
- gameswithgo 8y agoThe compiler would need to know what other programs are running on the core, or the computer (depending on CPU/instructions). It has no hope! This is why it is advisable for SIMD libraries or libraries that use SIMD to always offer some levers for this kind of thing. This makes it a pain as you may have to write the same function 2 or 3 times in different SIMD instruction sets, or use a library to do that for you: https://github.com/jackmott/simdeez https://github.com/jackmott/simdeez
- freeman478 8y agoIt seems like this should enable JIT compilers (java/.net/js) to realise even more of their theorical gains.
- Twirrim 8y agoAt the cost of some potentially significant complexity in the code. I wonder if the trade off would be worth it.
- Twirrim 8y agoThis situation gets absolutely awful when you consider that the Bronze and Silver Xeons do even more aggressive down throttling. Bronze speed plummets if only one core is doing AVX instructions. Compilers can't hope to realistically handle this. JITs at least have a chance, but adding in handling for this behaviour surely requires a lot more complexity than I'd imagine most runtime developers would want to add to their code.
- BeeOnRope 8y agoSince bronze and silver largely only have one AVX-512 FP unit, running AVX-512 in the L2 license is almost totally pointless: you'd often be better off running twice as many AVX/AVX2 instructions on the two 256-bit units since you run at a higher frequency and the FLOP/cycle is the same. The exception would be if your kernel can make some good use of other wide instructions such as memory access or shuffles.
- gnufx 8y agoFor a relevant JIT library, see libxsmm, once submitted here with no take-up.
- drb91 8y agoI wonder if one could generate multiple implementations of a code block and select at runtime which is the best one given the current state of the CPU. Obviously this would require some architecture changes to have a multi-address-select jump or whatever, but this fundamentally seems like a problem only solvable with information known at runtime. ...though, come to think about it, this would be pretty easy with a tracing JIT.
- muricula 8y agoIt’s called an indirect jump. If you use C++ virtual functions, C function pointers, or Go interfaces it happens all the time.
- drb91 8y agoThat is an entirely different functionality from what I am suggesting. The select would select based on some other internal state to the cpu than a register.