7 ms·
No it wouldn't because promotion of a process from an E core to a P core isn't free. Microsoft and Intel could literally do this today on existing hardware by t
by Veliladon 3y ago
No it wouldn't because promotion of a process from an E core to a P core isn't free. Microsoft and Intel could literally do this today on existing hardware by trapping unsupported instructions and then having the OS scheduler promote the process but they don't because it takes so stupidly long and requires an immense burst of resources. All this AVX10/128 is a necessary evil because the far better option is for an E core to be able to at least run the code, even if it's not performant.
- dralley 3y agoOr, intel could do like AMD and just design P cores that are space efficient enough to stuff them into a chip like E cores while still getting good performance.
- vlovich123 3y agoHow long is "stupidly long"? I get that it's expensive to transition back and forth repeatedly, but if it's on the order of even tens of milliseconds, you only need to do it once and then forever pin that process to P cores. And you could easily have a preamble in the init that checks the instructions you need to trigger that on boot if you wanted to front-load that cost. Heck, you could have it be a filesystem attribute that's automatically set when the promotion is noticed so that you just do it with 0 overhead after the very first execution.
- kimixa 3y agoExecutable formats like ELF and PE already allow pretty much arbitrary tags - just define one that says "This executable uses instruction set extension X" and you don't even need the cost of a trap or modifying the file/filesystem metadata at runtime. Sure, old apps might not then have that flag, but old apps aren't using avx512 either.
- vlovich123 3y agoI’m sure there are Avx512 apps that wouldn’t have this flag which poses a problem for Linux in terms of back compat.
- robot 3y agothis is already done by libraries and how new instructions are leveraged on cores that have them. During init a library checks availability of instructions and sets function pointers to relevant routines.
- kimixa 3y agoYeah, and you could set the process cpu masks at init time in your own code and do the equivalent. Hell, Intel could add that init step as a flag for their compiler. As mentioned elsewhere, I'm not sure this is hardware design being forced to adapt due to software implementation difficulties, so much as avx512 as a whole may not be worth the hardware area going forward. Even on the larger cores.
- nolist_policy 3y agoGreat then all processes will be running on the P-cores only, with the E-cores idling all day. Note that even simple things like memcpy, which every single program out there will use, will use AVX512!
- kimixa 3y agoGenerally it seems things that benefit from super wide simd also benefit from threading. If productivity apps want to leave performance on the table, so be it. But really I think the entire ecore/pcore split for avx512 is academic, as I'm not sure tying features and flexibility to wider registers makes sense even on pcores, as I'm not sure the hardware area cost is worth the benefit. I honestly wouldn't be surprised if newer architectures don't have the wider registers options on even their larger cores. You're chasing a pretty small market IMHO of people who have datasets large enough to benefit from large registers and wider alus, but not so large it's worth it to pass it over to an even more specialized accelerator. The benchmarks that tend to show 512-wide simd benefits may get even bigger benefits from running on a GPU.
- sylware 3y agoActually not, on "recent"/"not too old" x86_64 architecture, "rep movsX" instructions are now hardware accelerated for short and big memory blocks. Memcpy and other basic memory operations are getting their code generated directly by the compiler which shortcuts the OS runtime.
- diogenes4 3y ago> Microsoft and Intel could literally do this today on existing hardware by trapping unsupported instructions and then having the OS scheduler promote the process but they don't because it takes so stupidly long and requires an immense burst of resources. Why not just determine the right cpu to run on by examining the arch of the binary? Waiting for an instruction failure seems ridiculous.
- trifurcate 3y agoBecause you don't want any AVX512 binary to get stuck on a P-core forever.
- codetrotter 3y agoAlso, if application developers got to choose all of them would build their apps to request the most performance. And then how do we put anything on an efficiency core if every application claims to need to run on a P core. Back to square one it is.
- picture 3y agoWhy not use a prominent OS display "power efficiency" so the user can harass devs about a calendar app running on P core
- tedunangst 3y agoBecause your calendar app happens to use a jpeg decoding library that uses "high power" instructions.
- deadbeeves 3y agoBecause the binary might contain a JIT, or even more simply, it might load a DLL with new code.
- mjan22640 3y ago
- toast0 3y agoAVX 10/128 doesn't sound particularly useful, because you still need to do AVX2 at 256, so you're at best skipping half the 256-bit registers. But still doing some 256-bit calculations (even if it's double pumped) I'd think you could do some pretty good heuristics, like if the thread hit P-core instructions in the last time slice, don't schedule it on an E core for the next one. When a time slice starts on a P-core, leave the specialty instructions disabled, so you can monitor usage --- if you trap, enable it, and return; if that cost is still too high (which it might be), keep track of how many time slices in a row hit the trap, and maybe enable the P-core instruction preemptively for a few slices. Or just, give the program more information and let it decide. If you've got some cores with avx and some without, maybe the process wants to schedule only on the avx cores. Or maybe it can schedule some threads on any core and others need an avx core. As long as the possible permutations at run time aren't too crazy, it's reasonable.