5 ms·
You get a bunch of smart hardware guys into a room, they design this funky exotic architecture. Then the software goes "Allocate these threads to whatever is id
by Traster 4y ago
You get a bunch of smart hardware guys into a room, they design this funky exotic architecture. Then the software goes "Allocate these threads to whatever is idle" and suddenly you've completely lost any possible advantage and are thrashing around with no idea what you're doing. The big-little architecture from Apple was accompanied by software that basically handles that for you. From what I heard there were similar problems with Xeon Phi - great theoretical performance but a very difficult programming model and as a result very challenging sales for the Intel sales guys.
- pjmlp 4y agoConsole history is full of such examples.
- orra 4y agoI remember Mono, a .NET runtime, having to change its allocations to fit inside a smaller core. This was because a thread could start on a bigger core but then be shunted onto a smaller core. That fixed the crashing but obviously came with a loss of performance.
- masklinn 4y agoIt was a bit more complicated than that: the issue was that one of the early big.LITTLE designs (Samsung's Exynos 8890) had different cacheline sizes on the big and little cores. glibc's `__clear_cache` would cache the cacheline size on first call, so if the program was started on a big core then migrated onto a little core it would only flush every other cacheline. Which was an issue for any program needing to explicitely clear the caches, like most any JIT. And the mitigation was not to "change its allocation", it was to bypass libgcc and handroll cache clearing: https://github.com/mono/mono/pull/3549 https://github.com/mono/mono/pull/3549 Source: https://www.mono-project.com/news/2016/09/12/arm64-icache/ https://www.mono-project.com/news/2016/09/12/arm64-icache/ This issue didn't only affect Mono e.g. dolphin (https://github.com/dolphin-emu/dolphin/pull/4204 https://github.com/dolphin-emu/dolphin/pull/4204) and ppsspp (https://github.com/hrydgard/ppsspp/pull/8965 https://github.com/hrydgard/ppsspp/pull/8965) had been hitting the same issue and adopted mono's fix. But fundamentally this is the 8890 being broken: as the Mono post notes, technically nothing precludes core migration in the middle of clearing the cache, which would also lead to broken behaviour, with no mitigation.
- orra 4y agoIt wasn’t my intention to blame Mono. The solution seemed necessary albeit unfortunate. Nonetheless, I appreciate the corrections and extra context you have provided.
- cheschire 4y agoRight, this design model works only if there’s a certification process for software. If they can convince the virus scanning and HIPS companies to implement a certification standard, and then convince businesses that this certification process matters when purchasing enterprise numbers of clients, then it will start to make a dent.
- Klasiaster 4y agoSo true, my experience with the BIG.little ARM platform and a regular Linux kernel is that you always end up with the wrong scheduling and the system underperforms because it uses the little CPU for a compute-intensive single-thread task… I'll just avoid these kind of systems.
- linuxhansl 4y agoI thought that was fixed with Kernel 5.15.35 https://www.phoronix.com/scan.php?page=article&item=linux-51535-adl&num=1 https://www.phoronix.com/scan.php?page=article&item=linux-51... Did you have this kernel, or an earlier one? (Perhaps there're still many problems to fix)
- urthor 4y ago"What Andy giveth, Bill Taketh away." Big little is fundamentally the correct architecture. It's not Intel's fault that the software guys making the Linux Kernel haven't fully supported heterogeneous yet. In the long term, it'll be the optimal call for so many workloads. Thousands of server and laptop workloads care only about performance per watt, not peak single thread.
- Traster 4y agoIt's true that the software kills the theoretical performance that the hardware makes available. But I actually think this is an indictment of hardware not software engineers. It's basically "We've created this incredibly difficult problem for you, good luck" from the hardware engineers. How the hell is the OS meant to know whether this thread is going to turn out to be a rendering operation or a logging thread? With these architectures I think often what the management should really look at is: "We currently employ 100 hardware engineers, and 30 software engineers, with this architecture, we're going to need 50 hardware engineers and 500 software engineers. Do we still think this is the right call?"
- urthor 4y ago"How the hell is the OS meant to know whether this thread is going to turn out to be a rendering operation or a logging thread?" Because this problem is incredibly well solved (I believe?). Scheduling priorities in operating systems an enormously well explored topic. Put simply: you must incorporate heterogenous cores into the OS thread scheduler. There's three actors here: Hardware developers. OS developers. End user software engineers. The hardware developers are running into physical constraints. The fix is to shift the prioritisation between the different core types, big little, to the operating system scheduler.
- pohl 4y agoTrue, and not just the system software, but the larger software ecosystem as well. Apple paved the way for the M1 by years and years of goading developers to be honest with themselves about what QoS priorities they really need on their DispatchQueues, and now there are plenty of threads explicitly marking themselves as being good candidates for running on the E cores.
- zaptrem 4y agoHow did Apple do that? I've seen the documentation while building stuff with Xcode but I never saw anything that would have actually forced my app to behave (outside of iOS).
- pohl 4y agoIt sounds like you're not asking about technical enforcement — as opposed to prodding, which is what I was talking about. You're right, though, they do also apply pressure through technical means. They did this at design time, by thinking through what the default QoS will be when it is left unspecified.
- epsilon_greedy 4y agoAs part of alder lake intel includes thread director which handles thread dispatch via hardware and software, I believe. The main downside as far as I know is you have to run a very recent kernel to have access to this.
- rocqua 4y agoThread director gives hints to the scheduler I believe. Its up to the OS to actually use those hints.