5 ms·
From the title, not knowing what E and P cores are, I wondered if maybe it meant using software emulation to make a "virtual CPU core" could actually be more en
by phaedrus 4y ago
From the title, not knowing what E and P cores are, I wondered if maybe it meant using software emulation to make a "virtual CPU core" could actually be more energy efficient than running on a physical core. And now I kinda wonder if there is something to that idea...
For example I was just reading a different HN thread were people were complaining about Windows 10 randomly turning your laptop into a hairdryer when it's otherwise idle (pegging the CPU at 100% for an update or scan). It's clear a lot of software doesn't need to run at full hardware speed; maybe I'd rather these background scans use only 10% CPU and take 10 times at long. And I bet it wouldn't take 10 times as long, due to the way latency & etc. scale.
Another example is a lot of software is inefficient for stupid reasons. If we had automated ways to detect the most common stupid code, maybe we just not run it (if the answer makes no difference), or run an optimized replacement. Constrained to run on a physical core, you either need static analysis or dynamic recompilation and are limited in some ways, but running on an emulated core you could do whatever you want with dirty tricks that the process can't detect.
- eternityforest 4y agoLinux file scanners are the worst for this, or were until recently. I have no idea why it was so hard not to occasionally consume 100% CPU and tons of disk IO at inconvenient times, but multiple different indexers have exactly the same issue. I bet you could static analyze to find pure functions, then add a caching layer automatically. A lot of slowness seems to be because a cache wasn't used. Or because something stateless should have been stateful(As in fully recomputing stuff that depends on previous inputs rather than calculating one item at a time).
- smoldesu 4y agoMacOS' indexer is equally as bad. The amount of time's I've seen 10 mdworker threads pinning my CPU is staggering...
- als0 4y agoAFAIK this problem seems to have gone away with the M1 based Macs. Can anyone confirm?
- smoldesu 4y agoI'd imagine it's not an issue since Grand Central Dispatch will only assign it to a single core cluster, which leaves the other cores free to do their own processing.
- ceeplusplus 4y agoI believe indexing tasks only run on e-cores, at least from quick glances at Activity Monitor. I also see more kernel code run on e-cores when I'm compiling code, not sure if Apple is moving blocking syscalls onto e-cores and swapping threads or something (intuition tells me this shouldn't lead to perf gains due to context switch cost, but who knows).
- MBCook 4y agoI’ve been led to believe that all non-interactive tasks only run on the efficiency cores by default. Of course code can request alternate priorities. It’s entirely possible that Apple has written things like the spotlight indexer to explicitly make sure they run that way as well, or it may just be an intentional result of the policy.
- astrange 4y agoSpotlight runs at a low CPU priority, so it doesn’t use a lot of power (theoretically) and on Intel won’t increase P-states. The CPU % doesn’t mean much at all except it’s runnable.
- smoldesu 4y agoIn practice my CPU hits ~80c and makes the system nearly unusable. I just disabled Spotlight indexing altogether.
- sliken 4y agoThe major problems I see are one of two cases. Often just naive implementations run a filescan, run as fast as they can, the OS starts using more ram for caching, then at some point applications need more ram. That triggers a reclaiming pages from cache, writing dirty pages, reading pages from the disk for the application, all while the scan is keeping the I/O pipeline full. The user observes this as a slow/laggy machine, and if you are unlucky enough to have a i7/i9 in a laptop often a fan running flat out. A second case is where a vendor offers "unlimited" backups for a flat price. This creates an incentive for the vendor to make an INCREDIBLY inefficient backup client that consumes substantial ram and CPU resources locally and minimal storage remotely. Typically this means that you aren't allowed to use more efficient clients, and have to use the vendors client. Here's looking at you crashplan.
- Tagbert 4y agoAnd that is why many of us moved from Crashplan to Backblaze and never looked back.
- invalidator 4y agoI set this: echo 1 > /sys/devices/system/cpu/cpufreq/ondemand/ignore_nice_load My backups and indexing are already running nice, so now they loaf along at 1.6GHz. The CPUs only spin up to full speed when I run non-nice work. There's room to tune more. In the old days I would start large jobs (large builds, etc) with nice so they wouldn't interfere with interactive use. Now they would run slower instead of fast-but-yielding. I can always flip ignore_nice_load off if I want the old behavior for competing fast jobs. It's an easy 80% solution, especially because most indexers and backups come preconfigured to run nice.
- chlorion 4y agoOn Linux we have something called "cgroups" which is integrated with the kernel's scheduler. Cgroups can be used to control and limit resources and do exactly what you are describing! Cgroups can be used to limit the maximum amount of CPU bandwidth a certain group can use in a given period with fairly high resolution. For example, you can use the cpu.max variable to limit a group to run for a maximum of 1000 microseconds per second (1s = 1000000us). You can also limit what cores the group can run on, and the "weight" of the group when CPU bandwidth is being contested by other groups! Cgroups can also limit resources such as RAM, swap and IO in a similar way. Systemd has some features to put daemons in cgroups and provides ways to set the various control variables which is pretty handy, and it also has a way to launch regular programs with resource controls applied with systemd-run. There is also libcgroup and the sysfs interface for systems without systemd.