3 ms·
So if I understand correctly, M1 Pro has half the efficiency cores as the M1 (4 vs 2), and despite those cores running over double clock speed on the M1 Pro, th
by Infernal 5y ago
So if I understand correctly, M1 Pro has half the efficiency cores as the M1 (4 vs 2), and despite those cores running over double clock speed on the M1 Pro, the M1 Pro still takes longer to complete multi-threaded tasks at the lowest QoS (the QoS that limits threads to only ever execute on efficiency cores).
This is a nice test and good data, but it seems intuitively correct to me that things would work this way. I guess the unintuitive bit is "if you scale the QoS all the way back to minimum, regular M1 outperforms M1 Pro".
It would be interesting to see which one has the lower power usage during these tasks as well, given that regular M1 is a desktop machine (Mini in this case) and M1 Pro is a laptop.
- Kon-Peki 5y ago> It would be interesting to see which one has the lower power usage during these tasks as well The powermetrics data will have the info necessary to figure this out - it reports the power draw of each core cluster every time it takes a snapshot.
- martinald 5y agoThere's a tool which shows this easily: https://github.com/tlkh/asitop https://github.com/tlkh/asitop
- masklinn 5y agoNote that an important component is this is a followup to a previous article which showed that the M1+ (Pro or Max) E-cores can run above twice the frequency of the M1’s, thus usually yielding equivalent to faster performances on half the cores. This points out that (relatively logically) that only works if the cores are under active load computational load. > It would be interesting to see which one has the lower power usage during these tasks as well, given that regular M1 is a desktop machine (Mini in this case) and M1 Pro is a laptop. The form factor is not really a factor, the M1 is also present in laptops and I don’t think I’ve seen any evidence that the mini’s M1 is tuned in any way. That said the previous articles on the subject (“1 + 1 = 4” and “how M1 E cores win”) indicated that the E cluster of the M1+ reach 200mW at full residency and frequency (200% at 2GHz), while the E cluster of the M1 tops out around 165mW. Note that this is for the entire cluster of respectively 2 and 4 cores.
- my123 5y ago> and despite those cores running over double clock speed on the M1 Pro Note that it's a bit harder than that. The ecores on M1 and M1 Pro/Max run at around the same clock for an all-core workload, but the clock on regular M1 is halved when it's running a low priority job specifically.
- masklinn 5y agoOther way around. The E cores run at 1GHz nominal, when processing low-QoS tasks the Pro/Max will ramp theirs up to about 2 in order to provide similar throughput to the M1 with half the cores. The Pro/Max will ramp their e-cores even further when running high-QoS tasks (spilling over from the P-cores). The M1 can turbo slightly when only 1/2 e-cores are in use, but under full load the e cores remain a hair under 1GHz.
- foobiekr 5y agoThe most likely issue is MacOs has scheduling issues. In Go, it's pretty trivial to do an experiment where you load a bunch of stuff into memory and have a core per thread race through it, basically doing stream, and if you do this the UI will start stuttering prior to going completely non-responsive. Similarly, on the m1 macs, with nothing more than writes to an SMB mount, you can cause the entire system to wedge [this is trivial to reproduce using rclone against a SMB mount]. There's no reason these should ever happen. What can I say, I was curious if the advertised crazy levels of M1 pro memory bandwidth were genuine or conflating the all-in system-wide not-actually-visible-to-cores bandwidth hypothetically available.
- d1sxeyes 5y agoI asked a man with ten fast tractors and forty Ferraris to move 20 people from A to B while only driving tractors, and another man with twenty slow tractors and twenty Ferraris to do the same. Guess what, the guy with 20 tractors won! EDIT: I needed to fix the numbers a bit.