8 ms·
This article isn't accurate, in the sense that it's actually measuring how long the operating system takes to reach the CPU's maximum clock speed, rather than h
by theevilsharpie 4y ago
This article isn't accurate, in the sense that it's actually measuring how long the operating system takes to reach the CPU's maximum clock speed, rather than how long it physically takes the CPU to reach that speed.
Modern processors can switch frequencies very fast -- generally within microseconds. An example from the machine I'm currently using:
$ grep 'model name' /proc/cpuinfo | uniq; sudo cpupower frequency-info -y -m
model name : Intel(R) Core(TM) i5-4310U CPU @ 2.00GHz
analyzing CPU 0:
maximum transition latency: 20.0 us
Even very deep CPU idle states can be exited in <1 ms.
With respect to the operating system, the amount of time it takes to reach maximum frequency from idle depends on:
- How frequently the OS checks to see if the frequency should be increased
- Whether the OS will step through increasing frequencies, or go straight to the max frequency
- If the OS is stepping through increasing frequencies, how many it needs to step through
- Whether the core is already active, or if it needs to be awakened from a sleep state first
It looks like the OP is using Windows. Windows has a number of hidden (and poorly documented) tunables that control the above settings, and would have a significant impact on how quickly a CPU can reach max frequency. Microsoft used to have an extensive document (specifically, a Word document for Windows Server 2008) describing how to tune these parameters using the `powercfg` CLI tool and provided a detailed reference for each parameter, which I unfortunately can't find anymore. It looks like Microsoft's modern server performance tuning documentation describes a subset of these parameters.[1]
Linux has similar tunables for its `ondemand` CPU scheduler.[2] It looks like the default sampling rate to determine whether to adjust the frequency is 1000 times greater than the transition latency, or about 20 milliseconds on my machine.
I'm not familiar with macOS, but it likely has something similar (although it may not be user-accessible).
[1] https://docs.microsoft.com/en-us/windows-server/administration/performance-tuning/hardware/power/power-performance-tuning https://docs.microsoft.com/en-us/windows-server/administrati...
[2] https://www.kernel.org/doc/html/v5.15/admin-guide/pm/cpufreq.html https://www.kernel.org/doc/html/v5.15/admin-guide/pm/cpufreq...
- vardump 4y agoWhen a comment has the insight you expected from the article. Thanks!
- clamchowder 4y agoHey, author here. The Snapdragon 670, Snapdragon 821, and i5-6600K tests were done on Linux, and the rest were done on Windows. If Windows is delaying boost, it doesn't seem to be any different from Linux. And lack of intermediate states between lowest clock and max clock on all three of those (not considering the little cores) indicates the OS is not stepping through frequencies. Since people don't usually write their own OS, I don't think it's correct to take the "maximum transition latency" reported by the CPU to mean anything, because users will never observe that transition speed. Also, processors that rely on the OS (no "Speed Shift") can transition very fast if their voltages are held high to start with, suggesting most of the latency comes from waiting for voltages to increase. Please also read about Intel's "Speed Shift". While it's fairly new (only debuted in 2015), it means the CPU can clock up by itself without waiting for a transition command from the OS.
- theevilsharpie 4y ago> Since people don't usually write their own OS, I don't think it's correct to take the "maximum transition latency" reported by the CPU to mean anything, because users will never observe that transition speed. The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate (i.e., how frequently it checks whether to adjust the processor's frequency). > Please also read about Intel's "Speed Shift". While it's fairly new (only debuted in 2015), it means the CPU can clock up by itself without waiting for a transition command from the OS. Hardware-managed P-states don't need to wait for the OS to send a command to change the frequency, but the processor still performs periodic sampling to determine what the frequency should be (this happens ever millisecond on modern AMD and Intel hardware), the processor isn't necessarily going to choose the maximum frequency right away, and it's still subject to delays from the OS (e.g., Windows core parking). In any case, a multi-millisecond delay in switching frequency isn't because the processor is waiting for the voltage to increase.
- clamchowder 4y ago> The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate Yes, the governor can play a role. It's visible to the user, which is the point. Also, the ondemand governor is actually irrelevant to the article as the S821 and S670 used the interactive and schedutil governors respectively, and the i5-6600K was using speed shift. I think we're disagreeing because I really don't care about how fast a CPU could pull off frequency transitions if it's never observable to users. I'm looking at how it's observable to user programs, and how fast the transition happens in practice. > Processor still performs periodic sampling... Same as the above, that's not the point of the article. I'm not measuring "what could theoretically happen if you ignore half the steps involved in a frequency transition even though a user realistically cannot avoid them without serious downsides" (like artificially holding the idle voltage high and drawing more idle power, as in the Piledriver example) > In any case, a multi-millisecond delay in switching frequency isn't because the processor is waiting for the voltage to increase. Yes, there are other factors involved besides the voltage increase. I never said it was the only factor, and did mention speed shift taking OS transition commands out of the picture (implying that requiring OS commands does introduce a delay in CPUs without such a feature). If you want to test how fast a CPU can clock up, without influence from OS/processor polling, please do so and publish your results along with your methodology. I think it'd be interesting to see.