4 ms·
Hey, author here. The Snapdragon 670, Snapdragon 821, and i5-6600K tests were done on Linux, and the rest were done on Windows. If Windows is delaying boost, it
by clamchowder 4y ago
Hey, author here. The Snapdragon 670, Snapdragon 821, and i5-6600K tests were done on Linux, and the rest were done on Windows. If Windows is delaying boost, it doesn't seem to be any different from Linux. And lack of intermediate states between lowest clock and max clock on all three of those (not considering the little cores) indicates the OS is not stepping through frequencies.
Since people don't usually write their own OS, I don't think it's correct to take the "maximum transition latency" reported by the CPU to mean anything, because users will never observe that transition speed. Also, processors that rely on the OS (no "Speed Shift") can transition very fast if their voltages are held high to start with, suggesting most of the latency comes from waiting for voltages to increase.
Please also read about Intel's "Speed Shift". While it's fairly new (only debuted in 2015), it means the CPU can clock up by itself without waiting for a transition command from the OS.
- theevilsharpie 4y ago> Since people don't usually write their own OS, I don't think it's correct to take the "maximum transition latency" reported by the CPU to mean anything, because users will never observe that transition speed. The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate (i.e., how frequently it checks whether to adjust the processor's frequency). > Please also read about Intel's "Speed Shift". While it's fairly new (only debuted in 2015), it means the CPU can clock up by itself without waiting for a transition command from the OS. Hardware-managed P-states don't need to wait for the OS to send a command to change the frequency, but the processor still performs periodic sampling to determine what the frequency should be (this happens ever millisecond on modern AMD and Intel hardware), the processor isn't necessarily going to choose the maximum frequency right away, and it's still subject to delays from the OS (e.g., Windows core parking). In any case, a multi-millisecond delay in switching frequency isn't because the processor is waiting for the voltage to increase.
- clamchowder 4y ago> The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate Yes, the governor can play a role. It's visible to the user, which is the point. Also, the ondemand governor is actually irrelevant to the article as the S821 and S670 used the interactive and schedutil governors respectively, and the i5-6600K was using speed shift. I think we're disagreeing because I really don't care about how fast a CPU could pull off frequency transitions if it's never observable to users. I'm looking at how it's observable to user programs, and how fast the transition happens in practice. > Processor still performs periodic sampling... Same as the above, that's not the point of the article. I'm not measuring "what could theoretically happen if you ignore half the steps involved in a frequency transition even though a user realistically cannot avoid them without serious downsides" (like artificially holding the idle voltage high and drawing more idle power, as in the Piledriver example) > In any case, a multi-millisecond delay in switching frequency isn't because the processor is waiting for the voltage to increase. Yes, there are other factors involved besides the voltage increase. I never said it was the only factor, and did mention speed shift taking OS transition commands out of the picture (implying that requiring OS commands does introduce a delay in CPUs without such a feature). If you want to test how fast a CPU can clock up, without influence from OS/processor polling, please do so and publish your results along with your methodology. I think it'd be interesting to see.
- vardump 4y ago> I think we're disagreeing because I really don't care about how fast a CPU could pull off frequency transitions if it's never observable to users. I think we should care. What about interrupts that occur during that time? There are hardware devices that will just not work if it takes too long. Too long is usually 0.5 ms or so. However, 20 microseconds is just fine.
- clamchowder 4y agoThat's a different and unrelated topic. If you're concerned about how fast device driver code can respond, well you can get a lot done in 0.5 ms even with the CPU running at 800 MHz or whatever the idle clock is.
- vardump 4y agoSays someone who hasn't debugged a slow Windows graphics related ISR (interrupt), hogging the same CPU core where the your interrupt was supposed to be. (Also, whatever happened to that 50 µs ISR execution time limit? I guess it doesn't apply, if you're Microsoft. Then again, there was some GPU vendor code running as well in the call stack...) 0.5 ms is not really much on Windows.
- clamchowder 4y agoI meant it's unrelated to how fast a CPU clocks up. If something is taking longer than 0.5 ms, you shouldn't be doing it in the ISR. Queue up a DPC and do your longer running processing there, or send it to user space. And yeah it might not be your fault if another driver's ISR was hogging the CPU core. That's just a case of a badly written driver screwing up the world for everyone, because they're not supposed to be doing long running stuff in an ISR in the first place. https://docs.microsoft.com/en-us/windows-hardware/drivers/devtest/example-15--measuring-dpc-isr-time https://docs.microsoft.com/en-us/windows-hardware/drivers/de... says an ISR shouldn't run longer than 25 microseconds. 0.5 ms is an order of magnitude off. Not something going from 800 MHz to locked 4 GHz will fix.