4 ms·
> The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate Yes, the governor can play a role. It's visible
by clamchowder 4y ago
> The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate
Yes, the governor can play a role. It's visible to the user, which is the point. Also, the ondemand governor is actually irrelevant to the article as the S821 and S670 used the interactive and schedutil governors respectively, and the i5-6600K was using speed shift.
I think we're disagreeing because I really don't care about how fast a CPU could pull off frequency transitions if it's never observable to users. I'm looking at how it's observable to user programs, and how fast the transition happens in practice.
> Processor still performs periodic sampling...
Same as the above, that's not the point of the article. I'm not measuring "what could theoretically happen if you ignore half the steps involved in a frequency transition even though a user realistically cannot avoid them without serious downsides" (like artificially holding the idle voltage high and drawing more idle power, as in the Piledriver example)
> In any case, a multi-millisecond delay in switching frequency isn't because the processor is waiting for the voltage to increase.
Yes, there are other factors involved besides the voltage increase. I never said it was the only factor, and did mention speed shift taking OS transition commands out of the picture (implying that requiring OS commands does introduce a delay in CPUs without such a feature).
If you want to test how fast a CPU can clock up, without influence from OS/processor polling, please do so and publish your results along with your methodology. I think it'd be interesting to see.
- vardump 4y ago> I think we're disagreeing because I really don't care about how fast a CPU could pull off frequency transitions if it's never observable to users. I think we should care. What about interrupts that occur during that time? There are hardware devices that will just not work if it takes too long. Too long is usually 0.5 ms or so. However, 20 microseconds is just fine.
- clamchowder 4y agoThat's a different and unrelated topic. If you're concerned about how fast device driver code can respond, well you can get a lot done in 0.5 ms even with the CPU running at 800 MHz or whatever the idle clock is.
- vardump 4y agoSays someone who hasn't debugged a slow Windows graphics related ISR (interrupt), hogging the same CPU core where the your interrupt was supposed to be. (Also, whatever happened to that 50 µs ISR execution time limit? I guess it doesn't apply, if you're Microsoft. Then again, there was some GPU vendor code running as well in the call stack...) 0.5 ms is not really much on Windows.
- clamchowder 4y agoI meant it's unrelated to how fast a CPU clocks up. If something is taking longer than 0.5 ms, you shouldn't be doing it in the ISR. Queue up a DPC and do your longer running processing there, or send it to user space. And yeah it might not be your fault if another driver's ISR was hogging the CPU core. That's just a case of a badly written driver screwing up the world for everyone, because they're not supposed to be doing long running stuff in an ISR in the first place. https://docs.microsoft.com/en-us/windows-hardware/drivers/devtest/example-15--measuring-dpc-isr-time https://docs.microsoft.com/en-us/windows-hardware/drivers/de... says an ISR shouldn't run longer than 25 microseconds. 0.5 ms is an order of magnitude off. Not something going from 800 MHz to locked 4 GHz will fix.
- vardump 4y agoMy ISRs execute under 15 µs, some are as fast as 2 µs. I'm well aware of the DPC queueing. > ISR shouldn't run longer than 25 microseconds Weird, I think I read 50 microseconds somewhere else. Maybe I just remember it wrong? > 0.5 ms is an order of magnitude off. Not something going from 800 MHz to locked 4 GHz will fix. 0.5 ms is actually not that far fetched with higher priority interrupts masking and the delay for Windows ISR dispatching. There are also SMM missing time black holes occasionally. Windows isn't a real-time OS for sure! Although can't blame SMMs on Windows.
- chaps 4y agoThe interrupts are whatever, but the linux scheduler can do some really strange things with thrashing cpu frequencies, all while moving from core to core. So if I have a program running a core at 100%, the scheduler decides to move it to another core, meaning context switching, cache misses, interrupts, frequency scaling, etc. It adds up!
- zerohp 4y agoModern CPUs do not pause during the frequency transition. They switch to a stable clock (even if slower) while waiting for the main PLL to relock at the new frequency.
- rz30221 4y agoI read the article in full, and the data & information was interesting, but I have to say a lot of the points you're making in the comments now was not clear from the article text alone. Another important point is you're measuring the default behavior of the various control systems. One can change that which would allow the user to observe something else.
- clamchowder 4y agoYeah, didn't want to start an article with a five paragraph essay especially when wordpress pagination doesn't work, so I can't get an Anandtech style multi-page article up. And yep. You can even run a CPU at full clock all the time, meaning you will never observe a clock transition time. Cloud providers seem to do that.
- rz30221 4y agoI actually liked the article and I think the first image is really informative. I guess my point is the "clock frequency ramp time" is really due to the interplay of a bunch of different control systems, some in the OS and some not. And when those systems get mixed together, in a somewhat uncontrolled way (which is the case for most PCs), a huge amount of variability is the result and that's what the article did a good job quantifying but IMHO didn't make clear. But at the time scales in your plots "how quickly CPUs change clock speeds" is basically an implementation choice. Just my $0.02
- clamchowder 4y agoThanks :) I guess I can't reply to a 7th level comment, so hopefully this one shows up in the right place. I agree, there are multiple factors at play. But I don't think it's basically an implementation choice. Certainly it looks like it in some cases (S821 on battery, HSW-E and SNB-E). But it doesn't seem to be the case elsewhere. For example, speed shift lowers clock transition time by taking the OS out of the control loop.
- bee_rider 4y agoFor some reason, this site hides the reply button after a certain reply-chain length, but you can just click on the person's name. This will show all their posts, including the one you want to reply to (you may have to look for it), with the reply buttons present. I guess they must be trying to softly dis-incentivize really long chains, but not block them outright? It doesn't really make sense to me...