13 ms·
What should the CPU usage be of a fully-loaded CPU that has been throttled?
- jMyles 5y agoI agree with the author. His thesis: > While I sympathize with this point of view, I feel that reporting the CPU usage at 50% is a more accurate representation of the situation.
- bombcar 5y agoIt’s a matter of what question you’re asking - how much CPU is X using and how much headroom does it have vs why is my computer slow.
- toxik 5y agoMaldebrot set? Just look the word up my dude.
- sgerenser 5y agoI assume it was just a typo.
- boulos 5y agoAs long as we’re being pedantic, Mandelbrot is a name.
- toxik 5y agoNames are words.
- _Microft 5y agoNo, names can consist of more than one word. "New York City" for example is three words making up one name.
- toxik 5y agoNew York City is a proper noun and can for grammatical purposes be considered one word. It's not as clear cut as programmers would like to believe. https://www.quora.com/Are-names-considered-words?share=1 https://www.quora.com/Are-names-considered-words?share=1
- _Microft 5y ago"New York City" is a proper name but not a proper noun. You may want to look up the difference if you are convinced otherwise.
- toxik 5y ago"Current linguistics makes a distinction between proper nouns and proper names but this distinction is not universally observed and sometimes it is observed but not rigorously." Unbelievable how dense people play sometimes. It was clear what I meant when I said word, and in fact, I did refer to a single word if you want to play that game: Mandelbrot. "Maldebrot set" or whatever it said before wouldn't even give you the correct search results. It serves a purpose to correct this, whereas this is just grammatical bickering for no apparent reason other than "he started it". I pointed out an important mistake, you're just trying to find substance in an argument that holds as much water as a sieve.
- mercora 5y ago> Another theory is that this should report 50% CPU usage, because even though that CPU-intensive program is causing the CPU to consume all of its available cycles, it is not consuming all off the cycles that are potentially available. if they would be available wouldnt the CPU scaling kickin and ramp up the frequency as required? i think in this case it should really show up as 50%... otherwise its somewhat similar to battery charge status... apparently nobody wants to know what percentage of the initial full capacity is left but just how much is left from what is potentially available... so in a low power state were throttling is done permanently at some frequency i would just like to know how much capacity is actually left...
- buran77 5y agoThe CPU may not be able to scale up the frequency to "100%" at that time due to factors that cannot be overridden, like a high-temp situation. The problem is what is 100% when CPUs have gotten so complex around the frequency. What's the absolute reference, the base frequency, the single core turbo, the multi core turbo, overclock, AVX? For batteries you usually get 2 estimates, a capacity percentage of the total practically possible, and one for how much usage you can get out of that based on recent or estimated usage.
- wonnage 5y agoI think Apple's approach of attributing the capacity not available due to throttling (or various other reasons) to kernel_task might be the best. You can tell something is eating cpu, stuff isn't just stuck at 50% while the rest of the system looks idle. Although then you have users trying to figure out how to kill kernel_task, which isn't great either...
- bombcar 5y agoYeah until I learned how to see the throttling with “pmset -g thermlog” I didn’t realize I was throttling so much. I feel that should be made clear on the various CPU usage utilities.
- kuschku 5y agoThat’s actually an awesome idea, just create a virtual "cpu_throttling" task and attribute the throttling to this task.
- gervu 5y agoUsers trying to kill kernel_task sounds like an uncaught design problem. Replace the kill options in context drop-downs with "Why can't I kill kernel_task?" and have it pop open a short explanation, do the same as a feedback message if anyone tries to kill in the terminal. If the system doesn't do that already, either nobody considered that users would try that particular nonsensical thing, or nobody cared enough to let users find out immediately instead of wasting their time thinking the computer was at fault.
- orthoxerox 5y agoThat's actually a great approach, just rename the process to "thermal_throttle" and it all becomes clear.
- spullara 5y agoRelated, 50% CPU usage on a hyper-threaded CPU isn't 50%. It is usually closer to 80-90% depending on the workload. Something to watch out for when monitoring.
- undfg 5y agoInteresting. Does this mean that if you are not going to use all HT threads it's better to turn off HT?
- Dylan16807 5y agoNo, don't turn it off. One thread on a core will go full speed, and two threads on a core will do more work than one thread. It's just that your utilization graph will be misleading. A naive graph will assume that two threads do twice as much work as one, but the real improvement is much smaller. If you turn off hyperthreading and keep the same exact workload, then instead of "the graph says 50% but it's really more like 80-90%", you'll have "the graph says 100% and it's correct". The numbers now accurately represent your lower capacity.
- wtallis 5y agoDepends on the machine. The earliest implementations of HT worked by statically partitioning various caches and other resources in the processor core in half, which meant that a single-threaded process really could slow down by having HT enabled but not actively used. Newer desktop-class processors tend to have no significant downsides to leaving HT enabled, but there might still be some SMT implementations on niche products that don't handle this well.
- Dylan16807 5y ago> The earliest implementations of HT worked by statically partitioning Do you mean SMT in general? I don't think hyperthreading specifically has ever done that, but if I'm wrong I'd love to know more. (And AMD's version falls under "newer desktop-class processors")
- wtallis 5y agoWhen the OS has asked the CPU to slow down to more closely match the performance currently required by the software, then it is somewhat misleading to report that an application is using 90+% of the CPU time, even if the CPU is actually spending 90+% of its time running that application. However, when the CPU's speed has been reduced because it's too hot or the system is otherwise unable to allow the processor to sustain its full clock speed, you absolutely should see the Task Manager reporting 100% utilization. This scenario is what more often comes to mind when the term "throttled" is used. If a hardware platform's power management capabilities make it impractical for the operating system to satisfy both of the above goals, then it should favor the latter goal, and err on the side of not lying to the user when your system is truly operating at its limits.
- taneq 5y ago> However, when the CPU's speed has been reduced because it's too hot … Maybe just measure overall system load as current CPU temperature as a percentage of maximum?
- wtallis 5y agoThat's a decent method for some purposes, though it's not without its own flaws. For systems using integrated graphics or any laptop, the CPU's thermals are intertwined with the GPU's. You may also have a workload that sends a CPU core to its highest stable clock speed and safe voltage, but is still cooled well enough that the CPU doesn't get close to its temperature limits. Or you could have a situation where the CPU can tolerate a higher temperature, but it has to throttle so as not to burn the user because a fair chunk of that heat is being conducted up through the keyboard and down into the lap.
- MauranKilom 5y ago> but it has to throttle so as not to burn the user because a fair chunk of that heat is being conducted up through the keyboard and down into the lap. I've never used a mobile device that seemed to have such considerations. They all seem perfectly happy to make their surface scalding hot. Do manufacturers actually care about this?
- magila 5y agoA key problem with this proposal is that modern CPUs do not have a single definitive maximum frequency. You have the base frequency which is rarely relevant outside of synthetic workloads then you have a variety of turbo frequencies which interact in complex ways. AMD's latest CPUs don't even have clear upper bounds on their frequency scaling logic. It's a big black box and the results can vary depending on the workload, temperature, silicon quality, phase of moon, etc.
- awesomeusername 5y agoThe OS can keep a record over time of the maximum frequency each CPU core has ever hit. This will take into account machine to machine variance, and even environmental factors effecting maximum speed.
- wtallis 5y agoThat strategy guarantees that a process that runs for multiple minutes while consuming all available CPU cycles will be reported as using 100% CPU at most during the first few seconds, after which it will usually be reported as using somewhere less than 90%, and realistically could be reported as low as 65%. How is this helpful?
- jeffbee 5y agoEstimating x86 core frequency is a lot trickier than you've implied.
- toast0 5y agoOh, but it gets more fun. The same operation can take more clocks for various reasons... If I really undervolt my Zen 2 apu, power usage and benchmarks go way down, but clock frequency stays high; the CPU is clock stretching and it gets a lot less work done. Anyway, current processors run a separate clock per core, and maximum clocks are only available when a small number of cores are active; if all cores are busy, that should really be 100%, even if each core is only doing 80% of max for a single core. Mostly, I want to see % of time cpu is busy, and separately, stats on how throttled the cpu is, because it's hard to combine both into a coherent number. Maybe also some idea of how much of the core is being exercised, if it can be easily measured... I'd love to know when a program is keeping the cpu busy, but not making good use of it.
- tedunangst 5y agoWhat if my program has a large working set and the CPU spends 50% of its cycles waiting for memory fetches to complete? What is its CPU usage?
- annoyingnoob 5y agoI think we can have more nuance than a choice of 50% or 100%. I fall in the 'Show 100%' camp for sure, showing 50% but not indicating why isn't particularly helpful without knowing the background here. I think a separate indicator for an overall throttling condition would be helpful. Show a 100% usage and a throttle indicator together.
- buran77 5y agoDetermining how much is in use out of what is practically available is relatively easy. But determining what is 100% is still difficult to decide on, CPUs have too many frequency options to decide what the "real" 100% point is. Even picking the base frequency as the reference would be a pain to properly monitor and predict, and to take decisions based on that.
- annoyingnoob 5y agoIts more about determining what throttling means, 100% for conditions. In any case, the tools we are used to using have traditionally showed a percentage scale of usage - seems like a reasonable measure even if its hard to measure.
- Tempest1981 5y ago> but not indicating why Agree, you really want 2 indicators, vs 1 confusing hybrid.
- sundvor 5y agoAgree 100% (he!). The utilisation really needs to show busy as in % of presently available resources, and displaying what those are is a completely different concern. For Windows, if anything I'd love to see an additional metric, where the app resource usage is presented in % of available cores. Eg if I have a single thread app (i.e. a game) on let's say a 10 core CPU with SMP disabled, it might report 10% usage. But looking at the core graphs we see one at 100% with the other 9 idling. So the app is really maxed out at its architectural limit - ie 100%. It would be nice to have an additional column representing this - and maybe an idea of the foreground app's available % usage in other reporting (am I cpu limited, or GPU? people have bought Intel over AMD CPUs for this metric even though multi thread workload capacity has been far greater in AMDs). Taking a step back again, especially for Ryzen CPUs unless manually locked you'll have different per core frequencies - so the article's suggestion becomes an almost impossible mission: A Mandelbrot processor can be split across all threads rather equally for an unthrottled 10 core run say at 4ghz, but a core limited config may be able to run it at 4.5ghz prior to thermal throttling - what is the real max to report in a 10 core run following that then? And on a very cold day, with fans at a higher RPM, the core limited run might hit a higher 4.8ghz - now what? You just can't; reporting on current CPU capability is a separate concern altogether.
- woofie11 5y agoIt's the testing problem: Trying to report squish too many numbers into one. The right reporting is to report both the numerator and the denominator. I'm using 800 MIPS out of 1200 MIPS available.
- jzwinck 5y agoIf my processor is marketed as 1200 MIPS but can run at 2000 MIPS for up to 0.5 seconds out of every 15 second window (due to thermal constraints), but only up to 1700 MIPS when executing SIMD instructions for 0.5 seconds and 1000 MIPS thereafter, and only up to 900 MIPS when it gets too hot because my laptop is sitting on a blanket instead of a hard surface, what is the correct denominator?
- woofie11 5y agoThat's the point: it changes. One representation is a line graph with two lines -- one is showing the maximum MIPS (right now), and one is showing the actively-used MIPS. Another is a pie chart, where the size is the chart changes based on current capacity, and it is filled to how much of that capacity you're using. And so on. You have two dimensions. You don't want to squash that into a 1D representation.
- wmf 5y agoPower saving and "throttling" (usually it's un-turboing) are different cases that shouldn't be conflated; in one case the processor could run faster but in the other it can't. Ultimately we may want different metrics depending on what they're going to be used for. If you calculate relative to base frequency you will get utilization over 100% which is going to confuse some people. Linux has done some work in this area with frequency-invariant utilization tracking: https://lwn.net/Articles/816388/ https://lwn.net/Articles/816388/
- foobar33333 5y agoMacOS has a weird solution to throttling. They put a fake process in the process list which looks like it consuming x% of the resources but really it is just blocking some usage to allow the CPU to cool.
- bhj 5y agokernel_task is "real", but yeah, part of its responsibility is issuing NOOPs for thermal control: https://github.com/apple/darwin-xnu/blob/main/osfmk/kern/thread.h#L237 https://github.com/apple/darwin-xnu/blob/main/osfmk/kern/thr...
- pabs3 5y agoLinux has the same solution on my machine.
- leucineleprec0n 5y agoI love this tbh. It's a great indicator and allows me to easily use existing tracking/functions in Activity Monitor to assay thermal headroom at particular points etc
- _ph_ 5y agoI wished they would have a separate, properly named process for it. My old MB Pro (late 2015) was throtteling a lot and it took me very long to recognize, as it was just "kernel task" gobbling up cpu. Getting the machine cleaned out at a service center did help a lot, I wonder, whether redoing the thermal paste on the cpu would have improved it further, but I ended up replacing it with a current model.
- jeffbee 5y agoI don't like to speak about "throttling" because all modern CPUs are in a closed-loop control system where the capacity of one core-second varies. There is no question about whether your CPU is throttled. It is, always. That leads to all the uncertainty about the denominator. We know how many cycles passed while a certain thread had the CPU, but we don't have very good ways to estimate the number of cycles that were available. If you take the analysis one layer deeper, does a program that waits on main memory while chasing pointers randomly use 100% of the CPU, or does it waste 99% of it, since it's not using most of the execution resources? Such a program could be said to be using 100% CPU time, but it won't respond to higher CPU clock speeds. When waiting for loads it makes no difference if time passes at 4GHz or 400MHz. So anyway, it is complicated.
- SavantIdiot 5y agoThis is precisely why all of the large volume cloud server farms I worked with turn off throttling: they need 100% predictable CPU utilization. I worked on power control strategies at Intel for quite some time, and we would often joke in server (Xeon) parts that it was pointless because all of our work was disabled. Early throttles were 50% duty cycles, then L1 bubble injections, then V/F frequency scaling. The author only addresses the early mechanisms, but it gets even more complex with the PCUs in the later Xeons. It is not an easy question to answer, but I think the question can be modified. A single number doesn't solve the problem, you need to know utilization in the context of throttling (and magnitude). Then decide what you are trying to solve: scheduling or app-level throttling? Personally, the Hz denominator should change if it used in utilization , since that covers the majority of cases. Any other case should read both metrics. EDIT: Removed generalization and worded as anecdote.
- jeffbee 5y agoI have managed several extremely large fleets of computers and I can tell you they all use frequency scaling. I seriously doubt that your statement applies to "most server farms" when properly weighted.
- SavantIdiot 5y agoTrue, I should have more accurately said: all large volume cloud vendors I worked with.
- jeffbee 5y agoSo, you've never used GCE, EC2, or Azure? Because they all offer frequency scaling.
- SavantIdiot 5y agoYes, I use them all the time. Did you see the part where I was talking about my customers?
- 5y ago
- tumblewit 5y agoIf you show 50% usage then there are going to be plenty of customer service calls where customers complain about poor performance of the PC and then laptops will be compared based on this number by those that are not tech savvy …
- eevilspock 5y agoMaybe ditch relative (50%) for absolute (Hz)? That's what we do for other things that have no fixed upper bound, e.g. disk/network i/o. We even do this for memory now given dynamic virtual memory, etc.
- userbinator 5y agoAs far as I know, "CPU usage" of a process has always been "percentage of time spent in it" so I think anything else is overthinking/overcomplicating things. It makes perfect sense for CPU usage to go up if the frequency starts falling due to throttling. That's why a separate indication of current CPU frequency is necessary.
- dataflow 5y agoI find this intuitive for thermal throttling, but not for power throttling, though I'm not sure changing the definition makes sense now (it might not even be practical). I've frequently fired up some process manager and seen it at > 50% CPU despite the system not running anything interesting, and it always takes me a second to realize that it's throttled down to like 400 MHz and the process manager itself is consuming what's left of the CPU.
- temac 5y agoIt only makes sense to report e.g. 50% in case of throttling if you also report e.g. 130% in case of boosting (on a single core). Which could be useful. Now throw HT in the mix and loose your mind. I'm not sure there is a really better solution. Just document the one you choose, please!
- dataflow 5y ago> half-speed for whatever reason IMHO the reason actually does matter. Utilization should be relative to the maximum frequency the CPU could be running the same instructions at. If the CPU is throttling to save power, then it could be running the same instructions at a higher frequency, so utilization should be relative to the higher one. If it's throttling to lower the temperature, then it can't be running the same instructions at a higher frequency, so it's already maxed out at 100%.
- nightpool 5y agothat doesn't solve the "relative measures make the graph useless" problem though, does it?
- dataflow 5y agoHm I felt it does, but I might've missed some situation. Why do you feel it doesn't? Could you describe a scenario where it'd be misleading?
- nomel 5y ago> If it's throttling to lower the temperature… This gives the ability to go over 100% before thermals equalize, since max power would follow a Newton cooling curve.
- dataflow 5y agoI don't follow. Whatever frequency the CPU is currently running at is a lower bound on the maximum frequency it could be running at, so it shouldn't exceed 100%?
- nomel 5y agoBut the thermals of the system means it can, for short periods of time, run much faster than long periods of time. I don't think a constantly changing 100%, that changes with temperature, is very useful. I would be more interested in whatever the sustainable 100% is, for normal room temperature.
- formerly_proven 5y ago"CPU usage" as in "how many time-slices did this process eat?" is pretty easy to understand and generally points at the right things ("What's eating all that CPU?", "Is this application using the correct number of threads?" etc.) Trying to express "How much of the hypothetically available computational resources of the CPU did this application consume?" in a single number would seem like a futile exercise at best. VTune used to have something like this which IIRC was based on using all cores in parallel sections and IPC or something like that. It wasn't very meaningful, and is impacted by all sorts of factors.
- deleted 5y ago[deleted]
- eyesee 5y agoMaybe there should be a separate indication of throttling so we don’t conflate with CPU usage. Throttling is essentially determined by three factors: Energy usage, thermal saturation, and cooling. A bath tub analogy comes to mind: Energy usage represented by flow from a faucet, the heat sink is the water level in the tub, and the drain represents the cooling rate. Only energy usage could be plotted instantaneously while the others may have to be modeled and could change based on environmental factors.
- exabrial 5y agoIt should bea bar chart with the entire bar being the potential, and a red area showing the throttling as a percent. Throttling + actual execution total to 100%
- yuliyp 5y agoI don't think this makes much sense. The maximum throughput of a CPU in terms of instructions depends on so many factors (thermal, memory, instruction mix) that trying to summarize all of them into a linearly-scaling "utilization" metric is a bit tricky. You can measure the fraction of time CPU cores are busy, the number of instructions executed, frequencies, etc. to get an idea, but only experimentation will tell you what "100%" of a system's capacity is.
- bluedino 5y agoReporting the clockspeed along with the used percentage makes the most sense. “I’m using 100% of my cpu but I’m only at 1.2GHz, something is wrong”
- ineedasername 5y agoWhy not show the absolute % CPU usage and the % relative to the throttled capacity? No need to choose. For that matter, if throttled, also indicate the reason: heat, performance (any other reasons in might be?)
- randyrand 5y agoIf you want that approach, just add a line that says "thermal throttling: 50%" at the top. Then it still adds to 100% usage.
- hownottowrite 5y agoAfrican or European?
- webkike 5y agoIf you say “there are two perspectives on reporting this fact”, what you really mean to say is “there are two things to report here”
- 0-_-0 5y agoExactly, why not report both? I for one would want to know both.
- catern 5y agoWhichever one is easier to implement. Worse is better.
- kijin 5y agoLinux VMs have the concept of "steal". It represents CPU cycles that were supposed to be available, but were taken away by the hypervisor for various reasons. Steal appears in CPU usage stats alongside other types of wasted cycles such as interrupt handling and I/O wait. Perhaps that's something Microsoft can borrow and improve upon.
- RandomBK 5y agoAnother approach may be to take inspiration from the `cpu load` metric on *nix systems and go _above_ 100%. In this example, the CPU usage would be `200%`: The system would like to be doing twice as much as it's currently doing, but something's throttling it. Of course, this opens up other issues with how to aggregate multiple cores, what the benchmark for 'max' should be, etc. Perhaps the more fundamental answer is that there's no single metric that can sum up the situation for all use cases, in which case displaying '100%' would be more useful for a typical consumer while exposing multiple detailed metrics would be more useful for system admins and power users.
- axaxs 5y agoNot following. Linux reports each core as 100%. So an 8 core machine maxes at 800%. So seeing 200 load doesn't indicate to me that it's throttled, but that it's using the equivalent of two full cores. Or did I misunderstand?
- sharedfrog 5y agoThey're saying it can go over 1(00%) per core.
- viraptor 5y agoThe context was load average, not CPU usage percentage. The load of 1 means (more or less, I'm simplifying) "all the time there one task ready to be run, so no other task is starved". In most systems/situations you'd see load <1. But there are specific cases where it would be silly high. For example I've seen >400 on a VoIP conference server. It's more of a "pressure" / "need" measurement. And yeah, applying that to per-process CPU measurement would be interesting.
- rahimiali 5y agoYou might be thinking of the load average, not the cpu consumption of a particular process as reported, by, say, top or ps. The article is about the latter kind of tool.
- deleted 5y ago[deleted]
- sfisthemoon 5y agoI’d love to see how many watts my process is using. After all, energy is the resource each thread is really consuming.
- pabs3 5y agoOn Linux there is powertop for that: https://01.org/powertop/ https://01.org/powertop/
- sfisthemoon 5y agoMac and Windows have something similar, though not as specific. They all appear to be rough estimates based on the CPU time and the frequency at the time the process ran.
- johnklos 5y agoIt makes me sad that people at Microsoft don't seem to learn from history, because we have precedent that makes perfect sense: report percent (or fractions of 1) between work and what'd be 100% at that moment.
- moonchild 5y agoHere's another problem: what if a program is i/o-bottlenecked, and taking up 50% of the CPU's cycles. Because the CPU utilization is not 100%, it clocks down to 50% of its maximum clock rate, so now the program is taking up 100% of the CPU's cycles. This isn't throttling, it's just regular power management. Clearly that's a different kind of situation; how do you distinguish the two?
- dannyw 5y agoOn OSX, I recall there's a fake "throttled" daemon that reports the CPU usage lost due to thermal throttling. Name is probably wrong but it definitely exists.
- leucineleprec0n 5y agoya kernel_task
- 6510 5y ago% is just not the right unit.
- matco11 5y agoPerhaps, the system should report two numbers: the % usage of the CPU overall resources and the % usage of the throttled resources. There is no constraint that one has to use only one number. Perhaps, in 2021, CPUs are complex enough to deserve more than one number to give a useful representation of utilization, as it works with all sorts of factories/plants. Also, who would love to see in the metrics reported CPU usage per core?
- leucineleprec0n 5y agothank you. I'd hazard a guess that an additional step toward information demarcation would prove beneficial: Bifurcate the task manager into an "everyday" mode by default that displays /currentState where curentState = throttling, user-lowered or heightened TDP via TDP UP + ample cooling for sustained loads: Indicate currenState next to a colored icon set & standard percentages. Then, implement advanced mode that contains a deeply similar but mildly extended UI. Imagine Office with 1 ribbon tab, cleaned up, vs 3, filled to the brim with choice, tools. In principe, thinking of keeping the diff small but noticeable only slightly, like that.
- Const-me 5y ago> “dynamic frequency scaling”, a feature that allows software to instruct the CPU to run at a lower speed, commonly known as “CPU throttling” AFAIK that's not how it works in modern CPUs. Dynamic frequency scaling is a black box implemented in hardware. Software, and even OS kernels, have very little input over that. They can't subscribe for status updates. Even just getting the current frequency seems impossible, only indirectly by comparing RDTSC output (absolute time unscaled) with performance counters.
- CodeWriter23 5y agoI think this article asks the wrong question. The correct question IMO is why is the frequency throttled instead of maxed out when a program is saturating the CPU?
- concerned_user 5y agoIf it is thermal throttling then it means cpu is maxed out for the given conditions.
- CodeWriter23 5y agoI agree with that, but it seems like the use case specified is ordinary power conservation mode, not critical, thermally-induced power reduction mode. And also, why are systems shipped with inadequate thermal control?
- ksec 5y agoI talked about TDP Computing a lot. We are basically limited by Cooling of a product designed. How about another measure, % of CPU TDP? Or some form of TDP measuring. If it is 100% CPU TDP, I know it is pushing as hard as it can. ( But then when your cooling aged you will be running at 100% TDP but at lower clock speed without realising it. ) Thinking about it this simple subject is really complicated.
- wruza 5y agoI’d add “Energy saving / throttling” process and accounted it in another color on the graph (green vs blue), like in “world energy use over” google search. This way you’d have 50% green-busy CPU, and the other 50% available. Unthrottling would reduce green to lesser values, and 100% would always remain constant in absolute value.
- he0001 5y agoI don’t think the article’s solution is a good one. The CPU metric is on its own really hard to understand and complex. I’d rather have the OS to understand and report when the CPU is throttled and do the reporting accordingly. So the metrics are easier to interpret.
- lrizzo 5y agoI think there are two different use cases: - measurement at thread level (top and the like) should report instructions retired per unit of time, not %CPU. We want to know how much work is being done. The explanation on why so much/so little requires more information (scheduler? cache misses? cpu throttling?) - measurement at the CPU level (mpstat, etc) should report both %CPU (as "time active" vs "wall clock time") and active clock cycles per unit of time (dividing by clock stretching factor if used, and perhaps scaled to absolute max frequency and/or %CPU if one wants a percentage). %CPU tells us whether we are making full use of the available time, and perhaps suggests to schedule threads differently if appropriate Clock cycles tell us how much one CPU is affected by throttling (manual, automatic, due to C-state exits, etc), and this is an orthogonal indication on whether there is something at the system level that is making the CPU underutilized
- Neil44 5y agoI disagree. Seeing 100% CPU but a low speed in windows task manager is a strong easy signal that you need to look at thermal issues. Taking away that information makes troubleshooting harder. It is also ‘truth’ in that moment and adjusting it to some other value based on what the cpu could do (but isnt doing right now) only serves to obfuscate imo.
- nprateem 5y agoWhat? An African or European CPU?
- MauranKilom 5y agoFun tangential anecdote regarding how interconnected and unintuitive CPU performance can be: I once made something run 20% faster by spawning a thread that did nothing but spin (i.e. while (true);). I was trying to optimize some FEM code, toying with (hardcoded) solver parameters. On one console I had it spitting out the wall clock durations of time steps as the simulation was running, while on the other I was preparing the next run. I start compiling another version, and inexplicably the simulation in the other console gets faster. Like, 10%-20% less time taken per time step. "That must have been coincidence. There's no way the simulation got faster by compiling something in parallel." But curiosity got the better of me and I still investigated. Watching the CPU speed with CPU-Z, it turned out that the simulation was indeed getting down-clocked, and that compiling something in parallel made the CPU run faster, speeding up the simulation too. WTF? And indeed, I could make the entire simulation run significantly faster by calling std::thread([](){ while (true); }); at the start of main. Why? Well, the simulation happens to be extremely memory-bound (sparse mat-vec multiplication in inner loop). So the CPU is mostly waiting around for data to arrive. Apparently the CPU downclocks as a result. That would be fine, if not for the fact that the uncore/memory subsystem clock speed is directly tied to the current CPU speed. That's right: The program was memory-bound, hence the CPU clocked down, hence the uncore clocked down, hence memory accesses became slower. Knowing that feedback loop, it makes perfect sense that keeping the CPU busy with a spinning thread improves performance. But it's still one big wtf. This problem eventually went away as we parallelized more and more of the simulation, giving the CPU less reason to clock down. But for related reasons, the simulation still runs faster if you prevent hyperthreading (either by disabling it in BIOS or having num threads = num hardware cores). More threads don't improve memory bandwidth and the hyperthread pairs just step on each others toes.
- omegalulw 5y agoI'm confused, how is COU speed tied to memory? AFAIK, memory is tied to CPU base speed which is almost always 100 MHz. The CPU then just scales it's own multiplier.
- MauranKilom 5y agoNorthbridge frequency (as shown by CPU-Z) is correlated to CPU speed in my experiments. It's not one to one, but NB frequency definitely varies by a factor of two depending on CPU load. What exact mechanism controls this is not clear to me (and I'm actually not sure if it's clear to anyone outside of Intel - the one paper [0] I found at the time was based on reverse engineering experiments). Nevertheless, CPU clock speed definitely affects Northbridge speed, as proven by the latter increasing from spinning a thread that never touches memory. [0]: https://tu-dresden.de/zih/forschung/ressourcen/dateien/projekte/firestarter/2015_hackenberg_hppac.pdf?lang=en https://tu-dresden.de/zih/forschung/ressourcen/dateien/proje... See section V.A: > The results [...] indicate that uncore frequencies – in addition to EPB and stall cycles – depend on the core frequency of the fastest active core on the system. (That conclusion is fully in line with my own observations.) Also see the corresponding patent linked in the paper: https://patents.google.com/patent/WO2013137862A1/en https://patents.google.com/patent/WO2013137862A1/en
- barrkel 5y agoThrottling should show up as a pseudo-process, "consuming" peak performance.
- bbrks 5y agoThis is how MacOS presents thermal throttling in the Activity Monitor - there's a visible "kernel_task" process taking up CPU. The downside is of course, that users see this process name chewing up 500%+ of CPU, google "kernel_task cpu", and find the commands to disable the thermal throttling to "fix" the issue! https://eclecticlight.co/2019/02/25/playing-with-fire-dealing-with-slow-hot-macs/ https://eclecticlight.co/2019/02/25/playing-with-fire-dealin...
- tgtweak 5y agoThe other consideration is that you want to encourage application developers to use those lower power states, and not make it look like their program is artificially abusing available resources.
- rasz 5y agoreporting 50% utilization while CPU is heavily throttling sure makes Microsoft strategic partner Intel happy, a total coincident I imagine
- notacoward 5y agoI remember having debates in 1990, when SMP UNIX was still a new thing, about whether "load average" should be scaled by the number of processors. Things have only gotten messier since then. As usual, Brendan Gregg has a good take. https://twitter.com/brendangregg/status/1411654427304333313 https://twitter.com/brendangregg/status/1411654427304333313 Personally, what I'd want to see is the proportion of max achievable IPC, and if that means getting used to numbers well below 100% (even when I'm doing everything right) then so be it. I can adjust my expectations and targets.
- adrr 5y agoShouldn’t there be another metric that indicates what the capacity of the CPU is running at? It would be a important metric to monitor to detect issues like thermal based throttling or how often the CPU is going is utilizing boost features.
- NGRhodes 5y agoAll that matters to me is what % CPU is used compared to what is the maximum available on demand to a user at a given point in time.