4 ms·
Thanks! Would love to hear more about the counters that your interested in. We've exposed more in C5 than in previous instance types and we are trying to make
by aliguori 9y ago
Thanks! Would love to hear more about the counters that your interested in. We've exposed more in C5 than in previous instance types and we are trying to make more available over time in a safe way.
- KenoFischer 9y agoI have two use cases: - General performance analysis. For this more counters is generally incrementally better. - Running https://github.com/mozilla/rr https://github.com/mozilla/rr. This requires the retired-branch-counter to be available (and accurate - sometimes virtualization messes that up) The second one I actually care more about, because I've pretty much stopped trying to debug software when rr is not available, too painful ;). Feel free to email me (email is in my profile) for gory details.
- khuey 9y agoFor the benefit of anyone reading this, KVM and VMWare virtualization generally work. Xen has problems because of a stupid Xen workaround for a stupid Intel hardware bug from a decade ago. I can provide more details about that via email (in my profile) if desired.
- paulie_a 9y agoCan you please just post the info. Intel deserves to be shamed
- voidmain0001 9y agoIs this what khuey is referring to?: https://support.citrix.com/article/CTX136003 https://support.citrix.com/article/CTX136003
- khuey 9y agoOne of the things the performance monitoring unit (PMU) is capable of doing is triggering an interrupt (the PMI) when a counter overflows. When combined with the ability to write to the counters, this lets you program the PMU to interrupt after a certain number of counted events. Nehalem supposedly had a bug where the PMI fires not on overflow but instead whenever the counter is zero. Xen added a workaround to set the value to 1 whenever it would instead be 0. Later this was observed on microarchitectures other than Nehalem and Xen broadened the workaround to run on every x86 CPU. Intel never provided any help in narrowing it down and there don't seem to be official errata for this behavior too. This behavior is ok for statistically profiling frequent events but if you depend on exact counts (as rr does) or are profiling infrequent events it can mess up your day. https://lists.xen.org/archives/html/xen-devel/2017-07/msg02242.html https://lists.xen.org/archives/html/xen-devel/2017-07/msg022... goes a little deeper and has citations.
- irishcoffee 9y agoSeconding paulie_a, We're running a Xen stack right now and I haven't heard of this. We've worked around a few nasty bugs with Xen and linux doms already, but I'm wondering if we have this problem you're referring to and don't even know it.
- DSingularity 9y agoIm assuming rr is only unavailable for multithreaded apps? How frequently is rr available for your use?
- khuey 9y agorr works fine on multithreaded (and multiprocess) applications. It does emulate a single core machine though, so depending on your workload and how much parallelism your application actually has it might be painful.