10 ms·
63 Cores Blocked by Seven Instructions
- vkaku 7y agoIt's good to see how features turned on by default (System Restore) can have such a bad impact on performance. Thank you for doing the profiling!
- acdha 7y agoSystem Restore is commonly disabled for test systems or anyone who has a good automated deployment system. Now I’m wondering how likely it is that many engineers at Microsoft disabled it to save space or conserve every available IOP, especially in the era before large SSDs were widely available.
- pm7 7y agoDo you have any good list of such actions?
- userbinator 7y agoWindows has never been about absolute best performance, but "good enough"; which sometimes isn't. Otherwise a default install wouldn't have so much stuff running.
- snak 7y agoThat was a good read. In-depth but understandable. Thanks for sharing.
- mehrdadn 7y agoEdit: Never mind... I completely missed the word "empty" when reading the critical sentence. :(
- wjnc 7y agoAs per the article: "It is unclear why this code misbehaved so badly on this particular machine. I assume that it is something to do with the layout of the almost-empty 2 TB disk."
- saagarjha 7y agoWhy do the sample counts cluster so heavily on the jne, as opposed to the other instructions in the loop?
- jclulow 7y agoThe article links to another article which discusses this. In order to do stack sampling, the CPU doing the work must be periodically interrupted for the sampler to run and collect a stack. Modern CPUs are deeply speculative and often issue instructions out of strict program order. When you interrupt such a CPU it has to walk some of that back and decide where to leave off on all of the speculative execution; what the CPU was doing at the time is thus not completely captured by the stack trace alone.
- brucedawson 7y agoI talked to the author of that article. He hasn't done testing on AMD processors but his guess was: micro-up fusion means the seven-instruction loop is actually five micro-ops Zen2 processors can retire five instructions per cycle Therefore the loop runs at one iteration per cycle (wow!) The cmp [r8] instruction occasionally has cache misses This means that the seven instructions get synchronized such that the cmp [r8] instruction is the last of them to get retired in a seven-instruction block Therefore the next instruction is usually the jne TL;DR - the jne gets most of the samples because the cmp [r8] instruction is the most expensive.
- BeeOnRope 7y agoYeah I believe jne gets most of the samples because cmp [r8] is the most expensive, but there could be two separate reasons for that: Perhaps ETW shows you the precise instruction (i.e., "zero skid") that is slow to retire - this is not like a normal interrupt as described in the article but is available with some performance profiling events like 'cycles:ppp' on Linux perf (in particular, using the zero-skid PEBS events). In that case, the samples shows up on the jne, not the cmp likely because cmp/jne have fused, so basically get sampled as a single instruction and the samples show up pointing to the jump. The other scenario is that ETW shows you "skid 1" instructions, i.e., the instructions generally after the slow-to-retire ones (as described in the article), and cmp/jne didn't fuse (perhaps because a cmp with a memory source argument can't fuse on AMD?), and so it again points to the jne. I haven't looked at many ETW traces, so I couldn't tell you offhand - but for those who have, do the samples usually show on the expensive instructions (things like div and loads that miss are a giveaway), or on the one after? Added: Per Agner, I guess the "fusion + no skid" is the most likely (from the Ryzen section of microarchitecture.pdf): > A CMP or TEST instruction immediately followed by a conditional jump can be fused into a single μop. This applies to all versions of the CMP and TEST instructions and all conditional jumps, except if the CMP or TEST instruction has a rip-relative address or both a displace- ment and an immediate operand. That also lines up with the cmp having exactly 0 samples, unlike any other instruction of the 7: that's a common indication of fusion.
- KenanSulayman 7y agoTechnically this wasn't caused by those instructions but by the spinlocks waiting for the lock to be released. Also "blocked by seven instructions" sounds a bit click-baity.. you can lock the CPU or power off the computer with less than that amount of instructions :-)
- saagarjha 7y agoOr break it, depending on how old it is: f0 0f c7 c8: lock cmpxchg8b eax
- ncmncm 7y agoThe cause is obvious: they were building on Microsoft Windows, using the NTFS filesystem. Even Microsoft doesn't try to build on NTFS. Changing any single detail gives better results. Use a Samba share from a Linux filesystem. Run Mingw on a Linux system. Run MSVS in Wine on a Linux system. Windows is an execution environment for applications. There is no need for, and no value in, actually performing builds in your target execution environment. Use a system designed from the ground up for builds.
- youdontknowtho 7y agoThat's the first I've heard that MS doesn't try to build Windows on Windows with NTFS.
- dijit 7y agoI don't have first hand experience, but I know some people at MS who work on Xbox (which is a modified version of Windows+HyperV underneath); From what I understood from them, they do not use NTFS (they use SMB from a clustered filesystem) for builders, but they _do_ use a heavily modified version of windows, incidentally that modified version went on to become "windows nano". What the actual "Windows" team does is a mystery to me though, I would assume it was similar or the same.
- youdontknowtho 7y agoThat's super interesting. I wonder if they have moved to use SMBv3? I really liked the direction with Nano in 2016, but I guess it makes more sense as a container OS. Still, the latest version is what a lot of people wish they could start an operating system with. NT kernel, no wmi, no servicing, no activation.
- dooglius 7y agoAs a developer, you do some builds on your local box, so it's really up to you what filesystem you use.
- pitterpatter 7y ago
- markdog12 7y agoBruce Dawson does some of Microsoft's most valuable work for Windows. Doesn't even work for them.
- KIFulgore 7y agoHe worked for Microsoft years ago. Still an MVP all this time later.
- Darkphibre 7y agoI worked with the man for years in the Xbox Advanced Technology Group. Amazing individual. When he left the team, I conducted my own exit interview so I could learn from him, and walked away with pages and pages of insights on growing my own career and becoming a subject matter expert. He was on my interview loop at ATG, and I recount it as my favorite interview of all time. He pointed to a circuit diagram poster, and said to me "You have to write a game for that, what design considerations should you be aware of?" It looked something like this (can't find the actual poster, it's been a decade): https://qph.fs.quoracdn.net/main-qimg-9cdbc7bf35ef8126755175c99410013c https://qph.fs.quoracdn.net/main-qimg-9cdbc7bf35ef8126755175... A bit out of my league, but I identified the important aspects (multicore/hyperthreaded design, small L0/L1 cache and impacts to mispedictions, etc.) and spoke to what I could and where my uncertainties lay. Afterwards he gave up the rest of the time to let me ask questions about the team. One XFest he stood on stage giving a Powerpoint presentation on debugging and multithreaded concerns. An animation was slow, and he broke into it and started debugging Powperpoint live to demonstrate some of his techniques. A legend. A huge loss to Microsoft when we stepped away. I did and do hope him the best!
- anniely 7y agoWow. This was such a valuable post. Can you describe some of the insights you learned from the exit interview -- about growing your career and becoming a subject matter expert? I'm new to the field and I feel like that'd be immensely valuable to me and many others.
- Darkphibre 7y ago
- alexeiz 7y ago"loop running in the system process while holding a vital NTFS lock" It's not about the seven instructions. It's the lock that's been held while doing a busy loop.
- pbsd 7y agoFor each input source file, cl.exe creates at least 7 temporary files (with suffixes "gl", "sy", "ex", "in", "db", "md", "lk"). The churn of creating and deleting those, coupled with the slowness of performing checkpointing on a huge empty drive, seem to be the root cause here. This appears somewhat related to this bug report: https://developercommunity.visualstudio.com/content/problem/310131/clexe-creates-so-many-temp-files-it-freezes-the-sy.html https://developercommunity.visualstudio.com/content/problem/... Marking the temporary files as FILE_ATTRIBUTE_TEMPORARY could improve things, without having to go into significant Windows kernel changes.
- snagglegaggle 7y ago> Marking the temporary files as FILE_ATTRIBUTE_TEMPORARY could improve things, without having to go into significant Windows kernel changes. Having used Cygwin and MinGW (and less so WSL) NTFS is probably the main factor here. Especially when using Cygwin program compilation is very slow not just due to process creation but file access, potentially on disk. I could see checkpointing contributing, but having used backups/file history on Windows server, you will see CPU use irregularities that I think are similar in cause to this but you do not see CPU use irregularities as bad as described in the article. The pathological behavior of NTFS with many files is easy to prove once you encounter it.[0] This would be exposed to the kernel as well and is likely holding up NtfsCheckpointVolume. At least in my experience the problem goes deeper, and how NTFS structures are handled also contributes; for example, trying to enumerate and copy files in certain ways is extremely slow even if the files are easily enumerable. You can say that if disabling checkpointing gets rid of noticeable slowdown it is the right thing to do, but there are people who will rightly not want to disable it. [0] https://stackoverflow.com/questions/197162/ntfs-performance-and-large-volumes-of-files-and-directories https://stackoverflow.com/questions/197162/ntfs-performance-...
- pault 7y agoTime the difference between an npm install, with all its thousands of tiny files, on NTFS and ext4. It's excruciating.
- pizzazzaro 7y agoIs ninja python or C-based? I cant remember. Is "-pipe" in their c-compiler's make.conf ? Would that even matter when using ninja as the compiler? Im curious, and trying to think towards a solution.
- peter_d_sherman 7y agoExcerpt: "...I mean, how often do you have one thread spinning for several seconds in a seven-instruction loop while holding a lock that stops sixty-three other processors from running. That’s just awesome, in a horrible sort of way." I respectfully disagree. That's because everything in the universe that is percieved as negative -- turns out to have a positive use-case somewhere, sometime, in some context... In this case, I think the ability for one core to stop 63 other processor cores is purely awesome, because think of the possible use-cases! Debugger comes to mind immediately, but how about another if let's say there are 63 nasty self-resurrecting virus threads running on my PC? What about if you were doing some kind of esoteric OS testing where you needed to return to something like Unix's runlevel 1 (single user), but you'd rather freeze most of the machine (rather than destroying the context of everything else that was previously running?). Oh, here's the best one I can think of -- don't just do a postmortem, everything's dead core dump when something fails -- do a full (frozen!) "live" dump of a system that can be replayed infinitely, from that state! Now, because I take a contradictory position, doesn't mean we're not friends, or that I don't acknowledge your technical brilliance! Your article was absolutely great, and you are absolutely correct that for your use-case, "That’s just awesome, in a horrible sort of way.". But for my use-cases, it's absolutely awsome, in the most awesome sort of way! <g>
- lilyball 7y ago63 cores blocked on a single mutex is not at all like any of the scenarios you're describing. That's almost like describing the notre dame fire as having a positive use-case because what if you want to do a controlled demolition of a large building.
- paulddraper 7y agoI suppose it's awesome for that capability to exist, though it wasn't even close to what should have happened here. Awesomely horrible.
- Dylan16807 7y agoMaking all your processes wait on a lock is multithreading 101. The horrible part is the specific way this lock is getting held, which is not useful. And there are simpler ways to prevent all access to a drive.
- strictfp 7y agoSo, one busy process performs a file operation that triggers a system restore checkpoint, and the OS locks the entire drive during this file operation? Sounds strange to me. Is the problem that the checkpointing critical section has the same duration as the triggering file operation? I get that there must be some sort of critical section for setting a checkpoint, but I don't understand why it takes so long, and why it would be affected by how busy the userspace process that triggered it is. I would expect it to have a short barrier-style critical section; drain all outstanding writes, record some checksum or counter from a kernel data structure, and then release all writers again. In my mind this should be kernel code only, entirely unaffected by userspace, and if designed nicely, quite fast. So I guess I don't get what is going on here.
- kevingadd 7y agoMy guess would be that the system restore checkpointing functionality ends up holding a lock by necessity while it manipulates internal state. It would be really hard to make that sort of algorithm lock-free and preserve data integrity guarantees - I certainly wouldn't want to be responsible for writing it. Obviously the lock shouldn't be held so long and so often though...
- brucedawson 7y agoMy understanding is that the system restore checkpoints happen every five seconds. They hold a lock, which seems reasonable. The problem is that for some reason on this machine the checkpoint process was taking a really long time. I also don't understand why it was taking so long. It normally doesn't. Something went terribly wrong. > and if designed nicely, quite fast. Yep, should be. But it wasn't. If everything worked as it should then I'd never get to write any blog posts!
- jeffdavis 7y agoIt looks like this is a case where a process is holding lock A while waiting on lock B; and every other process is waiting on lock A. That's normal enough, though it seems like there are two mistakes: First: Never spin waiting on a lock for 3 seconds. If you expect a lock to be released very quickly, you spin K times and then, if you still don't have the lock, try something heavier that can deschedule your process. K should be small enough that your time slice is unlikely to expire while spinning, otherwise, it just causes confusion and wasted work because it looks like your process is doing work when it's not. Second: It seems dubious that using a feature like system restore causes all Write calls to wait for a lock held by a process in the middle of I/O. I'm sure there are some cases where that must happen (like if out of buffer space to hold the writes), but I would think it would be harder to hit. EDIT: Rephrased my comment in terms of two problems rather than just the first one.
- caf 7y agoLook again - that tight loop in RtlFindNextForwardRunClear isn't spinning on a lock - it's scanning forward through memory, 4 bytes at a time, looking for 4 bytes not equal to the pattern in %ebp. So it looks more like "process is holding lock A while doing a very long scan through memory". That would fit with the name of the function, too.
- brucedawson 7y agoCorrect. Nobody was spinning on a lock. Everybody was waiting politely. The problem was that the system process held the lock for too long, due to some inefficiency in system restore (root cause not yet understood by me).
- brucedawson 7y ago> Second: It seems dubious that using a feature like system > restore causes all Write calls to wait for a lock I agree that it seems dubious, but it is indisputably what was happening, repeatedly.
- Syzygies 7y agoSo when did you first realize he was discussing Windows, reading this? The "of course everyone is a straight white male" attitude that the OS need not be stated, so often seen in Windows posts, gave it away for me. However, my biases threw me for way too long: the level of sophistication meant this must be Linux, right? I should have recognized the graphics style in the screen grabs. Certainly not MacOS, but Linux can be all over the map stylistically. Does Windows really still look like that? Wow.
- ncmncm 7y agoI caught on when I realized this just wouldn't be happening anywhere else. Microsoft has, singly-handedly, got two generations of people used to computers working badly, convinced that it's not just unavoidable, but normal. If cars worked as badly, we would all see multiple explosions every day (and think it was awesome).
- CawCawCaw 7y agoThese posts by Dawson are always interesting. Now, if only he would investigate and remediate the performance deficiencies of other complex systems, such as ... Chrome?