7 ms·
The title is misleading — no code is broken, though it is true that .NET programs that exhibit lock contention will be slower. The root of the problem is that
by jstarks 8y ago
The title is misleading — no code is broken, though it is true that .NET programs that exhibit lock contention will be slower.
The root of the problem is that .NET implemented a rather baroque spin loop with several assumptions about the CPU and OS architecture. Unlike any other spin loop I have seen or written, they execute multiple pause instructions per loop. Arguably this was a poor choice.
It would be interesting to see the history of this design. It was probably done years ago when .NET was closed source.
- _wmd 8y agoThe stuff he quoted seems to make a strong case that the behaviour they implement was the intended behaviour prior to Skylake, why else specifically call it out in the newer manuals? Would be interested to know why you think the code was always broken. Perhaps Intel define a preferred algorithm they failed to implement?
- ptero 8y agoNot the parent, but I do not see that Intel is wrong here. New architecture makes new choices and, according to the redlined doc, they expect 1-2% performance improvements on threaded apps, which is a measurable improvement. They also explicitly call for possible degradation cases so folks who care can adjust. .NET apparently hits such a degradation case, but parent says this is due to .NET implementing spin loops in a way (almost) no one else does. If so, in this particular case Intel is doing the greatest good for the greatest number, which is reasonable. My 2c.
- arcticbull 8y agoAnd of course it's really easy for .Net to fix it for everyone in one simple patch, which it looks like they already have. I'm with you on this one.
- eloff 8y agoMy experience is to the contrary. Every spin lock I have seen uses more than one pause instruction. Some may start with one and then execute more. They all involve a loop over pause instruction(s). I don't think that's the problem per se. The problem is more likely the back-off logic.
- cesarb 8y agoAt least the Linux kernel spinlocks do a single pause instruction (which it calls cpu_relax()) before checking again to see whether the lock is free; it's a loop over the pause instruction, but the loop condition checks the lock value. I'd expect most spinlock implementations to do the same.
- eloff 8y agoThey basically all do that. I've seen some with multiple pause instructions in the loop body, and others with a switch or if/else if/else block where it starts off with one pause instruction, then goes to multiple, then moves to more drastic back-off techniques like giving up the thread time slice or micro sleeping.
- acqq 8y ago> I've seen some Then please quote at least one actually used in Linux kernel, including where it is used.
- __s 8y agoThere exist spink locks outside of the Linux kernel
- cesarb 8y agoYes, but the Linux kernel is the one I'm more familiar with. So I looked for another example: glibc's pthread_mutex_lock. The code is at nptl/pthread_mutex_lock.c and the "pause" instruction here is called atomic_spin_nop() (or BUSY_WAIT_NOP on older versions), defined by sysdeps/x86_64/atomic-machine.h. Yet again, it does a single "pause" instruction in the loop, and the loop condition tests the lock.
- gameswithgo 8y ago>with several assumptions about the CPU and OS architecture You aren't going to get a good spin lock without making some assumptions about the hardware and OS. Get new hardware, gotta add a new case.
- Matthias247 8y agoI think the issue is not only the .NET spin loop implementation, but also that .NET in general is heavy on shared memory concurrency and required locking. Examples: All the Concurrency primitives like Task<T> are thread-safe by default, and that has some associated cost. Other runtimes like javascript or other singlethreaded event loops don't have to pay for that cost, and might still be able to scale horizontally. Also the IO subsystem on .NET requires a lot of synchronization: Most IO callbacks (e.g. for sockets) will be made from an arbitrary Threadpool thread, which requires synchronization or marshaling to get back to another context. Again that will be cheaper on a runtime where the whole event sources and consumers are using the same thread. I certainly can't quantify it in numbers, but I guess .NET has some of the highest requirements regarding synchronization performance. At least the standard libs, there are obviously some efforts underway to improve that situation.
- runfaster2000 8y agoMuch of the context is in dotnet/coreclr #13388 [1]. At the risk of repeating the content of the blog post ... The issue was reported in August, 2017 and fixed later in the same month. The fix shipped in a .NET Core patch update in September. We attempted to get the fix into .NET Framework 4.7.2 but missed that scheduling. We're now looking at making the fix in a .NET Framework patch release. [1] https://github.com/dotnet/coreclr/issues/13388 https://github.com/dotnet/coreclr/issues/13388