3 ms·
I genuinely had not heard of anyone actually using a spinlock in production code until I started using LMAX Disruptor a few years ago. I was always told that t
by tombert 21d ago
I genuinely had not heard of anyone actually using a spinlock in production code until I started using LMAX Disruptor a few years ago.
I was always told that they were an anti-pattern, and I think that generally that is a pretty good rule of thumb, but I guess like most stuff in CS: there are always exceptions to "good rules of thumb".
I still haven't actually explicitly written a spinlock for anything in production, but Disruptor has shown me that there are cases for it.
- sedatk 21d agoIt’s one of the secret ingredients to avoid a Big Kernel Lock™.
- BoingBoomTschak 21d agoI think Linus says it well: https://www.realworldtech.com/forum/?threadid=189711&curpostid=189723 https://www.realworldtech.com/forum/?threadid=189711&curpost...
- markus_zhang 21d agoThanks for sharing. Can someone tell me what does this paragraph mean? > Use a lock where you tell the system that you're waiting for the lock, and where the unlocking thread will let you know when it's done, so that the scheduler can actually work with you, instead of (randomly) working against you. I have “implemented” a sleep lock in xv6. Is it what he meant? What does the Linux scheduler “know” about it and will do differently? (Trying to figure out what does “work with you” mean) Thanks in advance.
- moregrist 21d agoIn simplest terms: if you don’t tell the kernel that you’re waiting, the scheduler assumes you aren’t and will wake you up and let you spin, to the detriment of other threads that aren’t waiting. If the OS knows that a thread is waiting for a lock, the scheduler will not bother to schedule it until the lock is available. In general, it’s tempting when you’re bound by lock latency to skip the syscall overhead of sleeping. But a lot of the time that’s a code smell that there are other inefficiencies in the system and you should rethink how you’re scheduling work.
- rcxdude 21d agoThe basic idea is that a lock should be something the OS is aware of, so that while a thread is blocked on a lock, the scheduler never tries to wake it at all, and when it is unlocked, the scheduler can wake up the thread that's waiting on it immediately. If the scheduler isn't aware of the lock, it'll just try to wake up the thread periodically, often just wasting CPU when it's still blocked or kept asleep when it could be running.
- BoingBoomTschak 21d agoI think he means that if the OS knows about locking, it can manage a list of "waiters" to quickly know which thread to wake (resp. let sleep) once (resp. before) the lock is released. A bit like what classic UNIX does with wchan, but between the kernel and userspace this time. Related: https://rdmsr.github.io/writing/turnstiles/ https://rdmsr.github.io/writing/turnstiles/
- pamcake 21d ago> Note that even OS kernels can have this issue - imagine what happens in virtualized environments with overcommitted physical CPU's scheduled by a hypervisor as virtual CPU's? Yeah - exactly. Don't do that. Or at least be aware of it, and have some virtualization-aware paravirtualized spinlock so that you can tell the hypervisor that "hey, don't do that to me right now, I'm in a critical region". I can't be the only one who learned this the hard way by cramming too many vCPUs onto too few physical cores and initially wondering where the high load and latencies came from.
- ignoramous 21d ago> had not heard of anyone actually using a spinlock in production code Go stdlib sync.Mutex uses spins: https://victoriametrics.com/blog/go-sync-mutex https://victoriametrics.com/blog/go-sync-mutex / https://archive.vn/BIb7F https://archive.vn/BIb7F
- jmgao 21d agoOptimistically spinning for a bit before falling back to futex or equivalent is very different from a spinlock.
- mathisfun123 21d agonot all architectures have atomic cas
- loeg 21d agoReal architectures you'd run more than a single thread on? Such as?
- tom_ 21d agoQuad core ARMv8-A, e.g., Nintendo Switch.
- loeg 21d agoARMv8-A has atomic CAS, unless this is some silly definition thing where it's "a system that has the behavior of atomic CAS but is named something else."
- monster_truck 21d agoI don't think that came until 8.1
- tom_ 21d agoMaybe that's the trap I'm falling into? From memory it doesn't have any actually atomic operations, only an ll/sc sort of mechanism - which I've never really thought of as an atomic operation? Though it's true you can end up with the same end result, in that you can just keep trying the operation until you accidentally do a read-modify-write that's ended up - well, "atomic" is a valid way to describe it. So maybe it is good enough to count, though personally I'm still not quite convinced.
- loeg 20d agoYeah. I would call ll/sc atomic. Wikipedia's current verbiage: > Load-link returns the current value of a memory location, while a subsequent store-conditional to the same memory location will store a new value only if no updates have occurred to that location since the load-link. Together, this implements a lock-free, atomic, read–modify–write operation. https://en.wikipedia.org/wiki/Load-link/store-conditional https://en.wikipedia.org/wiki/Load-link/store-conditional
- bob1029 21d agoTo be really pedantic, it's a spin wait, not a spin lock in disruptor. You are waiting for a sequence, not mutually excluding some resource. Many threads can watch the same volatile at the same time without blocking each other.
- nly 21d agoIf you have an application where your threads are pinned to dedicated cores, and those cores are all isolated from general OS scheduling, then it's the lowest latency means to synchronize arbitrary things between threads Entering the kernel with a futex wait or wake under contention costs a couple of microseconds, whereas a spinlock will cost you double digit to low triple digit nanos depending on cores/sockets etc
- kazinator 21d agoBefore we had futexes in the Linux kernel, spinlocks were used to boostrap the implementation of everything else in the user space threading library. If you have futexes you can try to grab a lock with an atomic operation and if that fails, go wait on the futex via system call, so there is no need to spin. Spinlocks then remain useful as an optimization, because there are situations in which it is cheaper to spin around a bunch of times until the thread on another processor gives up the lock, than to take a trip into the kernel. You can also spin, but with a scheduler yield in the loop; we don't normally think of that as a spinlock. That's what you fall back on after spinning some number of times and failing to get the lock. In the Linux kernel, spinlocks are the low level primitive. They are very efficient because unlike user space threading, they are not faced with guesswork about scheduling. They are "surgical".
- BobbyTables2 21d agoTell a kernel developer that spin locks aren’t for production code. Bring a wind turbine with you because the laughing will be quite intense…
- ChickeNES 21d agoHeh, I lost it myself when reading it, so you're not wrong.
- tombert 21d agoTotally fair. I haven’t done much kernel stuff (outside of very basic toy stuff for QEMU). At the level I work (which is generally server/distributed stuff), I have always used mutexes that are built into the platform. Or more realistically, if I am the one writing the code, I just avoid mutexes and make my code ridiculously convoluted to do so.
- yxhuvud 21d agoVery different situation as stuff inside the kernel presumably cannot be preeempted at any time. Preemption does a number on spinlocks.
- loeg 21d agoKernel can prevent preemption (mostly). Userspace can not (mostly).
- foldr 21d agoOne use case I’ve found is for a lock that you don’t need to acquire. For example, you need a lock to read a cache entry, but if you can’t acquire the lock after a few spins, you can just proceed without the cache. For fine-grained locking, a spin lock can have a significantly lower memory overhead than a full futex.