4 ms·
You could easily find examples of this - for example in any textbook threaded socket server design (thread pool contending on accept())
by _wmd 8y ago
You could easily find examples of this - for example in any textbook threaded socket server design (thread pool contending on accept())
- lallysingh 8y agoYeah, those textbooks aren't talking about good implementations that scale up to that many cores.
- spacenick88 8y agoBut wouldn't that use a non spinning OS lock. I.e. the scheduler knows which threads are waiting on the accept and can randomly pick one to directly wake up?
- mickronome 8y agoIf your design have multiple threads waiting for accept on the same socket, then leaving it to the OS to schedule said threads on accept should at least be the first to try. If nothing else because it's the least surprising thing to do.
- smaddox 8y agoDoes textbook equate to real-world, though? Don't real world applications either use a thread per socket (i.e. per TCP stream) or something like epoll? Genuine questions, because I really don't know.
- the8472 8y agoReal world implementations would use OS-specific socketopts to create multiple listen sockets sharing the same port and then use one socket per thread.
- _wmd 8y agoYou're among the 1% of engineers who both care enough to know and actually know about these kinds of facilities. For the rest of them, whatever they got at school or off Stack Overflow is as far as it goes, unless some business priority demands something different - and even in that case "throw more hardware at it" is often the route taken
- gpderetta 8y agoThe rest of them is probably, hopefully, using some library abstraction that implement this optimization.
- blattimwind 8y agoI dunno about the BSDs but SO_REUSEPORT is a fairly recent addition. Was there a way to achieve this before then?
- dragontamer 8y agoThat's not a spinlock. The "Socket textbook example" is either a semaphore or a mutex. The textbook example of a spinlock is maybe a synchronized counter (iterations++), or maybe a producer/consumer queue.
- _wmd 8y agoThe accept call on Linux: - invokes the accept() function of the socket family: https://github.com/torvalds/linux/blob/be779f03d563981c65cc7417cc5e0dbbc5b89d30/net/socket.c#L1635 https://github.com/torvalds/linux/blob/be779f03d563981c65cc7... - which invokes the accept() function of the protocol: https://github.com/torvalds/linux/blob/be779f03d563981c65cc7417cc5e0dbbc5b89d30/net/ipv4/inet_connection_sock.c#L425 https://github.com/torvalds/linux/blob/be779f03d563981c65cc7... - which invokes lock_sock_nested(), which spins https://github.com/torvalds/linux/blob/be779f03d563981c65cc7417cc5e0dbbc5b89d30/net/core/sock.c#L2847 https://github.com/torvalds/linux/blob/be779f03d563981c65cc7...
- dragontamer 8y agoThe accept call sleeps if no sockets are available. The thread goes to sleep entirely and stops executing on the core completely. Yes, there are spin-locks involved in the process. But I'm talking about these lines of code: > error = inet_csk_wait_for_connect(sk, timeo); > mutex_acquire(&sk->sk_lock.dep_map, subclass, 0, _RET_IP_); Etc. etc. You know, the stuff that causes milliseconds to multiple-seconds worth of delay, as opposed to "pause / spinlocks" which is measured in nanoseconds.
- _wmd 8y agoThe parent context asked what code could have 40 cores spinning, I picked the first example I could think of. You said it wasn't a spinlock, it was. Sure there is a followup lock that sleeps for the empty-queue case, but that doesn't relate to OP's question You can find similar behaviour in the filesystem APIs, anything touching the VM (mmap_sem IIRC is also a spinlock), pretty much any OS API where threads are all banging at the same shared resource that isn't expected to require a long wait. struct file also contains a spinlock, but doesn't look like it's used in normal operation