3 ms·
I played around with a go server to do some simple scaling numbers - looking at possibly using go to implement a large-number-of-idle-connections notification s
by jbert 14y ago
I played around with a go server to do some simple scaling numbers - looking at possibly using go to implement a large-number-of-idle-connections notification server.
I found the (good) result that I could spawn a new goroutine for each incoming connection with minimal (~4k) overhead. This is pretty much what you'd expect since a goro just needs a page for it's stack if it's doing no real work. I had something like 4 VMs each making ~30k conns (from one process) to the central go server with something like 120k conns.
I found one worrying oddity however. Resource usage would spike up on the server when I shut down my client connections (e.g. ctrl-C of a client proc with ~30k conns).
Reasoning about things a bit, I think this is due to the go runtime allocating an OS thread for each goro as it goes through the socket close() blocking call. I think it has to do this to maintain concurrency. So I end up with hundreds of OS threads (each only lives long enough to close(), but I'm doing a lot at the same time).
Can anyone comment:
- is this guess as to the problem likely to be correct?
- is this "thundering herd" a problem in practice?
- are there ways to avoid this? (Other than not using a goro-per-connection, which I think it the only idiomatic way to do it?)
My situation was artificial, but I could well imagine a case that losing, say a reverse proxy, could cause a large number of connections to suddenly want to close() and it would be a shame if that overwhelmed the server.
- aaronblohowiak 14y ago> I think this is due to the go runtime allocating an OS thread for each goro as it goes through the socket close() blocking call. I think it has to do this to maintain concurrency I highly doubt that it is creating a thread per goro on client disconnect. If you have a minimalish example of this, the golang mailing list would be very interested in working with you to identify what went wrong and create a patch if it is an issue with the Go implementation.
- evmar 14y agoIn case you didn't see it, I commented above with a link to a bug.
- tptacek 14y agoBlocking system calls spawn OS threads in Go, which can be cached and recycled for new goroutines. You don't see this if you code to pkg/net because it multiplexes i/o with a select/kqueue goroutine, but you'll see it right away if you code directly to the syscalls. Close isn't a blocking call, though.
- _stephan 14y agoClose is annotated as a blocking syscall in http://golang.org/src/pkg/syscall/syscall_linux.go#L810 http://golang.org/src/pkg/syscall/syscall_linux.go#L810 Is there a guarantee on Linux that close can't block? (I suppose it depends on the file type and the definition of blocking.)
- tptacek 14y agoOnly if SO_LINGER is set.
- evmar 14y agoIt seems so. https://code.google.com/p/go/issues/detail?id=4056 https://code.google.com/p/go/issues/detail?id=4056 An interesting point raised there is that if they instead used a limited thread pool for all goroutines to share when making OS calls you could produce deadlocks.
- mattgreenrocks 14y agoI'm curious what calls could induce deadlock. I figured everything used a nonblocking API internally and the runtime conferred blocking semantics.
- jbert 14y agoAs I understand it, the runtime has to clone a new thread for each sync call into the OS. That is the (only) mechanism by which a goro can perform such a sync operation without risking blocking the process. It's entirely possible under some workloads that those blocking OS calls require other goros to make progress (think pipes), so this could result in deadlock (although for some workloads it wouldn't I guess).
- kibwen 14y agoI'm groping blindly here, but from the references in that message to cgo I think it might have to do with calling into foreign code. Presumably Go's scheduler can't yield when you're inside C code (I'm not aware of any lightweight thread system that can achieve this) so you have to block there, which means you can no longer schedule other goroutines on that thread, and thus you have to spawn a new thread if you want to get any work done. But I'm just guessing here.