7 ms·
Goroutines Under the Hood (2020)
- geodel 4y agoI mean it has many diagrams and logical explanation of Goroutines and concurrency concepts in general but it is definitely not under the hood descriptions.
- geertj 4y agoI came here to write the same thing. Things I had hoped this would go into is how goroutines grow their stacks, and how they are preempted.
- ignoramous 4y agoThere you go: Vicki Niu, Goroutines: Under the hood, https://www.youtube-nocookie.com/embed/S-MaTH8WpOM https://www.youtube-nocookie.com/embed/S-MaTH8WpOM (2020).
- deleted 4y ago[deleted]
- aaronbwebber 4y agoI love Go and goroutines, but... > A newly minted goroutine is given a few kilobytes a line later > It is practical to create hundreds of thousands of goroutines in the same address space So it's not practical to create 100s of Ks of goroutines - it's possible, sure, but because you incur GBs of memory overhead if you are actually creating that many goroutines means that for any practical problem you are going to want to stick to a few thousand goroutines. I can almost guarantee you that you have something better to do with those GBs of memory than store goroutine stacks. Asking the scheduler to handle scheduling 100s of Ks of goroutines is also not a great idea in my experience either.
- uluyol 4y agoWhy is spending GB on stack space a bad thing? Ultimately, in a server, you need to store state for each request. Whether that's on the stack or heap, it's still memory that necessarily has to be used.
- guenthert 4y agoDespite popular belief, not everything is a (web) server. I can imagine many threads to be appealing in e.g. simulations.
- uluyol 4y agoSure, but my point is that if you want to run 100k-1m things concurrently, you need to store state for them. That has a memory cost no matter what.
- astrange 4y agoI definitely can’t hire anyone in this thread to work on cell phone performance. We fight for 10 KBs of memory and yes, we are still doing this in 2022. Even on a server, you may have TBs of RAM but you don’t have that much L1 cache nor that much memory bandwidth.
- uluyol 4y agoWhy would you need hundreds of thousands or millions of goroutines for a cell phone app/daemon? I would expect the number (and corresponding memory usage) to therefore be low.
- astrange 4y agoThe only numbers in programming are one, two, and many. So if you’re not very careful and nothing stops you, it’s pretty easy to create an unbounded amount of anything.
- 4y ago
- captainmuon 4y agoOne thing that really goes against my intuition is that user space threads (lightweight treads, goroutines) are faster than kernel threads. Without knowing too much assembly, I would assume any modern processor would make a context switch a one instruction affair. Interrupt -> small scheduler code picks the thread to run -> LOAD THREAD instruction and the processor swaps in all the registers and the instruction pointer. You probably can't beat that in user space, especially if you want to preempt threads yourself. You'd have to check after every step, or profile your own process or something like that. And indeed, Go's scheduler is cooperative. But then, why can't you get the performance of Goroutines with OS threads? Is it just because of legacy issues? Or does it only work with cooperative threading, which requires language support? One thing I'm missing from that article is how the cooperativeness is implemented. I think in Go (and in Java's Project Loom), you have "normal code", but then deep down in network and IO functions, you have magic "yield" instructions. So all the layers above can pretend they are running on regular threads, and you avoid the "colored function problem", but you get runtime behavior similar to coroutines. Which only works if really every blocking IO is modified to include yielding behavior. If you call a blocking OS function, I assume something bad will happen.
- gpderetta 4y agoJust the interrupt itself is going to cost a couple of order of magnitude more than the whole userspace context switch.
- garaetjjte 4y agoOne of the issues is that OS schedulers are complex, and actually much more expensive than context switch itself. You can mitigate this with user-mode scheduling of kernel threads: https://www.youtube.com/watch?v=KXuZi9aeGTw https://www.youtube.com/watch?v=KXuZi9aeGTw
- masklinn 4y ago> And indeed, Go's scheduler is cooperative. It hasn't been cooperative for a few versions now, the scheduler became preemptive in 1.14. And before that there were yield points at every function prolog (as well as all IO primitives) so there were relatively few situations where cooperation was necessary. > Without knowing too much assembly, I would assume any modern processor would make a context switch a one instruction affair. Any context switch (to the kernel) is expensive, and way more than a single operation. The kernel also has a ton of stuff to do, it's not just "picks the thread to run", you have to restore the ip and sp, but also may have to restore FPU/SSE/AVX state (AVX512 is over 2KB of state), traps state. Kernel-level context switching costs on the order of 10x what userland context switching does: https://eli.thegreenplace.net/2018/measuring-context-switching-and-memory-overheads-for-linux-threads/ https://eli.thegreenplace.net/2018/measuring-context-switchi... > LOAD THREAD There is no load thread instruction
- jeffbee 4y agoIt didn't really lift the hood at all, unfortunately. Luckily for us the runtime is extensively commented, e.g. https://github.com/golang/go/blob/master/src/runtime/proc.go https://github.com/golang/go/blob/master/src/runtime/proc.go
- yellow_lead 4y ago(2020) But so high level it's still relevant.
- qwertox 4y agoI've fallen in love with Python's asyncio for some time now, but I know that go has coroutines integrated as a first class citizen. This article (which I have not read but just skimmed) made me search for a simple example, and I landed at "A Tour of Go - Goroutines"[0] That is one of the cleanest examples I've ever seen on this topic, and it shows just how well integrated they are in the language. [0] https://go.dev/tour/concurrency/1 https://go.dev/tour/concurrency/1
- throwaway894345 4y agoHaving used both in production for many years, Go’s model is waayyyy better, mostly because Python’s model results in a bunch of runtime bugs. The least of which are type error things like forgetting to await an async function—these can be caught with a type checker (although this means you need to have a type checker running in your CI and annotations for all of your dependencies). The most serious are the ones where someone calls a sync or CPU-heavy function (directly or transitively) and it starves the event loop causing timeouts in unrelated endpoints and eventually bringing down the entire application (load shedding can help mitigate this somewhat). Go dodges these problems by not having sync functions at all (everything is async under the covers) and parallelism means CPU-bound workloads don’t block the whole event loop.
- zarzavat 4y agoI’m not a Go programmer, doesn’t having everything async make your code riddled with race conditions? It seems like it would make ordering very hard to reason about.
- throwaway894345 4y agoNo, only mutable shared state is subject to race conditions. Mutable shared state is rare, and you can use a mutex to guarantee exclusive access to it. Note also that only the implementation is async, but the programmer interface is synchronous (in other words, the programmer doesn’t need to type “await” all over the place).
- rounakdatta 4y agoWhile the blog is a great introductory post, https://www.youtube.com/watch?v=KBZlN0izeiY https://www.youtube.com/watch?v=KBZlN0izeiY is a great watch if you're interested in the magical optimizations in goroutine scheduling.
- bogomipz 4y agoIn the conclusion the author states: >"Go run-time scheduler multiplexes goroutines onto threads and when a thread blocks, the run-time moves the blocked goroutines to another runnable kernel thread to achieve the highest efficiency possible." Why would the Go run-time move the blocked goroutines to another runnable kernel thread? If it is currently blocked it won't be schedulable regardless no?
- kangda123 4y agoI haven't read the article but, generally, main thread pool is designed to utilise the whole processor. So on a 12-core system there will be 12 OS threads. No point in having more. But, if the program opens just 12 big files all threads become blocked and that's obviously tragic: other routines are starved, network sessions time out, timers don't run - chaos! Thus, whenever a thread in the main pool is blocked, you move the thread outside and spawn a fresh new one that can keep going through the compute. There's also a subtler benefit. Each user thread has a context, e.g. its local run queue. Now, if thread blocks, others need to help it out and steal its work. Go improves that by having a nice handover system, so no random stealing is necessary. Further, by taking context off blocked threads, it keeps all the tasks more centralised. There is probably at most a few tens of processors on any common hardware but there could be thousands threads. It's better to tie runqueues to the former rather than the latter.
- hoosieree 4y agoThere's no reason to move the blocked task. But any other tasks queued up behind the blocked task could be moved to another OS thread where they'll have a chance of running.
- bogomipz 4y agoYeah I thought it was an odd statement to make especially as part of the conclusion.
- guenthert 4y agoSo it's just coroutines on top of n:m scheduling, similar to what SysV offered a while ago?
- klodolph 4y agoSysV? Are you thinking Solaris? There have been various implementations of M:N threads for some time now. The concept is simple, but the devil is in the details.
- deleted 4y ago[deleted]