5 ms·
Why are you parallelizing these in such different ways? For example, in Go you start 800 goroutines, but in others you start just 4 threads/tasks as worker pool
by james4k 13y ago
Why are you parallelizing these in such different ways? For example, in Go you start 800 goroutines, but in others you start just 4 threads/tasks as worker pool. I would imagine indeed these would give quite different results.
- logicchains 13y agoIt's done differently in Go because someone else wrote it, and I found it to be faster than the version I wrote using a worker pool of 4 tasks. If you look closely you'll see it doesn't actually ever run more than 4 goroutines simultaneously.
- james4k 13y agoI noticed that, but you're still allocating for 800 goroutines in the end. I guess in the scope of the benchmark, that is still relatively cheap. It just was a red flag for me.
- logicchains 13y agoNo worries. A commenter here submitted a faster version, so it uses that now instead anyway.
- kid0m4n 13y agoThanks :P And now an even faster version for your perusal: https://github.com/logicchains/Levgen-Parallel-Benchmarks/pull/11 https://github.com/logicchains/Levgen-Parallel-Benchmarks/pu... > 120 ms improvement over the origin implementation
- kid0m4n 13y agoThe original was running upto 8 goroutines (my runtime.NumCPU() is 8.)
- vanderZwan 13y agoI can't say anything about the other languages, but as far as Go is concerned, I suspect it's to achieve better load balancing. See Dmitri's explanation here: https://groups.google.com/d/msg/golang-nuts/CZVymHx3LNM/esYkA_YoB-MJ https://groups.google.com/d/msg/golang-nuts/CZVymHx3LNM/esYk... (Dmitri is one of the main people currently responsible for the further development of the Go scheduler). If you only have 8 goroutines running at a time with 8 cores, and one stalls, that wastes CPU usage. If you have, say, 80, the scheduler will put another one to work so the CPU is idle less often.
- kid0m4n 13y agoTrue, but that would apply only when you are dealing with tasks which are likely to stall. In this particular case, we are doing CPU intensive calculations in tight loops and no I/O. I scaled the number of goroutines from 1 to 8 (total cores in my MBP w/ Hyperthreading), and got best performance with 8. Performance started regressing beyond 8