6 ms·
It must be off by 3 of 4 orders of magnitude. Python with asyncio does about 120_000 switches per second with a trivial fibre/coroutine here: import time i
by erdewit 8y ago
It must be off by 3 of 4 orders of magnitude. Python with asyncio does about 120_000 switches per second with a trivial fibre/coroutine here:
import time
import asyncio
REPS = 1_000_000
async def coro():
for i in range(REPS):
await asyncio.sleep(0)
t0 = time.time()
asyncio.run(coro())
dt = time.time() - t0
print(REPS / dt, 'reps/s')
- blattimwind 8y agoasyncio eats about one order of magnitude. A more minimalistic async-approach gives something on the order of ~900000 resumptions per second. (Single-thread hardware performance differences would not account for a factor 10)
- blattimwind 8y agoSlightly modified "benchmark" from above: asyncio: import time import asyncio REPS = 1_000_000 async def coro(): for i in range(REPS): await asyncio.sleep(0) t0 = time.perf_counter() asyncio.run(coro()) dt = time.perf_counter() - t0 print(REPS / dt, 'reps/s') ~180k asynker (https://github.com/enkore/asynker https://github.com/enkore/asynker): import time import asynker REPS = 1_000_000 async def coro(): for i in range(REPS): await asynker.suspend() t0 = time.perf_counter() sched = asynker.Scheduler() sched.run_until_complete(coro()) dt = time.perf_counter() - t0 print(REPS / dt, 'reps/s') ~900k The excellent curio (https://github.com/dabeaz/curio https://github.com/dabeaz/curio): import time import curio REPS = 1_000_000 async def coro(): for i in range(REPS): await curio.sleep(0) t0 = time.perf_counter() curio.run(coro()) dt = time.perf_counter() - t0 print(REPS / dt, 'reps/s') ~180k Not that this number is hugely important.
- wenc 8y agoWould uvloop change these benchmarks? https://github.com/MagicStack/uvloop https://github.com/MagicStack/uvloop
- blattimwind 8y agoIt actually does, but again, this is a "how fast can you do nothing" microbenchmark. It's a bit of trivia for almost all intents and purposes. (The number's 500k)
- ovi256 8y agoGolang seems to do around 350k. There's a chance I'm missing some tricks, but the code is so short it's not probable I'm missing that much. See: https://repl.it/repls/GrimyChiefGame https://repl.it/repls/GrimyChiefGame package main import ( "fmt" "time" ) func coro(reps int) { for i := 0; i < reps; i++ { go time.Sleep(0 * time.Nanosecond) } } func main() { REPS := 5000000 start := time.Now() coro(REPS) dt := time.Since(start) fmt.Printf("The call took %v to run.\n", dt) fmt.Printf("REPS / duration %v\n", (REPS*1e9)/int(dt)) }
- weberc2 8y agoI don’t think you measured context switches per second, but rather goroutines forks per second. Am I mistaken?
- coder543 8y agoI don't think you're measuring context switching. if I remember correctly, Go's scheduler has a global queue and a local queue per worker thread, so when you spawn a goroutine it probably has to acquire a write lock on the global queue. Allocating a brand new goroutine stack and doing some other setup tasks has a nontrivial overhead that has nothing to do with context switching, regardless of global locks. To properly benchmark this, I think I would start with just measuring single task switching by measuring how long it takes main to call https://golang.org/pkg/runtime/#Gosched https://golang.org/pkg/runtime/#Gosched in a loop a million times. This would measure how quickly Go can yield a thread to the scheduler and have it be resumed, although this includes the overhead of calling a function. Then I would launch a goroutine per core doing this yield loop and see how many switches per second they did in total, and then launch several per core, just to ensure that the number hasn't changed much from the goroutine per core measurement. Since Go's scheduler is not bound to a single core, it should scale pretty well with core count. I might run this benchmark myself in awhile, if I find time.
- coder543 8y agoI wrote my own quick benchmark: https://gist.github.com/coder543/8c1b9cdffdf09c19ef61322bd26d2e44 https://gist.github.com/coder543/8c1b9cdffdf09c19ef61322bd26... The results: 1 switcher: 14_289_797.08 yields/sec 2 switchers: 5_866_478.94 yields/sec 3 switchers: 4_832_941.33 yields/sec 4 switchers: 4_604_051.57 yields/sec 5 switchers: 4_268_906.99 yields/sec 6 switchers: 3_982_688.58 yields/sec 7 switchers: 3_799_103.41 yields/sec 8 switchers: 3_673_094.58 yields/sec 9 switchers: 3_513_868.07 yields/sec 10 switchers: 3_351_813.00 yields/sec 11 switchers: 3_325_754.64 yields/sec 12 switchers: 3_150_383.56 yields/sec 13 switchers: 3_037_539.31 yields/sec 14 switchers: 2_435_807.77 yields/sec 15 switchers: 2_326_201.72 yields/sec 16 switchers: 2_275_610.57 yields/sec 64 switchers: 2_366_303.83 yields/sec 256 switchers: 2_400_782.51 yields/sec 512 switchers: 2_408_757.26 yields/sec 1024 switchers: 2_418_661.29 yields/sec 4096 switchers: 2_460_257.29 yields/sec Underscores and alignment added for legibility. It looks like the context switching speed when you have a single Goroutine just completely outperforms any of the benchmark numbers that have been posted here for Python or Ruby, as would be expected, and it still outperforms the others even when running 256 yielding tasks for every logical core. The cost of switching increased more with the number of goroutines than I would have expected, but it seems to become pretty constant once you pass the number of cores on the machine. Also keep in mind that this benchmark is completely unrealistic. No one is writing busy loops that just yield as quickly as possible outside of microbenchmarks. This benchmark was run on an AMD 2700X, so, 8 physical cores and 16 logical cores.
- ciconia 8y agoNot really equivalent since it does not involve an event reactor, only context switching: X = 1_000_000 f = Fiber.new do loop { Fiber.yield } end t0 = Time.now X.times { f.resume } dt = Time.now - t0 puts "#{X / dt.to_f}/s" ~ 780K on my machine
- ioquatix 8y agoHi, I'm the author of the article, and it was a while since I measured that number. Let me check it and fix it if it's wrong. Update: I wrote an addendum https://www.codeotaku.com/journal/2018-11/fibers-are-the-right-solution/context-switching https://www.codeotaku.com/journal/2018-11/fibers-are-the-rig...