5 ms·
Their explanation for why Go performs badly didn't make any sense to me. I'm not sure if they don't understand how goroutines work, if I don't understand how go
by latch 2y ago
Their explanation for why Go performs badly didn't make any sense to me. I'm not sure if they don't understand how goroutines work, if I don't understand how goroutines work or if I just don't understand their explanation.
Also, in the end, they didn't use the JSON payload. It would have been interesting if they had just written a static string. I'm curious how much of this is really measuring JSON [de]serialization performance.
Finally, it's worth pointing out that WebSocket is a standard. It's possible that some of these implementations follow the standard better than others. For example, WebSocket requires that a text message be valid UTF8. Personally, I think that's a dumb requirement (and in my own websocket server implementation for Zig, I don't enforce this - if the application wants to, it can). But it's completely possible that some implementations enforce this and others don't, and that (along with every other check) could make a difference.
- vandot 2y agoThey didn’t use goroutines, which is explains the poor perf. https://github.com/matttomasetti/Go-Gorilla_Websocket-Benchmark-Server/blob/master/go-gorilla_websocket-benchmark-server.go#L58 https://github.com/matttomasetti/Go-Gorilla_Websocket-Benchm... Also, this paper is from Feb 2021.
- windlep 2y agoI was under the impression that the underlying net/http library uses a new goroutine for every connection, so each websocket gets its own goroutine. Or is there somewhere else you were expecting goroutines in addition to the one per connection?
- donjoe 2y agoWhich is perfectly fine. However, you will be able to process only a single message per connection at once. What you would do in go is: - either a new goroutine per message - or installing a worker pool with a predefined goroutine size accepting messages for processing
- jand 2y agoAnother option is to have a read-, and a write-pump goroutine associated with each gorilla ws client. I found this useful for gateways wss <--> *.
- initplus 2y agohttp.ListenAndServe is implemented under the hood with a new goroutine per incoming connection. You don't have to explicitly use goroutines here, it's the default behaviour.
- necrobrit 2y agoYes _however_ the nodejs benchmark at least is handling each message asynchronously, whereas the go implementation is only handling connections asynchronously. The client fires off all the requests before waiting for a response: https://github.com/matttomasetti/NodeJS_Websocket-Benchmark-Client/blob/master/connection.js#L122-L133 https://github.com/matttomasetti/NodeJS_Websocket-Benchmark-... so the comparison isn't quite apples to apples. Edit to add: looks like the same goes for the c++ and rust implementations. So I think what we might be seeing in this benchmark (particularly the node vs c++ since it is the same library) is that asynchronously handling each message is beneficial, and the go standard libraries json parser is slow. Edit 2: Actually I think the c++ version is async for each message! Dont know how to explain that then.
- josephg 2y agoWell, tcp streams are purely sequential. It’s the ideal use case for a single process, since messages can’t be received out of order. There’s no computational advantage to “handling each message asynchronously” unless the message handling code itself does IO or something. And that’s not the responsibility of the websocket library.
- necrobrit 2y agoGood point!
- ikornaselur 2y agoYeah I thought this looked familiar.. I went through this article about a year and a half ago when exploring WebSockets in Python for work. With some tuning and using a different libraries + libuv we were easily able to get similar performance to NodeJS. I had a blog post somewhere to show the testing and results, but can't seem to find it at the moment though.
- tgv 2y ago> I'm curious how much of this is really measuring JSON [de]serialization performance. Well, they did use the standard library for that, so quite a bit, I suppose. That thing is slow. I've got no idea how fast those functions are in other languages, but you're right that it would ruin the idea behind the benchmark.
- bryancoxwell 2y agoAre you referring to Go’s stdlib?
- tgv 2y agoYes, the repo uses encoding/json from the standard library.
- klabb3 2y ago> Their explanation for why Go performs badly didn't make any sense to me. To me, the whole paper is full of misunderstanding, at least the analysis. There's just speculation based on caricatures of the language, like "node is async", "c++ is low level" etc. The fact that their C++ impl using uWebSocket was significantly slower than then Node, which used uWebSocket bindings, should have led them to question the test setup (they probably used threads which defeats the purpose of uWebSocket. Anyway.. The "connection time" is just HTTP handshake. It could be included as a side note. What's important in WS deployments are: - Unique message throughput (the only thing measured afaik). - Broadcast/"multicast" throughput, i.e. say you have 1k subscribers you wanna send the same message. - Idle memory usage (for say chat apps that have low traffic - how many peers can a node maintain) To me, the champion is uWebSocket. That's the entire reason why "Node" wins - those language bindings were written by the same genius who wrote that lib. Note that uWebSocket doesn't have TLS support, so whatever reverse proxy you put in front is gonna dominate usage because all of them have higher overheads, even nginx. Interesting to note is that uWebSocket perf (especially memory footprint) can't be achieved even in Go, because of the goroutine overhead (there's no way in Go to read/write from multiple sockets from a single goroutine, so you have to spend 2 gorountines for realtime r/w). It could probably be achieved with Tokio though.
- Svenskunganka 2y agoThe whole paper is not only full of misunderstandings, it is full of errors and contradictions with the implementations. - Rust is run in debug mode, by omitting the --release flag. This is a very basic mistake. - Some implementations is logging to stdout on each message, which will lead to a lot of noise not only due to the overhead of doing so, but also due to lock contention for multi-threaded benchmarks. - It states that the Go implementation is blocking and single-threaded, while it in fact is non-blocking and multi-threaded (concurrent). - It implies the Rust implementation is not multi-threaded, while it in fact is because the implementation spawns a thread per connection. On that note, why not use an async websocket library for Rust instead? They're used much more. - Gives VM-based languages zero time to warm up, giving them very little chance to do one of their jobs; runtime optimizations. - It is not benchmarking websocket implementations specifically, it is benchmarking websocket implementations, JSON serialization and stdout logging all at once. This adds so much noise to the result that the result should be considered entirely invalid. > To me, the champion is uWebSocket. That's the entire reason why "Node" wins [...] A big part of why Node wins is because its implementation is not logging to stdout on each message like the other implementations do. Add a console.log in there and its performance tanks.