5 ms·
I wonder if Tokio is also a reason for worse performance compared to Go's concurrency runtime.
by mvelbaum 4y ago
I wonder if Tokio is also a reason for worse performance compared to Go's concurrency runtime.
- raverbashing 4y agoI wonder how many people are using 'async' just for the sake of it without a real need for it and shooting themselves in the foot while at it
- pkolaczk 4y agoActually in this case async is the only way to get sane performance and both drivers deliver excellent performance thanks to async. I've been using Scylla Rust driver in my C* benchmarking project and it is an order of magnitude faster than the tools which use threads. https://github.com/pkolaczk/latte https://github.com/pkolaczk/latte
- raverbashing 4y agoCool, good to know. I know threads are a limiting factor, but sometimes people jump into async while the problem is somewhere else
- pkolaczk 4y agoIn this case each request is very tiny amount of work on the client, so waking up a thread to do that work just to immediately block waiting on the response from the server is very wasteful. With async you can send hundreds of requests in a simple loop, on a single thread. It's not only more efficient but also actually easier to write.
- insanitybit 4y agoThere are probably a bunch of reasons, which is why I want an easy "run benchmarks" command that I can use. I'd even be fine using infra so long as I had pulumi/terraform to set it all up for me. I just don't want to spin up EC2 instances manually, get the connections all working, make sure I can reset state, etc. I already have a fork of Scylla where I removed a lot of unnecessary cloning of `String` but no way I'm gonna PR it without a benchmark. I also opened a PR to replace the hash algorithm used in their PreparedStatement cache, which gets hit for every query, but they wanted benchmarks before accepting (completely fair) and I have none. `ahash` is extremely fast compared to Rust's default - https://github.com/tkaitchuck/ahash https://github.com/tkaitchuck/ahash and with the `comptime` randomness (more than sufficient for the scylla use case) you can avoid a system call when creating the HashMap. There are also some performance improvements I have in mind for the response parsing, among other things.
- ianpurton 4y agoSo would the ideal solution be if ScyllaDB had a github action to run benchmarks against PR's? Not sure how decent a benchmark would be without running up servers in the cloud. So I guess provisioning infra would be a requirement? So perhaps this could be run manually. But certainly possible - Pulumi up infra - Run benchmarks - Collect results - Attach to PR.
- insanitybit 4y agoI'd be happy with a few things: 1. Benchmarks of "pure" code like the response parser, which I could `cargo bench`. I may actually work on contributing this. 2. Some way to run benchmarks against a deployed server. I wouldn't recommend a Github action necessarily, a nightly job or manual job would probably be a better use of money/resources. If I could plug in some AWS creds and have it do the deployment and spit out a bunch of metrics for me that'd be wonderful.
- indiv0 4y agoI just did a comparison between almost every hashing algorithm I could find on crates.io. On my machine t1ha2 (under the t1ha crate) beat the pants off of every other algorithm. By like an order of magnitude. Others in the lead were blake3 (from the blake3 crate) and metrohash. Worth taking a look at those if you’re going for hash speed. I don’t have the exact numbers on me right now but I can share them tomorrow (along with the benchmark code) if you’re interested.
- insanitybit 4y agoThe PR I have lets you provide the algorithm as the caller, although I did benchmark against fxhash and I think it would be a good idea to suggest `ahash`. I'm certainly interested. `ahash` has some good benchmarks here: https://github.com/tkaitchuck/aHash/blob/master/FAQ.md https://github.com/tkaitchuck/aHash/blob/master/FAQ.md
- ComputerGuru 4y agoFYI Small hashes beat better quality hashes for hash table purposes.