5 ms·
Thanks for providing your feedback. As Redis Manifesto states - our goal is to fight against complexity. antirez - you are our inspiration and I seriously take
by romange 4y ago
Thanks for providing your feedback. As Redis Manifesto states - our goal is to fight against complexity. antirez - you are our inspiration and I seriously take your manifesto close to heart.
Please allow the possibility that Redis can be improved and should be improved.
Otherwise other systems will eventually take its market apart.
I appreciate your comments very much. I've wrote about you in my blog. I am an engineer and I disagree with some of the design decisions that were made in Redis and I decided to do something about it :)
to your points:
1. DF provides full compatibility with single node Redis while running on all cores, compared to Redis cluster that can not provide multi-key operations across slots.
2. Much stronger point - we provide much simpler system since you do not need to manage k processes, you do not need to *provision* k capacities that managed independently within each process and you do not need to monitor those processes, load/save k snapshots etc. Our snapshotting is point in time on all cores.
3. Due to pooling of resources DF is more cost efficient. It's more versatile. We have a design partner that could reduce its costs by factor of 3 just because he could use x2gd machine with extra high memory configuration.
Regarding your note about memcached - while we provide similar performance like memcached our product proposition is anything unlike memcached and it's more similar to Redis. Having said that - I will add comparison to memcached. I do believe that memcached as performant as DF because essentially it's just an epoll loop over multiple threads.
Re you comment about snapshotting. We also push the data into serialization sink upon write, hence we do not need to aggregate changes until the snapshot completes.
The complex part is to ensure that no key is written twice and that we ensure with versioning. I do agree that there can be extreme cases where we need to duplicate memory usage for some entries but it's only for the entries at flight - those that are being processed for serialization.
Update: re versioning and memory efficiency. We use DashTable that is more memory efficient that Redis-Dict. In addition, DashTable has a concept of bucket that is comprised of multiple slots (14 in our implementation). We maintain a single 64bit version per bucket and we serialize all the entries in the bucket at once. Naturally, it reduces the overhead of keeping versions. Overall, for small value workloads we are 30-40% more efficient in memory than Redis.
- reconditerose 4y agoAs one of the folks that currently works on Redis, I want to highlight the "Redis can be improved and should be improved". There is a lot of really good ideas put forth that are likely worth consideration in the Redis project as well. There has been a lot of conversations about renewing multi-threading, especially to address the point of simplifying management and better resource utilization. Glad to see you guys made a lot of progress, although a little disappointing you chose to go down the path of building yet another source available DB and not contributing to open source.
- tayo42 4y agoCan someone/outsider realistically show up and start working on making redis multi threaded?
- reconditerose 4y agoYes, the core folks in Redis now have outlined a plan for moving towards multi-threaded awhile ago. It's honestly not made a lot of progress because raw performance matters a lot less than is argued here. As was succinctly mentioned by antirez already, Redis scales comfortably to 100s of millions of QPS with cluster mode. So, it's really building a lot of custom functionality to support better vertical scaling. Which is useful, especially when the vertical scaling keeps you on a single process. The conversation happened here, https://github.com/redis/redis/issues/8340 https://github.com/redis/redis/issues/8340, and it's not like the most pressing issue for the project. It's also not as complex as what was implemented for dragonfly, which basically has native support from the ground up for concurrent programming during command execution. It would be hard to do in C as well.
- tayo42 4y ago> raw performance matters a lot less than is argued here It matters a lot for where i work, we believe multithreading is holding redis peformance back. > Redis scales comfortably to 100s of millions of QPS with cluster mode. is there somewhere i can read more about that? curious about the server needs to do that. i worth key/value clusters a little larger then that. if possible it would be cool to use redis for it.
- kristoff_it 4y agoIt seems to me you didn't address the main point from parent: did you benchmark your multithreaded implementation vs a single core Redis? Nevermind the amazing advantages that having to spawn 1 process vs N brings, the question is how does your software compare when Redis is used as inteded.
- romange 4y agoI benchmarked DF vs single core Redis. If there is a constructive suggestion for a different benchmark that compares similar product propositions I will happily oblige and do that. i.e. what do you mean by using Redis as intended?
- Game_Ender 4y agoSo two options I am curious about: - If it's a normal configuration partitioning a single large node with multiple instances using Redis cluster - A cost equivalent cluster of machines with a similar memory size running on Redis cluster
- romange 4y ago1. I think this is how Redis the company designed their enterprise solution. You can find architecture documentation on their site. 2. Based on my knowledge it should be more or less equivalent. The reason they put it on the same machine (i am guessing here) is because shards on the same node are behind their Redis proxy, that kinda hides the complexity of connecting to each node separately. it's like a gateway to that machine and to all its redis processes.
- antirez 4y agoThanks for the nice words romange. The complexity here can be seen in two ways: complexity of deploying more Redis instances, or complexity of the single instance. It's a trade off. But I think that Redis may go fully threaded soon or later, and perhaps your project may accelerate the process (I'm no longer involved, just speculating). 1. Your point about Cluster, I addressed it many times: the point is, soon or later even with multi-threading you are going to shard among N machines. So I believe that to have this problem ASAP is better and more "linear". 2. Already addressed in "1" and my premise. 3. Yep there are advantages in certain use cases related to cloud costs and so forth, that's why maybe Redis will end fully threaded as well. About memory efficiency, what I meant is that to have versioned data structures, that is an approach to do user-space copy on write even in the case of multiple changes to large single keys (big sorted set example), you need more memory likely, to augment the data structure. Otherwise the trick is to copy the whole value, that has other issues. It's a tradeoff.
- phamilton 4y ago> soon or later even with multi-threading you are going to shard among N machines In a world where cloud providers offer instances with terabytes of memory and 128 vCPUS (e.g. aws x2iedn.32xlarge family maxes out at 4TB, gcp m2 family maxes out at 12TB) is that really inevitable? Applications serving 10s of millions of users likely won't come anywhere close to that limitation.
- romange 4y agoIt's interesting comment, phamilton. I agree with antirez and I agree with you. full discloser - I am ex-googler. I believe that "horizontal scale" movement staryted from Google. You could see it in their GFS and Mapreduce papers from early 2000s. And by that time they were completely right, of course. Geniuses of Jeff Dean and Sanjay Gwattemat put Google year ahead compared to other state of the art for more than a decade. There was not even one system developed in Google that is not horizontally scalable (i am omitting acquisitions here). We used to joke in 2009 that the most expensive server we have is our perforce server. Nowdays it's internally developed source control system that is backed by Bigtable. So of course, antirez is right, of course! If you need infinite scale - you must go horizontally. But the reality is that most companies and most use-cases do not need terrabytes of data. I would say that today the comfort zone for Dragonfly is upto 512GB per instance (1). So dragonfly solves the issue for... I would say 99% percent of the use-cases. Only the last percentile would need horizontal scale, and probably their business is already big enough, so that they can affort a high-quality eng team to work with horizontal clusters. (1) We need to improve some things (mainly around serialization format of rdb) to reach another magnitude of 4TB. Nobody wants to wait for days to load a 4TB snapshot.