4 ms·
This comment may be a bit snarky, but it's important to understand that the TL;DR for this is not: you need to learn how to debug Redis and understand a lot of
by randomdrake 10y ago
This comment may be a bit snarky, but it's important to understand that the TL;DR for this is not: you need to learn how to debug Redis and understand a lot of internal things about it and the libraries using it. It's that you can really save yourself a ton of time and headache with a better understanding of application architecture.
Step 1) Throw everything at Redis because you don't know how to architect. Use it as a cache. Use it as a temporary data store. And use it as a persistent data store at the same time. What could go wrong, right?
Step 2) Watch it break because you're using it wrong.
Step 3) Write a blog post about everything you discovered about how you can sort of get it working while continuing to do it wrong.
Step 4) Only concede at the very end of the blog post that you were just doing it wrong™ from the beginning.
"The root causes of these errors were more a conceptional[sic] issue. The way we used Redis was not ideal for this kind of traffic and growth."
Unless I'm missing something, there's really no reason that all three use cases (caching, temporary data store, persistent data store) should all be sitting on a single Redis server so far away from the metal of the application server.
I do not want my search cache waiting on temporary data to be read and flushed from the cache. I do not want data I'd like to persist competing with search cache results. I don't want any of these things! That's not how this works. That's not how any of this works!
- omn1 10y agoIn any bigger project, architecture grows organically. It's easy to say "I would have done better". In hindsight, every flawed system looks like a bad idea. It takes a visionary to foresee all possible error cases. Things go wrong. I think it's an honest post with lots of practical tips.
- randomdrake 10y ago> In any bigger project, architecture grows organically. It's easy to say "I would have done better". In hindsight, every flawed system looks like a bad idea. It takes a visionary to foresee all possible error cases. Kudos to the team on their findings and for their honest post with some tips for their odd architecture. I hope my comment did not imply that "I would have done better," or that hindsight is not, 20/20. I have enough experience to understand both of those things. However, as architecture grows organically, it doesn't take too much understanding or experience to know that using a single Redis instance for three different use cases would probably run into issues. That's more than likely not a symptom of "organic growth," but of poor decision making. Trivago received a $4,000,000,000 valuation. Why does their engineering deserve the benefit of the doubt here? At some point we (technologists and developers) have got to concede that maybe it's not just a few folks trying to keep up with organic growth, but that some actual bad decisions were made and should be accepted and highlighted more than just covering up for mistakes. I think the post could have been more useful if it would have conceded the point earlier that they made bad decisions from the get-go, that they had to play catch up with those decisions, and that the advice was for attempting to catch up to these decisions. It would be very interesting to understand why they made the choices they did. Why are they running memcached and Redis at the same time? Why wasn't a regular database suitable for data persistence or temporary data? Why wasn't the cache for search installed on the application server to remove network latency altogether? Why weren't they running php-fpm from the get-go? I made the comment because doing some research or having some understanding from the beginning, might have avoided the post altogether. I think it's important for us to be able to own our mistakes and honestly approach how we made those mistakes so they aren't repeated by others. I think that is far more useful than a writeup on how to adjust some settings to deal with some fairly basic mistakes. Especially for a technical blog post from a company with a $4B valuation.
- jlg23 10y ago> However, as architecture grows organically, it doesn't take too much understanding or experience to know that using a single Redis instance for three different use cases would probably run into issues. That's more than likely not a symptom of "organic growth," but of poor decision making. If one always had time to reflect upon problems during development you were right. But during crunch time one makes a lot of decisions without proper evaluation just based on a (more or less) educated guess. > Trivago received a $4,000,000,000 valuation. Why does their engineering deserve the benefit of the doubt here? Because a lot of readers here have worked in (potentially over-evaluated) environments that constantly underestimate the time required to build something. That someone reflected upon their mistakes and published a post mortem is a good sign someone found the time to finally re-visit all those previous decisions with a rested mind and some distance to the frantic time of "just getting things done".
- andygrunwald 10y agoOriginal author here. Thanks for your comment. I don't think you mean everything bad or snarky here. I like to read critical feedback. Thats how you learn. But i guess you mix things up. So yes. This architecture has grown over the time. And the whole story didn't happen yesterday. It starts September 3, 2010. So >6 years back. The first _Real_ problems appeared in 2013. So ~4 years back. With our current knowledge, i agree that those three use cases should not fit in one redis instance. > I think the post could have been more useful if it would have conceded the point earlier that they made bad decisions from the get-go, that they had to play catch up with those decisions, and that the advice was for attempting to catch up to these decisions. This post is written as a kind of story. Of course it would be possible to conduct this post to learnings. > It would be very interesting to understand why they made the choices they did. Feel free to ask every question you want to know. I am here and happy to answer. > Why are they running memcached and Redis at the same time? Because this are two different systems with two different concepts for different usecases. If you use caching, i agree. Both can cache data. But on use cases where you need master slave replication (maybe across data centers), memcached may not fit. IMO it is hard to say both systems are the same. E.g. if you deal with caching data that vary a lot in size per entry (same data concept), one memcache instance / consistent hashing ring might be the wrong solution, because of the slap concept of memcache. > Why wasn't a regular database suitable for data persistence or temporary data? This question is not answered in general. But i use one use case. We had a MySQL table running on InnoDB that was really read heavy with ~300.000.000 lines. The indexes were set for the read patterns. Everything fine. An normal insert in this table took some time. And we wanted to avoid to spent this time during a user request. This would slow down the request. One option the dev team considered is to write it in redis and have a cronjob that reads this data out of redis and stores them in the table. With this the insert time was moved from the user back to the system. > Why wasn't the cache for search installed on the application server to remove network latency altogether? There is a cache on the application server. But we have several cache layer in our architecture. This was just one of them. > Why weren't they running php-fpm from the get-go? We are working on a switch to php-fpm from the typical prefork model. Such a task sounds easy. But in a bigger env this can get quite a challenge. > I made the comment because doing some research or having some understanding from the beginning, might have avoided the post altogether. Agree. The challenge here was, that the people who introduced Redis at that time has left the company. So in short: We had the problem, and had to fix it. With our current knowledge, we would fix this in a different way and maybe choose different approaches. But yeah, i assume this is a normal learning process. Anyhow. Thanks for your feedback.