4 ms·
So you make sure that you've multiple RabbitMQ servers or a RabbitMQ cluster so there is no single point of failure :)
by bhaisaab 13y ago
So you make sure that you've multiple RabbitMQ servers or a RabbitMQ cluster so there is no single point of failure :)
- mnutt 13y agoOk, more simply, what about a network partition where the redis machine gets cut off from the RabbitMQ servers? My only point is really just that redis-as-a-queue has some bad failure modes if you aren't careful.
- bhaisaab 13y agoSure, I agree and I understand things go wrong, so we've monitoring tools (munin, pingdom, pagerduty etc.) and we check them often or they post notifications often. There are at least two folks on-call 24x7. Based on present workloads we've calculations on how much time it would take to exhaust resources and we plan our servers accordingly, this buys us time to react.