7 ms·
Hate to be a party pooper, but I'd like to give people here a more generic mental tool to solve this problem. Ignoring Elixir and Erlang - when you discover yo
by jondot 10y ago
Hate to be a party pooper, but I'd like to give people here a more generic mental tool to solve this problem.
Ignoring Elixir and Erlang - when you discover you have a backpressure problem, that is - any kind of throttling - connections or req/sec, you need to immediately tell yourself "I need a queue", and more importantly "I need a queue that has a prefetch capabilities". Don't try to build this. Use something that's already solid.
I've solved this problems 3 years ago, having 5M msg/minute pushed _reliably_ without loss of messages, and each of these messages were checked against a couple rules for assertion per user (to not bombard users with messages, when is the best time to push to a a user, etc.), so this adds complexity. Later approved messages were bundled into groups of a 1000, and passed on to GCM HTTP (today, Firebase/FCM).
I've used Java and Storm and RabbitMQ to build a scalable, dynamic, streaming cluster of workers.
You can also do this with Kafka but it'll be less transactional.
After tackling this problem a couple times, I'm completely convinced Discord's solution is suboptimal. Sorry guys, I love what you do, and this article is a good nudge for Elixir.
On the second time I've solved this, I've used XMPP. I knew there were risks, because essentially I'm moving from a stateless protocol to a stateful protocol. Eventually, it wasn't worth the effort and I kept using the old system.
- Vishnevskiy 10y agoI think you misunderstand the problem we are solving here. We are not trying to solve this because our system can't handle it. We are protecting it from when Firebase decides to slowdown in a way that causes data to backup and OOM the system. Since these are push notifications that have a time bound on usefulness we don't care about dumping to an external persisted queue like RabbitMQ or Kafka (we rather deliver newer notifications faster, than wait for the backed up buffer to flush). Firebase also only allows 1000 concurrent connections per senderId with 100 inflight pushes (that have not received an ack) which means that only 100,000 can be inflight. Ultimately if a remote service is providing backpressure because it is having a struggle no amount of auto scaling on your end is going to help you. This service buffers potential pushes for all users being messages, that then watches the presence system to determine if they are on their desktop or mobile (this is millions of presence watchers and 10s of millions of buffered messages), and users are constantly clearing these buffers by reading on the clients and finally when a user is offline or goes offline we emit their pushes to them (which is what this article talks about). This service was evolved from our push system from the game we worked on and when it just did pushes only and no other logic it could push at 1m/sec in batches, but its responsibility has changed. Context matters :)
- metafunctor 10y agoCould you not reach pretty much the same result with a queue, though? For example, workers could discard messages older than some threshold, quickly emptying the queue if there are expired messages. Clients might not even queue messages if the queue is currently too long, perhaps even providing a convenient signal for them to back off from their most chatty behaviour. Some messages will not be delivered on time if there is significant backpressure. There is not much you can do about it, apart from avoiding choking yourself. Perhaps the queue could work with a LIFO policy, to help at least some messages go through in time instead of having most messaged delayed near to the expiration threshold.
- jondot 10y agoContext definitely matters - thanks for the background info. I understood what you're trying to do, incidentally I did the same (all these details sound familiar to me). At the time, Storm helped me batch, break batches and validate discrete units, re-batch, aggregate, repeat that process how many times I wanted, and finally, batch the stream with a strategy I wanted (number of users, messages, or balance number of connections) and deliver to Google. Then I would just say, XMPP is new and fancy, but consider the old fashioned stateless HTTP interface. When I was implementing my own service, I was worried Google is not going to handle the load. Since we were partners with Google for a good while I was able to climb the ladder of people to get an answer, and plow through their closed-door policy for questions such as "Will you guys handle this load? (5M msg/min)". I wrote a huge email explaining every edge case and what I'm doing. The answer was "We will handle it.". No detail, no context, no buts. I wasn't confident at all. But in the end, they did handle it :)
- teacpde 10y ago> You can also do this with Kafka but it'll be less transactional. Could you explain why using RabbitMQ is more transactional?
- chillydawg 10y agoI've no idea about kafka but rabbit offers you message ack/nack and publisher confirms. Generally, you can build very solid things on top of it, depending on whether you need to distribute rabbit or not.
- di4na 10y agoKnowing that RabbitMQ is full erlang, why bring a really big dependency if you have all the things in your everyday language anyway ?
- jondot 10y agoAn all encompassing answer would be - the abstractions. Why would you use an operating system if really you have your CPU documented and know all of the instruction sets?
- gregpardo 10y agoYeah this... Hey guys you don't need erlang to solve this problem... just use this tool built in erlang.
- neiled 10y agoAnyone have any nice links to describe more about queue prefetching as described in this case? My google skills are failing because of all the CPU related articles.
- jondot 10y agohttp://www.rabbitmq.com/consumer-prefetch.html http://www.rabbitmq.com/consumer-prefetch.html