12 ms·
A few pieces of advice based on running https://github.com/mozilla-services/autopush-rs https://github.com/mozilla-services/autopush-rs, which handles tens of m
by sciurus 7y ago
A few pieces of advice based on running https://github.com/mozilla-services/autopush-rs https://github.com/mozilla-services/autopush-rs, which handles tens of millions of concurrent connections across a fleet of small EC2 instances.
1) Consider not running the largest instance you need to handle your workload, but instead distributing it across smaller instances. This allows for progressive rollout to test new versions, reduces the thundering herd when you restart or replace an instance, etc.
2) Don't set up security group rules that limit what addresses can connect to your websocket port. As soon as you do that connection tracking kicks in and you'll hit undocumented hard limits on the number of established connections to an instance. These limits vary based on the instance size and can easily become your bottleneck.
3) Beware of ELBs. Under the hood an ELB is made of multiple load balancers and is supposed to scale out when those load balancers hit capacity. A single load balancer can only handle a certain number of concurrent connections. In my experience ELBs don't automatically scale our when that limit is reached. You need AWS support to manually do that for you. At a certain traffic level, expect support to tell you to create multiple ELBs and distribute traffic across them yourself. ALBs or NLBs may handle this better; I'm not sure. If possible design your system to distribute connections itself instead of requiring a load balancer.
2 and 3 are frustrating because they happen at a layer of EC2 that you have little visibility into. The best way to avoid problems is to test everything at the expected real user load. In our case, when we were planning a change that would dramatically increase the number of clients connecting and doing real work, we first used our experimentation system to have a set of clients establish a dummy connection, then gradually ramped up that number of clients in the experiment as we worked through issues.
- dylanz 7y agoCould you explain or post a link to something about #2? I’ve never heard of that before!
- thecopy 7y agoI would also be interested in this. First time I hear about such a limitation.
- sciurus 7y agoI can not. AWS, for whatever reason, does not want this publicly documented. I would write up what we learned in our testing in more detail, but I have been asked not to. You can find some discussion of this behavior in places like https://forums.aws.amazon.com/thread.jspa?threadID=231806 https://forums.aws.amazon.com/thread.jspa?threadID=231806. I originally became aware of the issue, before hitting in production, from the HN comment at https://news.ycombinator.com/item?id=18314138 https://news.ycombinator.com/item?id=18314138
- toast0 7y agoI'm glad you found my comment helpful! I myself found out about to avoid the connection tracking in this thread https://news.ycombinator.com/item?id=15724072 https://news.ycombinator.com/item?id=15724072 It's super frustrating that Amazon doesn't document this in a more approachable manner.
- QuinnyPig 7y agoTheir response displeases me.
- runT1ME 7y agoI can corroborate this, at Verizon we had problems scaling out ELBs for a very similar use case.
- toredash 7y agoThis is what you need to know: "Not all flows of traffic are tracked. If a security group rule permits TCP or UDP flows for all traffic (0.0.0.0/0) and there is a corresponding rule in the other direction that permits all response traffic (0.0.0.0/0) for all ports (0-65535), then that flow of traffic is not tracked. The response traffic is therefore allowed to flow based on the inbound or outbound rule that permits the response traffic, and not on tracking information."
- chatmasta 7y agoVery interesting. It makes a lot of sense from a Linux perspective. The conntrack table requires memory, so there is some physical limit on its size. That also explains why it scales per instance type.
- toredash 7y agoThis is also why "endless scaling" is a fud. At some point or some level, there is either a physical, soft, standard or some kind of limit
- gchamonlive 7y agoI recently played around with Athena for load balancer logs, and with sql it is easy to cross reference many connection entries from the load balancer. Do you believe this could help in getting more visibility to spot bottleneck problems stated in 2 and 3?
- sciurus 7y agoNot really, because there won't be logs for connections that were never established.
- johne20 7y agoHow are you keeping track of those?
- sciurus 7y agoAs I recall, when we hit the ELBs connection limits the failed connections were reflected in the SpilloverCount metric. I believe we only hit the security group limits in load testing. There, we could see connections fail to an instance when that instance's established connections as reported by netstat or similar tools hit a certain threshold.
- Thaxll 7y agoIf you have a lot of trafic you shoudn't use an ELB in the first place, should be using an NLB by default it comes with a 40Gb pipe and multi millions connections. https://docs.aws.amazon.com/elasticloadbalancing/latest/network/introduction.html https://docs.aws.amazon.com/elasticloadbalancing/latest/netw...
- arrty88 7y agoYou can also employ DNS round robin with health checks in Route53
- nathancahill 7y agoThat's the first tool I grab for balancing _a lot_ of connections.
- Thaxll 7y agoELB is pretty old and should be avoided tbh, use ALB and if you can't should switch to NLB ( for pure tcp LB ).
- windlep 7y agoSame issue as NLB, they charge for units based on concurrent connections and it gets very expensive very quickly.
- song 7y agoDNS round robin has the disadvantage of not guaranteeing that the traffic will be equally distributed. This can be a problem if there's a lot of automated traffic. But, with that said, in term of price, it's unbeatable.
- sieabahlpark 7y agoThe point is at scale you can't guarantee equal but something around there
- gravypod 7y ago> 1) Consider not running the largest instance you need to handle your workload, but instead distributing it across smaller instances. This allows for progressive rollout to test new versions, reduces the thundering herd when you restart or replace an instance, etc. I came here to say this. Horizontal compute is a miracle.
- verttii 7y agoDistribution itself might be a world of pain depending on the nature of your application. For example stateful apps where low latency is expected.
- semiotagonal 7y agoHow meaningful is the benchmark for handling idle connections? Handling more is better, of course, but if the server melts down when 0.1% of them have any activity, maybe maximizing idle connections isn't the right place to expend optimization effort?
- cortesoft 7y agoDepends on your workload
- sciurus 7y agoFor us, the experiment (https://blog.mozilla.org/services/2018/10/03/upcoming-push-shield-study/ https://blog.mozilla.org/services/2018/10/03/upcoming-push-s...) with mostly idle clients was very helpful, since it flushed out problems with out load balancing layer and in the end gave us confidence that it, our app, and our persistence layer could safely handle many more connections. If we had begun having problems once we started sending more push messages, we would have simply stopped using the new service (https://github.com/mozilla-services/megaphone https://github.com/mozilla-services/megaphone) responsible for that until we worked through them.