8 ms·
What's the differences of using HAProxy or Envoy between using the cloud load balancers of AWS or Google Cloud?
by vecio 6y ago
What's the differences of using HAProxy or Envoy between using the cloud load balancers of AWS or Google Cloud?
- speedgoose 6y agoMore control and pricing I guess.
- ublaze 6y agoCloud load balancers can be sneakily expensive. Few months ago, we spent a few weeks replacing an ELB with naive client side load balancing via round robin, which saves us > 200k/year. ELBs charge per byte transmitted, which seems reasonable, but can end up really expensive.
- user5994461 6y agoThey charge per byte and per request if I remember well, which can be really expensive for serving both small API call and large files. Another limitation is that the ELB only works in AWS to AWS instances in the same location. Gotta use something else for geographic load balancing and for other datacenters.
- cutemonster 6y ago> client side load balancing In the browser? Or a mobile app? They send 1 api req to server 1, then 1 to server 2 and so on? What about any session cookies maybe tied to a specific server?
- Nextgrid 6y agoPresumably round-robin DNS. A DNS response would only return a handful of servers, of which the client will itself only pick one at random for the duration of the session. Now this approach has drawbacks (DNS responses are cached, and the DNS record picked initially by the client will typically be cached until the app/browser is restarted) but if they are acceptable to you then it's an easy, proven solution.
- cutemonster 6y agoHmm I'd guess they have DNS cnames like api1.x.com and api2 and 3, 4 And then the client picks one, and if that server is offline, picks another Seems as simple as DNS based? And works with broken server(s)
- boring_twenties 6y agoExcept you don't need that because you can just return all four IP addresses for one record, e.g. api.x.com
- cutemonster 6y agoI think if such a DNS/ip based round robin server is down, or replies 500 error, the client won't try another server Unless there's a way to get all ip addrs in js? By custom client code that queries the DNS system?
- Nextgrid 6y agoDNS resolution is handled by the DNS server and the browser. JS isn't involved, it's just telling the browser to connect to a certain hostname and the browser itself decides which IP to map it to (based on its DNS cache). If the DNS server is down the website wouldn't load at all, but this is an acceptable trade-off considering DNS is a very simple system (not many things can go wrong) and servers can be redundant.
- cutemonster 6y agoThere's a misunderstanding. That reply is to me off topic, I knew about those DNS things already Thanks anyway for replying
- jrockway 6y agoI've found that the cloud load balancers lag behind the state of the art in features and that their assumptions and configurations can be pretty brittle. I haven't used Amazon's ALB, but with the legacy ELB, they can't speak ALPN. So that means, if you use their load balancer to terminate TLS, you can't use HTTP/2. Their automatic certificate renewal silently broke for us as well; whereas using cert-manager to renew Let's Encrypt certificates continues to work perfectly wherever I use it. (At the very least, cert-manager produces a lot of logs, and Envoy produces a lot of metrics. So at the very least, when it does break, you know what to fix. With the ELB, we had to pray to the Amazon gods that someone would fix our stuff on Saturday morning when we noticed the breakage. They did! But I don't like the dependency.) I have also used Amazon's Network Load Balancer with EKS. It interacts very weirdly. The IP address that the load balancer takes on changes with the pods that back the service. The way the changes happen is that the NLB updates a DNS record with a 5 minute TTL. So you have a worst case rollout latency of 5 minutes, and there is no mechanism in Kubernetes to keep the old pods alive until all cached DNS records have expired. The result is, by default, 5 minutes of downtime every time you update a deployment. Less than ideal! For that reason, I stuck with ELB pointing to Envoy that terminated TLS and did all complicated routing. The ALB wouldn't have these problems. It's just a HTTP/2 server that you can configure using their proprietary and untestable language. It has some weak integration with the Kubernetes Ingress type, so in the simplest of simple cases you can avoid their configuration and use a generic thing. But Ingress misses a lot of things that you want to do with HTTP, so in my opinion it causes more problems than it solves. (The integration is weak too. You can serve your domain on Route 53, but if you add an Ingress rule for "foo.example.com", it's not going to create the DNS record for you. It's very minimum-viable-product. You will be writing custom code on top of it, or be doing a lot of manual work. All in all, going to scale to a large organization poorly unless you write a tool to manage it, in which case you might as well write a tool to configure Envoy or whatever.) In general, I am exceedingly disappointed by Layer 3 load balancers. For someone that only serves HTTPS, it is completely pointless. You should be able to tell browsers, via DNS, where all of your backends are and what algorithm they should use to select one (if 503, try another one, if connect fails, try another one, etc.) But... browsers can't do that, so you have to pretend that you only have one IP address and make that IP address highly available. Google does very well with their Maglev-based VIPs. Amazon is much less impressive, with one IP address per AZ and a hope and a prayer that the browser does the right thing when one AZ blows up. Since AZs rarely blow up, you'll never really know what happens when it does. (Chrome handles it OK.)
- ryan_lane 6y agoOne of the reasons Envoy was built was because ELB/ALBs have opaque observability and fail in ways you can't control.