13 ms·
The GitHub Load Balancer
- otoburb 10y agoGiven this is based on HAProxy and seems to improve the director tier of a typical L4/L7 split design, I'm led to believe GLB is an improved TCP-only load balancer. But they also talk about DNS queries, which are still mainly UDP53, so I'm hoping GLB will have UDP load-balancing capability as gravy on top. I excluded zone transfers, DNSSEC traffic or (growing) IPv6 DNS requests on TCP53 because, at least in carrier networks, we're still seeing a tonne of DNS traffic that still fits within plain old 512-byte UDP packets. Looking forward to seeing how this develops. EDIT: Terrible wording on my part to imply that GLB is based off of HAProxy code. I meant to convey that GLB seems to have been designed with deep experience working with HAProxy as evidenced by the quote: "Traditionally we scaled this vertically, running a small set of very large machines running haproxy [...]".
- otoburb 10y agoEDIT: I read the post more carefully - the reference to DNS was merely to highlight that often times a single public [V]IP can get overloaded and common mitigation strategies.
- dcgudeman 10y agoWhere did it say that it is based on HAProxy?
- deleted 10y ago[deleted]
- bogomipz 10y agoIf you look under "stay tuned" it says: "Now that you have a taste of the system that processed and routed the request to this blog post we hope you stay tuned for future posts describing our director design in depth, improving haproxy hot configuration reloads and how we managed to migrate to the new system without anyone noticing." That leads me to believe it involves HAProxy.
- Ianvdl 10y agoGiven the title and the length of the post I was expecting a lot more detail. > Over the last year we’ve developed our new load balancer, called GLB (GitHub Load Balancer). Today, and over the next few weeks, we will be sharing the design and releasing its components as open source software. Is it common practice to do this? Most recent software/framework/service announcements I've read were just a single, longer post with all the details and (where applicable) source code. The only exception I can think of is the Windows Subsystem for Linux (WSL) which was discussed over multiple posts.
- emilburzo 10y agoI don't know if it's common practice, but I've noticed this trend starting to gain momentum. See the product unveiling from Apple, GoPro and probably others that I haven't been following. Hype and slow release, it's the new clickbait.
- deleted 10y ago[deleted]
- falsedan 10y agoeBay released a series of posts on how they manage their Jenkins installations [0][1][2] (part four is missing). I think people choose this pattern when: * engineers implement something cool * management want engineering content to promote the company/drive recruitment * the engineers are pressed for time/not professional technical writers Some bigger companies stagger the context when it's really hefty and makes sense to chunk it up, plus it drives repeat visitors. For smaller companies, the schedule usually slips, and the first post is usually "look how hard this problem is! wow it's really really hard! isn't is amazing that we even tried to fix it? OK see you next time!". Yes, I think GitHub and eBay are small companies. [0]: http://www.technology-ebay.de/the-teams/mobile-de/blog/taming-the-hydra-part-1.html http://www.technology-ebay.de/the-teams/mobile-de/blog/tamin... [1]: http://www.technology-ebay.de/the-teams/mobile-de/blog/taming-the-hydra-part-2.html http://www.technology-ebay.de/the-teams/mobile-de/blog/tamin... [2]: http://www.technology-ebay.de/the-teams/mobile-de/blog/taming-the-hydra-part-3.html http://www.technology-ebay.de/the-teams/mobile-de/blog/tamin...
- 10y ago
- tadelle 10y agoGithub!! Fortress of Open-Source!!! But why they are not open-sourcing GitHub just like GitLab?
- nathankleyn 10y agoThey mention in the post that they will be open-sourcing this code over the coming weeks.
- tadelle 10y agoI'm talking about entire GitHub not components or Load Balancer...
- meddlepal 10y agoWhy do you think they should open source it?
- tadelle 10y agoWhy not? They could release community version at least.
- SteveNuts 10y agoBecause if you say to management "hey, why don't we release the code that our entire business is built off of for anyone to host for themselves for free!" you'd be laughed out of the building on your way to the psych ward.
- gm-conspiracy 10y agoTheir entire business is already built upon software that anyone can host themselves for free.
- 10y ago
- lifeisstillgood 10y agoI love using GitHub and appreciate the impact it is and has had. But this post is what is wrong with the web today. They have taken a distributed-at-it's-plumbing technology, and centralised it so much that now we need to innovate new load balancing mechanisms. Years ago I worked at Demon Internet and we tried to give every dial up user a piece of webspace - just a disk always connected. Almost no one ever used them. But it is what the web is for. Storing your Facebook posts and your git pushes and everything else. No load balancing needed because almost no one reads each repo. The problem is it is easier to drain each of my different things into globally centralised locations, easier for me to just load it up on GitHub than keep my own repo on my cloud server. Easier to post on Facebook than publish myself. But it is beginning to creak. GitHub faces scaling challenges, I am frustrated that some people are on whatsapp and some slack and some telegram, and I cannot track who is talking to me. The web is not meant to be used like this. And it is beginning to show.
- bm1362 10y agoThis reminds me of the criticism of Dropbox [1] when it was first announced on HN - I think you're not the norm. [1] https://news.ycombinator.com/item?id=8863 https://news.ycombinator.com/item?id=8863
- deleted 10y ago[deleted]
- jimjag 10y ago++1!!
- RodericDay 10y agoI started coding 3 years ago, and only started playing around with VPSs a year ago. It's hard to explain how... unaware I was, that I could just have a little "plot of land" on the internet, not managed by anyone other than me (and DO).
- auxbuss 10y agoThis is very enlightening. One of things that everyone on HN think as obvious, yet is far from that to most folk; even those with a fair degree of technical knowledge.
- Scaevolus 10y agoRelated presentations/papers about large scale load balancing: Facebook: https://www.usenix.org/conference/srecon15europe/program/presentation/shuff https://www.usenix.org/conference/srecon15europe/program/pre... Google: http://static.googleusercontent.com/media/research.google.com/en//pubs/archive/44824.pdf http://static.googleusercontent.com/media/research.google.co...
- squiguy7 10y agoI know they mentioned their SYN flood tool but I recently saw a similar project from a hosting provider and thought it was neat [1]. It seems like everyone wants their own solution to this when it is a very common and non-trivial problem. [1]: https://github.com/LTD-Beget/syncookied https://github.com/LTD-Beget/syncookied
- jimjag 10y agoI am increasingly bothered by the "not invented here" syndrome where instead of taking existing projects and enhancing them, in true open source fashion, people instead re-create from scratch. It is then justified that their creation is needed because "no one else has these kinds of problems" but then they open source them as if lots of other people could benefit from it. Why open source something if it has an expected user base of 1? Again, I am not surprised by this. They whole push of Github is not to create a community which works together on a single project in a collaborative, consensus based method, but rather lots of people doing their own thing and only occasionally sharing code. It is no wonder that they follow this meme internally.
- bogomipz 10y agoI'm not sure I agree, they mention HAProxy and Foo over UDP so they are leveraging existing open source technologies. Custom additions to suit ones particular use case isn't necessarily the same as NIH syndrome.
- jimjag 10y agoThey mention HAProxy, but it doesn't look like it's based on it.
- logicalstack 10y agoJoe from GitHub here, we'll talk about it later posts but GLB is based on a number of open source projects including, haproxy, iptables, FoU and pf_ring. Many existing open source solutions are optimized for short lived HTTP requests and don't address the long running connection issue (like a large git clone). We wanted something better for our use case.
- jimjag 10y agoThat is good to know. Thx.
- halifaxbeard 10y ago
- bogomipz 10y agoDo the Directors use Anycast then? That wasn't clear to me.
- jssjr 10y agoAnycast usually implies traffic will be directed to the nearest node advertising that prefix. The GLB directors leverage ECMP which provides the ability to balance flows across many available paths.
- bogomipz 10y agoAnycast and ECMP work together in the context of load balancing. ECMP without Anycasted destination IPs would be pointless for horizontally scaling your LB tier. What Anycast means is just that multiple hosts share the same IP address - as opposed to unicast. When all the nodes sharing the same IP are on the same subnet "nearest" is kind of irrelevant. So the implication is different.
- jssjr 10y agoSure. Feel free to call it anycast then. I usually hear anycast routing used in the context of achieving failover or routing flows to the closest server/POP, but there is probably a more formal definition in an RFC that I'll be pointed to shortly. =) We are using BGP to advertise prefixes for GLB inside the data center to route flows to the directors. In our case all of the nodes are not on the same subnet (or at least not guaranteed to be) which is one of the reasons why we chose to avoid solutions requiring multicast. I expect Joe and Theo will get into more details about that in a future post though.
- bogomipz 10y agoAre you running Quagga or Bird on the director instances then? I'm looking forward to reading more about it.
- logicalstack 10y ago
- deleted 10y ago[deleted]
- alsadi 10y agoI never like github approach, they alway use larger hammers
- NatW 10y agoI'm curious if they looked into pf / CARP as part of their research into allowing horizontal scalability for an ip. See: https://www.openbsd.org/faq/pf/carp.html https://www.openbsd.org/faq/pf/carp.html
- logicalstack 10y agoCARP and similar systems require an active/passive configuration which we did not want since it needs at least twice as many hosts, half of which are not doing any work. We had similar issues with our former Git storage system based on DRDB (http://githubengineering.com/introducing-dgit/ http://githubengineering.com/introducing-dgit/). pfsync, lvs and etc uses multicast to share connection state which we also wanted to avoid.
- treve 10y agoI half expect a comment here explaining why Gitlab does it better ;)
- sytse 10y ago:) We're not doing this better. We're struggling with our load balancers right now. We're using Azure load balancers and then HAproxy. But the Azure ones sometimes don't work. Luckily the new network type on Azure supports floating IPs so we can set something up ourselves https://gitlab.com/gitlab-com/infrastructure/issues/466 https://gitlab.com/gitlab-com/infrastructure/issues/466
- dogismycopilot 10y agoI would love to see the solution. We're also desiring to run HAProxy in Azure with keepalived (even in unicast mode). The black-box "windows based" load balancer that Azure offers is quite limited.
- sytse 10y agoCool, we're trying to do all the infrastructure work in the open under https://gitlab.com/groups/gitlab-com https://gitlab.com/groups/gitlab-com
- pcarranza 10y agoInfrastructure issues are here: https://gitlab.com/gitlab-com/infrastructure https://gitlab.com/gitlab-com/infrastructure The one about the load balancer is https://gitlab.com/gitlab-com/infrastructure/issues/467 https://gitlab.com/gitlab-com/infrastructure/issues/467 but we don't have a lot of data published just yet as we are currently figuring it out.
- tadelle 10y agoWhat about Docker Swarm?
- jdc0589 10y agodo you guys have a blog post or anything talking about why you settled on Azure? I've used both azure and aws a decent amount now, and GitLab falls outside the "right place for the right tool" mentality I've been operating with. Most of our infrastructure is in Azure, but it's also mainly windows (azure has the best windows VM pricing with a moderate was agreement, which isn't surprising). There are some regulatory considerations for us too, buts that's anither conversation.
- gumby 10y agoThey talk about running on "bare metal" but when I followed that link it looked like they were simply running under Ubuntu. Is it so much a given that everything is going to be virtualized? When I think of "bare metal" I think of a single image with disk management, network stack, and what few services they want all running in supervisory mode. Basically the architecture of an embedded system.
- wmf 10y agoYes, it is assumed that all startups are running in EC2 us-east-1 and "bare metal" is the accepted term for non-virtualized systems.
- gumby 10y agoThanks for the clarification. Yet another dilution of a technical term, sigh. I was quite excited when I read it, and felt quite let down when I followed up.
- colemickens 10y agoI don't get it. What else is bare metal meant to mean? "bare metal" = "embedded system"? What does "embedded system" mean then? I guess my age / cloud-nativeness is showing?
- unwind 10y agoIn the embedded space, there often isn't any type of OS or kernel between the application code and the hardware resources ("the metal"). If I want to send out a character through the board's serial debugging port, I don't do an open()/write()/close(), I poke the UART's transmit register. When they said "bare metal", I too thought they ran without OS which had been kind of cool.
- malodyets 10y agoI wondered the same thing: "Wow, they have their own git kernel?" But no.
- deleted 10y ago
- NicoJuicy 10y agoI notice a lot of negativity arround here. Don't know why that is... But i'll take my 5 cents on it. NIH - Not invented here and redoing an opensource project. - Github said they used HAProxy before, i think the use case of github could very well be unique. So they created something that works best for them. They don't have to re-engineer an entire code base. When you work on small projects, you can send a merge request to do changes. I think this is something bigger then just a small bugfix ;). Totally understand them there for creating something new - They used opensource based on number of open source projects including, haproxy, iptables, FoU and pf_ring. That is what opensource is, use opensource to create what suits you best. Every company has some edge cases. I have no doubt that Github has a lot of them ;) Now, Thanks GitHub for sharing, i'll follow up on your posts and hope to learn a couple of new things ;)
- p1mrx 10y agoGitHub only speaks IPv4, so I would be extra-skeptical about using any of their networking code to support a modern service.
- gwright 10y agoWhile I understand that NIH syndrome is a real thing, it is very dissapointing to read many of the comments here. I think very few HN readers are really in a position to have an informed opinion regarding Github's decision to build new piece of software rather than using an existing system. Personally I find this area quite interesting to read about because it is very difficult to build highly available, scalable, and resilient network service endpoints. Plain old TCP/IP isn't really up to the job. Dealing with this without any cooperation from the client side of the connection adds to the difficulty. I look forward to hearing more about GLB.
- jedberg 10y agoAwesome. The whole time I was reading I was thinking "they need Rendezvous hashing". And then bam, last paragraph mentions that is in fact what they are using.
- contingencies 10y agoI am intrigued by their opening statement of multiple POPs, but the lack of multi-POP discussion further in the system description. My understanding is that the likes of, for example, Cloudflare or EC2 have a pretty solid system in place for issuing geoDNS responses (historical latency/bandwidth, ASN or geolocation based DNS responses) to direct random internet clients to a nearby POP. Building such a system is not that difficult, I am fairly confident many of us could do so given some time and hardware funding. Observation #1: No geoDNS strategy. Observation #2: Limited global POPs. Given that the inherently distributed nature of git probably makes providing a multi-pop experience easier than for other companies, I wonder why Github's architecture does not appear to have this licked. Is this a case of missing the forest for the trees?
- yladiz 10y agoI'm of two minds about this. Part of me agrees with many of the commenters here, in that Not Invented Here syndrome was probably in effect during the development of this. I don't really know Github's specific use case, and I don't know the various open source load balancers outside of Haproxy and Nginx, but I would be surprised if their use case hasn't been seen before and can be handled with the current software (with some modification, pull requests, etc.). On the other hand, I would guess Github would research into all of this, contact knowledgeable people in the business, and explore their options before spending resources on making an entirely new load balancer. Maybe it really is difficult to horizontally scale load balancing, or load balance on "commodity hardware". That being said, why introduce a new piece of technology without actually releasing it if you're planning to release it, without giving a firm deadline? This isn't a press release, this is a blog post describing the technical details of the load balancer that is apparently already in production and working, so why not release the source when the technology is introduced?
- madmulita 10y agoWe are in the process of moving all of our infrastructure to OpenStack, OpenShift, Ansible, DevOps, Microservices, Docker, Agile, SDN and what not. There are some brainiacs pushing these magic solutions on us and one of the promises is load balancing is not an issue, even better, it's not even being talked about. Please, please, tell me there's something I'm missing.
- lamontcg 10y agoWhy not just use DNS load balancing over VIPs served by HA pairs of load balancers? Back in the day we did this with Netscalers doing L7 load balancing in clusters, and then Cisco Distributed Directors doing DNS load balancing across those clusters. It can take days/weeks to bleed off connections from a VIP that is in the DNS load balancing, but since you've got an H/A pair of load balancers on every VIP you can fail over and fail back across each pair to do routine maintenance. That worked acceptably for a company with a $10B stock valuation at the time.
- manigandham 10y agoCompany stock value has nothing to do with their scaling, performance and customized processing requirements.
- deleted 10y ago[deleted]
- wtarreau 10y agoDid people really read the article ? For me it was pretty clear, maybe it involves some regular load-balancing terms that people are not familiar with, because I'm seeing a lot of bullshit written in the comments, but here is what is described there : - in a traditional L4/L7 load balancing setup (typically what is described in my very old white paper "making applications scalable with load balancing"), the first layer (L3-4 only, stateless or stateful) is often called the "director". - the second level (L7) necessarily is based on a proxy. For the director part, LVS used to be used a lot over the last decade, but over the last 3-4 years we're seeing ECMP implemented almost in every router and L3 switch, offering approximately the same benefits without adding machines. ECMP has some drawbacks (breaks all connections during maintenance due to stateless hashing). LVS has other drawbacks (requires synchronization, cannot learn previous sessions upon restart, sensitivity to SYN floods). Basically what they did is something between the two for the director, involving consistent hashing to avoid having to deal with connection synchronization without breaking connections during maintenance periods. This way they can hack on their L7 layer (HAProxy) without anyone ever noticing because the L4 layer redistributes the traffic targeting stopped nodes, and only these ones. Thus the new setups is now user->GLB->HAProxy->servers. And I'm very glad to see that people finally attacked the limitations everyone has been suffering from at the director layer, so good job guys!