5 ms·
(HashiCorp Co-Founder) First, I love this blog post and I tweeted about it, so please don’t interpret any of my feedback here negatively. > If fly.io found th
by mitchellh 5y ago
(HashiCorp Co-Founder)
First, I love this blog post and I tweeted about it, so please don’t interpret any of my feedback here negatively.
> If fly.io found the upper bound on a Consul scale out, what do you think a reasonable threshold looks like for a smaller system?
I don’t know the exact scale of Fly.io’s Consul usage, but I would imagine they’re far, far from the upper bound of Consul scale out. We have documented some exact scale numbers here[1]. And of course these are only the customers we can talk about.
I didn’t read this post as talking about scale limits. Instead, it discusses the tradeoffs of certain Consul deployment patterns, and considers whether their particular usage of Consul is the best way to use a system with the properties that Consul has. And this is a really important question that everyone should be asking about all the software they use! I appreciate Fly sharing their approach.
To answer your second part (what is a reasonable threshold), we have documented recommended hardware requirements and scale limits here: https://learn.hashicorp.com/tutorials/consul/reference-architecture https://learn.hashicorp.com/tutorials/consul/reference-archi... This is the same information we share with our paying customers.
[1]: https://www.hashicorp.com/case-studies https://www.hashicorp.com/case-studies
- mrkurt 5y agoThis is true. We did not have scaling issues with Consul. It scaled very well from "nugget of an idea" to "whoah maybe this is going to be a big company". I think the best software buys time for the users. Consul bought us years.
- chrisweekly 5y agobug company -> big company
- tptacek 5y agoI don't think we're anywhere close to the limit of Consul's ability to scale, but I think we're abusing Consul's API. If I had to pinpoint an "original sin" of how we use the Hashistack here, it's that we need individual per-instance metadata for literally all our services, everywhere. I can't imagine that's how normal teams use Consul.
- datalopers 5y agoPurely as an ignorant outsider here, but now I've seen Roblox and Fly.io have either crippling outages or an inability to scale due to issues in Consul. It's not a good look.
- ikiris 5y agoDo you also blame guns when people shoot themselves in the foot when they kept it loaded and the safety off?
- datalopers 5y agoFly.io committed a bug fix back to Consul, and Roblox’s 3-day outage was due to flaws in Consul streaming.
- sammy2244 5y ago
- oceanplexian 5y agoI worked at a place that ran Federated Consul in 60 DC's across ~40,000 machines running consul-agent. Originally, the largest DCs had about 8,000 nodes which caused some problems that we had to work through. But I'm of the thought that you shouldn't have 8,000 of anything in a single "datacenter" without some kind of partition to reduce blast radius.
- politician 5y agoHow many people were dedicated to keeping that configuration running?
- politician 5y ago> To answer your second part (what is a reasonable threshold), we have documented recommended hardware requirements and scale limits here: That's (really) good documentation, but doesn't directly address the Fly.io situation nor my situation: multiple data-centers in multiple jurisdictions around the globe. > https://learn.hashicorp.com/tutorials/consul/federation-gossip-wan https://learn.hashicorp.com/tutorials/consul/federation-goss... > To start with, we have a single Consul namespace, and a single global Consul cluster. This seems nuts. You can federate Consul. But every Fly.io data center needs details for every app running on the planet! Federating costs us the global KV store. We can engineer around that, but then we might as well not use Consul at all. I think a better way to ask my question might be: Is there a threshold below which can we safely run Consul in a single global cluster like Fly.io tried before it got too unwieldy?