5 ms·
I'm in kind of the opposite position, as I feel confused about why I keep seeing so much resistance to higher levels of abstraction here. > It boggles my mind
by tene 5y ago
I'm in kind of the opposite position, as I feel confused about why I keep seeing so much resistance to higher levels of abstraction here.
> It boggles my mind that people look at something like Kubernetes and decide "You know what? We need more layers. On top of this."
I'm not involved with this project at all, but yes, I look at the massive Kubernetes deployments I've worked with, and conclude that toil and various other kinds of problems could be reduced with higher-level abstractions for declaratively managing configuration for all of these clusters.
If you wanted to run workloads on 100k servers, how would you do it? Would you have a team manually configure each cluster individually, or would you consider that there are a lot of similarities between clusters, and you only want differences that are chosen intentionally, and look into some abstraction to help keep the complexity down?
If 100k isn't enough for you, is there any number of nodes at which you'd consider building some tools support for your work?
What about data center count and change rate? There are a lot of business models that benefit from running workloads closer to customers, so it's useful to have deployments in many datacenters and clouds across the world, and for this to grow over time. Do you really not see why some people are interested in being able to bring up a cluster in a new DC by updating a few configs rather than starting from nothing by hand with kubeadm init?
I apologize if I've misrepresented your position by misunderstanding something. I feel like I'm missing or misunderstanding something from your perspective, but I don't know what it is.
More generally, I think it's extremely normal for people to build higher levels of abstraction as things grow. It's usually worth considering abstraction or architectural changes in a system every time it grows by about 10x.
Some of the most painful platforms I've supported in my career have been those that didn't do this, and instead tried to grow through brute force. Something can be fine to do once, troublesome to do 10 times, and overwhelming to do 100 times.
- smichel17 5y agoI think the comment you are replying to was objecting to the number of layers, rather than the level, of abstraction. More concretely, I'd guess they would not object to a tool that replaces k8s and offers a higher level api, but they do object to building said tool on top of k8s. For myself, I think it depends on how leaky the abstractions are. The leakier, the fewer layers the house of cards can stand. That is, it's more about where you slice it than how many times.
- throwaway894345 5y agoThat’s something that comes in time, if ever. If Kubernetes started by replacing the OS later of abstraction, say by requiring unikernel applications, then it would have virtually no adoption because we as an industry don’t know how to build and operate unikernels in production. You have to meet the market/industry where it’s at and then co-evolve with the industry into some tighter, more optimal form.
- tene 5y agoI mean, I agree in that I'd like such a tool too, but given that we don't have this hypothetical more-scalable kubernetes, how is it such an unimaginably shitty idea to build tooling for declarative configuration of clusters? As you say, there's still a fundamental amount of complexity involved in operating a system, and the best abstractions limit the amount you need to care about in a given context. Something I like about declarative configuration specifically is how well it helps you move information from ephemeral human memories and habits into something more-reliable. Most places I've worked usually had some weird ephemeral things, oral history about what needs to be treated differently, many systems that needed conversations to understand, etc. The more you limit your use of global mutable state (interactively changing things in production), the less room there is for important stuff to live in people's heads. When I can spin up on a new service just by learning how a tool works and reading their configs, then I have all of the "what" and "how" and "where" there to look up at any time in a consistent way. There's much less room for weird state or needing to do rituals or broken staircases or whatever. When I talk with people, I can focus much more on the "why" questions. I might be more sensitive to this than some people, because I've got a shitty memory, but the more I'm able to just check a thing the better. Theoretically documentation can fill a lot of the same role, but I've never worked anywhere with consistently up-to-date and comprehensive documentation for what's deployed. When it's the only way to deploy at all, it's no longer optional. It feels like the same sort of thing as a good static type system to me. You can do local reasoning with just what you see, and limit unexpected action at a distance. The types are mandatory, and machine-checked, so you can't get them wrong in certain ways, and you can't just skip writing them like you do sometimes for unit tests that would cover similar verification otherwise.
- windexh8er 5y ago> If you wanted to run workloads on 100k servers... But... Most don't run 100k servers. Most don't run 10k servers. I was working with a very large consumer insurance company recently and while they are really large for their segment, they're well under 15k servers in production. And containers aren't much of a thing. There's a lot of overengineered architecture going into not-so-hyperscaler/FAANG production environments compared to what Kubernetes' original use case was designed to solve for. The common consequence appears to be increased complexity and less operational system knowledge .
- tene 5y agoI don't follow where you're trying to go here, or how you intend this to contradict anything I said. You're right that most people don't run 100k servers. In fact, most people don't run any servers at all, and don't use kubernetes. Some people do, and some of those people use approaches like this. The comment I was responding to, by my reading, seems to express incredulity that anyone would do this, and seems to imply that this is a bad, counterproductive idea that nobody should ever implement. The most notable quote is "I feel like I landed in crazy-land". Do you read it differently from me? My intent in my reply was to describe some of the use-cases where this is helpful. It's not only FAANG that run more than a small handful of clusters; there are plenty of smaller businesses that need to run large numbers of clusters. I don't know anything about the company that wrote this article, but I found https://zitadel.ch/usecases https://zitadel.ch/usecases on their website, and given that description, it sounds quite plausible that they run many small clusters in many clouds and datacenters. They probably don't need a ton of compute, but I wouldn't be surprised if they had latency needs for being close to their customers, or other specialized requirements that motivate dedicated clusters. I also don't know anything about the insurance industry, so I'm curious to hear if I've guessed incorrectly, but it doesn't seem like the kind of business with especially high compute needs. I believe you when you say that your company has no use for this. Can you believe me when I say that there are genuine problems these systems are trying to address? I keep hearing about companies who fund their employees building unnecessary pointless over-engineered production automation, but I have yet to ever encounter one in real life. Every SRE job I've had so far has been on a team that really wanted to invest time in improving infrastructure and automation, but couldn't get time away from the toil for it. If anyone can recommend a company that over-invests in infrastructure automation and architecture instead of under-invests, I'd dearly love to try it and see if it's as bad as I've heard. Even without high compute needs or large numbers of clusters, there's still a lot of benefits you can get when you can move things to declarative configurations and away from global mutable state and human minds. I agree that this can go wrong, but that doesn't mean it's categorically worthless to try. On the other hand, if you're just looking to gripe about it having gone wrong for you, could you share some war stories? I may be feeling idealistic from frustration with low investment in tools support and production automation lately, so maybe I could use some horror stories of it going badly to scare me straight. :)