3 ms·
I see issues with going full on kubernetes-as-controlplane due to how operators do reconciliation. The represented resources are more or less eventually consist
by twalla 3y ago
I see issues with going full on kubernetes-as-controlplane due to how operators do reconciliation. The represented resources are more or less eventually consistent which present challenges if you have workflows that need explicit orchestration and strong requirements for certain dependencies to be met before progressing. If the logic for a particular operator (quality varies here) doesn't handle missing pre-reqs or downstream resources being in a weird state gracefully you end up having to chase down tons of events and logs that are distributed across many components with no real way to quickly and effectively pinpoint where things went sideways (this comes with experience for most human cluster operators but most application developers are not going to know you need to go check the logs of these 5 controllers and check the status of their attendant custom resources to see what needs unwedging)
I understand Flux has some capabilities to pause or perform incremental deployment until certain dependencies are met, which is a step in the right direction.
- akhayam 3y agoI hear you, and am not advocating a full, one-size-fits-all, controller solution for control plane and workers. My take was mainly on using operators for short-cycle control loops inside workers which are inherently simpler in their business logic. I do agree with you on deployment systems taking up more responsibility in understanding app properties and abstracting infrastructure away from app developers. However, it's notoriously hard to get deployment systems to do this correctly. At Amazon, this was a somewhat solved problem but took their best brains, many years of pain, and multiple iterations to get right. If you haven't already, I would recommend reading this thread from Joe Magerramov where he talks about the challenges of deployment systems and how many stars need to line up to get deployments done right: https://twitter.com/_joemag_/status/1587283479448150016 https://twitter.com/_joemag_/status/1587283479448150016. Both Flux and Argo are doing splendid work to solve this problem. I just wish we can democratize the known wisdom, failures and patterns/anti-patterns from cloud operators, so others don't have to repeat the same mistakes. I'll see if I can put together some blog posts on this topic.