3 ms·
"Added sqlite3 as the default storage mechanism. etcd3 is still available, but not the default." I want to use and support this product for this reason alone.
by johnmarcus 8y ago
"Added sqlite3 as the default storage mechanism. etcd3 is still available, but not the default."
I want to use and support this product for this reason alone. Etcd was always an un-necessary complexity added for god knows what reason. Later cluster management solutions have abstracted etcd creation and management away (thankfully), but it's always irksome that it is there. Thank you to the K3s development team for taking on that challenge!
- geofft 8y agoAIUI it gets you away from a single point of failure, right? Unless you have a reliable NFS server (and non-SPOF NFS servers are rare and pricy), running k8s on SQLite sounds like you can only have one master. Of course that's totally fine for many k8s deployments and might even increase reliability for some use cases, but still, moving from a distributed system to a local one is a significant change.
- johnmarcus 8y agomeh. Practically speaking, we're talking about a single point of disk failure - which certainly happens but at a rate that is sufficiently low. Plus, the actual amount of data stored is tiny, you can replicate it in seconds. Amazon has solutions for this if it's truly of concern, I would guess google does as well. IMHO, the operation of etcd - and the fact the data became unreadable if you lost quorum - was a much higher risk factor than possible disk failure. It was impractical to backup as well, you either have quorum or you don't. Even without NFS, I could backup that sqlite db every 5 minutes via a cron job and have most of my cluster state perfectly preserved.
- geofft 8y agoDisk failure happens quite frequently for me at scale, but so do other things like RAM going bad or network cards dying or entire mainboards just acting weird or top-of-rack switches silently dropping packets because of memory corruption (all of these have happened to machines I'm responsible for in the last six months). Again I think this is a matter of scale. If you've got enough machines that disk failure is a concern, you can also run a 9-node etcd cluster and have a big enough pager rotation that keeping 5 of them up 99.999% if not 100% of the time isn't a challenge. If you have less than a rack of machines and you're not at the point where you're worried about having a SPOF in your network switches or power supplies, running etcd a bunch of overhead for a problem you don't have and you are genuinely better served by a robust non-distributed system whose availability is the availability of your hardware.
- philips 8y agoAlternatively the K3s authors could have embedded a single node etcd process into Kubernetes using the embed package instead of introducing sqlite. https://godoc.org/github.com/etcd-io/etcd/embed https://godoc.org/github.com/etcd-io/etcd/embed This is something the Kubernetes community might consider as well.
- tasubotadas 8y agoAs a someone that used etcd during v1 times (CoreOS + Fleet) I can only agree. I am not that familiar with etcd3 but v1 and v2 were horrible: difficult to tune and find a working set of "timeout" parameters, picky to CPU availability, developers giving zero f*cks to bugs/docs, crappy documentation, and good luck if you lose quorum.