6 ms·
Lots of good information here but it is still not enough for a production setup IMHO. There is a great need of good source of production setups. In open source
by anemic 10y ago
Lots of good information here but it is still not enough for a production setup IMHO.
There is a great need of good source of production setups. In open source software this seems to be the secret that no one is willing to reveal. I tried to setup a kubernetes cluster from scratch a while ago and soon I was browsing the source code for answers. Openstack is the same, you need to understand a lot about the inner workings before you can even attempt to setup something for production.
There is always a simple "this shell script starts your own <name your tech here> cluster in vagrant" but it is still not a production setup.
And even if this article is the "hard way" it describes:
> This is being done for ease of use. In production you should strongly consider generating individual TLS certificates for each component."
And it does not mention that the crucial part is the common name field in the certificate maps to the user name that is the magic information that I once needed.
I sincerely appreciate this article but production setup is still a long learning experience away.
- olalonde 10y agoI just used `./cluster/kube-up.sh` to setup my cluster on AWS. I am now wondering what's missing for a production setup. It seems to be working OK so far (though I just have 3 minions and a few pods). One thing I wish I knew how to do is how to safely upgrade the cluster without re-creating it from scratch. Care to elaborate a bit?
- lobster_johnson 10y agoThe problem is that it's not declarative. You can't tweak the config and run it again to converge. Kubeup is designed to run once, unlike systems such as Puppet and Terraform that declaratively set up the world to fit your specification. Kubeup also does a lot of mysterious stuff. By using it, you don't have a clear idea of which pieces have been set up and how they slot into each other. It is, in short, opaque and magical. For comparison, I set up Kubernetes with Salt on AWS. It was, by all means, "the hard way", and took me a few days to get running and a couple of weeks to run completely correctly (a lot of stuff, like kubeconfig and TLS behaviour, is still undocumented), but as a byproduct I now have the entire setup in a reproducible, self-documenting, version-controlled config.
- DigitalJack 10y agoHave you by chance open sourced your setup? I started going down this route with terraform, but ended up stopping and just using the kube-up script due to time constraints. However, now that I have a cluster up and running, I can take the time to build a parallel cluster with more understanding, and migrate the services to it. I found a terraform example, but it declares itself out of date, and looked more complicated than I thought it should be... that was just a gut feel though. I have not used Salt, but I always like to learn new things, especially if they make my life easier. I'd be very interested in checkout out your setup, and any lessons learned you have to share. Thanks in advance
- lobster_johnson 10y agoI haven't, but I would be happy to. I just need to generalize it a little bit. Email me and I will send you a link once it's done, sometime next week (I'm on vacation).
- nathan_f77 10y agoHi, I would just like to second the request for an open-source, production-ready implementation. They seem to be rarely shared, which is a big shame. I've been playing with https://github.com/kz8s/tack https://github.com/kz8s/tack lately, but your implemention sounds like even more comprehensive.
- marcoceppi 10y agoWe've opensourced our setup: http://www.jorgecastro.org/2016/07/29/ubuntu-kubernetes-v1-dot-3-3-ready-for-testing/ http://www.jorgecastro.org/2016/07/29/ubuntu-kubernetes-v1-d...
- DigitalJack 10y agoThis project looks very interesting, and once I'm more comfortable at a lower level of abstraction, I may use this.
- 10y ago
- timtadh 10y agoHave you tried shooting a node in the head and seeing what happens? Always a good exercise to run. Run a few disaster recovery exercises and see if you can get it back. I recommend doing that on non-production of course!
- olalonde 10y agoThanks for the tip. I did yesterday actually by manually shutting down the node from SSH (sudo shutdown). It seemed to "just work" without having to do anything else. There might have been a tiny period of unavailability to one of my services but not enough for me to notice. Luckily, I don't have crazy high availability requirements yet.
- anemic 10y agoThat is perfectly fine if it suits your use case! I have to deal with industry certifications and unfortunately using a ad-hoc certificate authority is not an option or running in insecure ports. Also I was setting it up on coreos and baremetal servers. It should be possible to run pods in google container engine or similar very easily, but would there be any fun in that?
- olalonde 10y agoRight, no industry certification to follow here and pretty loose availability requirements. You just had me worried for a minute that everything would suddenly grind to an hal or that there were glaring security holes! But my "production" requirements are definitely not as strong as yours.
- philips 10y agoThe CoreOS team has worked a ton to improve the baremetal installation experience for Kubernetes. You can read more about it here: https://coreos.com/kubernetes/docs/latest/kubernetes-on-baremetal.html https://coreos.com/kubernetes/docs/latest/kubernetes-on-bare... And the installation flow that builds on top of that for Tectonic: https://tectonic.com/blog/tectonic-1-3-release.html https://tectonic.com/blog/tectonic-1-3-release.html
- meta_AU 10y agoIf you move etcd to separate t2.nano then upgrading is easy. The only stateful part is etcd.
- ChartsNGraffs 10y agoIf you're looking for something that is a little more flexible for deploying Kubernetes, I recommend either KOPS[1] or kube-aws[2]. kube-aws is tethered to AWS but is much more flexible than the standard kube-up.sh script. KOPS is the heaviest lifting tool I've found for deploying Kubernetes. It's short for Kubernetes Ops and (I believe) it can even generate Teraform configs so you can get the upgrades without re-creating everything. [1] https://github.com/kubernetes/kops https://github.com/kubernetes/kops [2] https://github.com/coreos/coreos-kubernetes/tree/master/multi-node/aws https://github.com/coreos/coreos-kubernetes/tree/master/mult...
- roman_sf 10y ago(Currently work in progress, but working. Some of these statements are forward-looking.) Lol
- Rapzid 10y agoKOPS is pretty awesome in that: * Actually pretty much works for what's in scope.. * It's got some nice configuration options that are discoverable and not hidden away in envars... * Some good prelim docs explaining how kubernetes is bootstrapped * Cluster management seems to function properly * Updating/upgrading What's missing IMHO(from an AWS user's standpoint including kops and k8s): * SUPER unapproachable codebase ATM for KOPS and friends * More flexible cluster dns naming so we can leverage real wildcard certs accross dev environments * Running kubernetes in private networks * Passing in existing networks created through other tools(terraform, cloudformation, custom etc) * Responsibility for stuff seems spread out across projects and is unclear which lies where(also leading to an unapproachable-ness for contributions) * AWS controllers that don't seem to fully leverage the AWS API's (traffic balanced to all nodes and then proxy'd via kube proxy; no autoscale life cycle event hooks) * Unclear situation on the status of ingress controllers; are they even in use now or is it all the old way?! * No audit trails * IAM roles for pods * Stuff I'm probably missing It's very frustrating TBH. On one hand AWS ECS has IAM roles for containers now, for the new Application Loadbalancer, and private subnet support. On the other hand they DON't have pet sets, automatic EBS volume mounting(WTF), a secrets store, configuration API, etc. Also frustrating is I feel the barrier to contribute is a too high ATM even though I have the skills necessary.. It's SO close though. If I can get private, existing subnet support I can probably start running auto provisioned clusters that are of use for some of our ancillary services in production. From there I might be able to help contribute to KOPS and AWS controllers. Right now it looks like there is just this one guy doing most of the work on AWS and KOPS; probably quite overloaded.
- pat2man 10y agoRedHat has a production Ansible script for OpenShift: https://github.com/openshift/openshift-ansible https://github.com/openshift/openshift-ansible most of which is Kubernetes setup. Its amazing how much actually goes on here. Firewall rules, certificates, Docker storage configuration, etc. Its definitely not something that you can just thrown in a VM and assume everything will work.
- zwischenzug 10y agoThere is also a chef version a colleague of mine wrote: https://github.com/IshentRas/cookbook-openshift3 https://github.com/IshentRas/cookbook-openshift3
- anemic 10y agoansible/puppet/chef is the way to go. Single line of DSL is worth 1000 lines of text on a blog, imho.
- atsaloli 10y agoCFEngine, too! :) I'm still in business consulting and training on CFEngine. Although I'm branching out into Git as that has a wider user base.
- timtadh 10y agoDon't get me started on setting up production hdfs and hadoop that is a nightmare. It is comparatively well documented with lots of people running it! As soon as you get into the weeds with kerberos and HA mode forget about the documentation explaining anything properly. Cargo culting from random blog posts, reading the source, and just playing around with config files is the name of the game. There are some weird interactions between KDC settings and some of the daemons that are not documented at all.
- anemic 10y agoSpot on! Once you get everything running you're so exhausted documenting it by writing a blog post is not on your mind, sleeping for a week is...
- zwischenzug 10y agoOh God yes, this nonsense is most of my life - kerberos, ad, and my favourite un covered area of enterprise integration, storage. Nfs4 plus Kerberos anyone?
- zaptheimpaler 10y agoThis has been my experience with Spark too! On one hand, its exciting working on new things changing so fast that the "best way" isn't common knowledge yet. On the other hand, having to grep through source to find out what a config option really does is just painful.
- deleted 10y ago[deleted]
- marcoceppi 10y agoHave you tried to juju to set up things like hdfs and hadoop and the things that plug into it?
- kwmonroe 10y agoI'm happy to see marcoceppi mentioning juju here - i'm one of the enablers of juju big data. We've worked really hard to make it simple to stand up hadoop on clouds, containers, and metal (https://jujucharms.com/hadoop-processing/ https://jujucharms.com/hadoop-processing/). Juju brings the modeling, Bigtop brings the core apps. Scaling, observing, and integrating are old news; HA is landing now; your post and others like it have put security on our -next radar. Having read a few kdc/hdfs stories, i think i'm going to miss the days when dfs.permissions.enabled was good enough ;)
- sytse 10y agoI think that not having a good production setup for an open source project is a combination of a couple of things: 1. Documentation is not satisfying work, maybe because it is as absolute as code? 2. Contributing documentation doesn't get you as much recognition as code 3. If you set it up differently yourself there is no need to maintain a fork (unlike code changes) 4. For open source projects that are company backed there is a perverse incentive to keep the documentation vague if they only make money though support
- mariojv 10y agoOpenStack has an attempt at Ansible playbooks for production deploys: https://github.com/openstack/openstack-ansible https://github.com/openstack/openstack-ansible I have no idea how effective they are, but it's something to consider when setting up a deployment from scratch.
- continuational 10y agoI propose a new term: Consultancy Driven Development. It goes like this: - If it's too easy to set up, nobody will hire us to make it work. - Implement a kickass setup dirt cheap for some big-name company, so we can claim they use it in production. Yeah we tweaked it so it bears little resemblance to the original product, and only fits an incredibly narrow use case, but nobody stands to benefit from blogging about that. - Better ship with a configuration file that isn't production ready. - Did I say one? Better have three configuration files, each duplicated in distribution-dependent directories (in some cases), needing manual sync between servers to prevent catastrophic data loss. - Remember not to publish the program that checks for errors in configuration; half of our income would disappear. - Benchmark with a configuration file that nobody would use in production, but looks really impressive when taken out of context. - People want transactions, remember to claim support (and if you must, explain somewhere in the fine print that a transaction can only span a single operation on exactly one document, and btw. is precisely none of A, C, I or D). - Somewhere on the front page, it should say how we can support petabytes of data (and it performs very well, as long as you write all your data in one batch, never modify it, keep it all in memory, and turn off persistence). - Never give away answers online. Answer every question about configuration with "it depends". - Don't release a new version without renaming a few configuration options. Be "backwards compatible" by ignoring unknown, obsolete and misspelled options.
- sytse 10y agoI agree there is perverse incentive for open source companies to do this if their business model depends on support. Freely after Upton Sinclair: It is difficult to get a person to document something, when his salary depends upon his not documenting it!
- ymse 10y agoWhat you are describing is basically Openstack. Although Hanlons razor applies, none of the current actors stands to benefit from improving the situation. * Extremely difficult to set up. * Claims that half of Fortune 100 uses it (read: many are required to support it; the rest have one guy with a toy installation in some branch office). * Consists of dozens of components, each with several-thousand lines config files (actually Python code) that must be kept in sync between all nodes (yet have node-specific data). * Claims to be "modular", but have complex interdependencies between each of the components. * Upgrading is not officially supported, but some companies will help you. * Will break in mysterious ways, and require you to backport bugfixes since you're stuck on an unsupported version after a year. * Have unhelpful error messages (e.g. throw Connection Refused exception when you're actually receiving an unexpected HTTP return code). * Write documentation in a way that appears OK to new users, but vague enough to be useless for those who are looking for specific information.
- hosh 10y agoI set up a K8S cluster from scratch using the CoreOS tutorials and several other articles (like Kubernetes From the Ground Up series). What's missing is this: 1. Architecture for your specific needs. This is a design process and not easily condensed into a tutorial. It still requires critical thinking on the part of the designer. 2. How the different components fit together and why they matter. For (1) to be commoditized, there needs to be sufficient number of installs where people try different things and come up with a best practices that the community discovers. There are not enough of that for that to take place. For example, I put brought up a production cluster on AWS. I also had to decide how this was all going to interact with AWS VPS and availability zones. How do I get AWS ELB to talk to the cluster? The automated scripts are only the starting point because they assume a certain setup, and I wanted to know what the consequences of those are. This is where the consultants and systems design comes in. On the other hand, Kesley Hightower probably has a lot of that knowledge in his head. By getting it out there, I think more people will try this, understand the principles, and collectively, we'll start seeing these best practices emerge. Maybe Hightower will eventually write a book. In the meantime, if you want to know how to design and deploy a custom setup, you do need to know the building blocks, how they are put together, and how you can compose them in a way for specific use-cases. It's no different than choosing a framework, like Rails or Phoenix, and then learning how to compose things that the framework offers you in order to do what you need to do. You get that knowledge from playing with it and experimenting. Having said all of that, while I'm glad I do have a good foundation for Kubernetes, if I want to use it in production, I'm probably just going to use GKE.
- parasubvert 10y agoHe is writing a book: http://shop.oreilly.com/product/mobile/0636920043874.do http://shop.oreilly.com/product/mobile/0636920043874.do
- philips 10y agoThere is a ton of work going on upstream to make Kubernetes easier to install and manage in production environments. A big chunk of that work is what is being called "Self-Hosted Kubernetes". The idea is that once you bring up a single machine running a Kubelet you can bootstrap the other services that make up a Kubernetes cluster from there. You can learn more about that here: https://coreos.com/blog/self-hosted-kubernetes.html https://coreos.com/blog/self-hosted-kubernetes.html As far as TLS there is ongoing work upstream to add a CSR system for the "agents" called Kubelets. This will allow people to automate the TLS setup and simplify the management. Details are tracked here: https://github.com/kubernetes/features/issues/43 https://github.com/kubernetes/features/issues/43 Also, there are more discussions happening to improve the first install experience. https://github.com/kubernetes/kubernetes/pull/30361#issuecomment-240494793 https://github.com/kubernetes/kubernetes/pull/30361#issuecom... https://github.com/kubernetes/kubernetes/pull/30360 https://github.com/kubernetes/kubernetes/pull/30360 Kubernetes is really focused on not just making it easy to install. Which is the trivial scripting part, as you point out. But, to make Kubernetes easy to manage over the lifecycle of the cluster. Which is where work like self-hosted, TLS bootstrap, etc start to come in.
- willejs 10y agoThis all makes me laugh, and cry at the same time. It makes me laugh because everyone wants to run k8s for no real reason, they havent got scale, traffic or many woes. Please just run some vms, CM, unattended upgrades, capistrano and packer. Mostly the loose reasoning is 'simplicity', and its the new shiny. This is perceived by people thinking that deployment, service discovery, config etc. is provided for free in kubernetes, and one boot script will solve all. On top, everyone thinks its trivial to operate this, maintain it, and no one understands what 'production ready is'. I almost think people think it replaces ops, but it does the opposite. It makes me cry because, running k8s is hard, ops is hard, and so is telling people they might be wrong. K8s consists of half a dozen components, they have dozens of config flags, and much functionality is buggy, in beta, or flux. To top this off k8s is based on etcd. Etcd is barely production ready by their own admissions (remember /production.md in github?) but if you have run it you will understand the bugs, and vague docs coupled with reading the source constantly when problems arise. K8s consists of many components, kubelet, proxy, controller, scheduler and more. These you have to install and configure, and many scripts do this badly in a one size fits all approach, and many CM methods do this barely in an ok manner currently. I cry, because of overlay networking too, its a nightmare, and the alternative cloud permissions are scary.
- jdubs 10y agok8s is still way easier to run than mesos & marathon with all the accompanying service that make it usable.
- SEJeff 10y agoAgreed (as someone who has spent the last 1.5 years working on/with Mesos and will soon be looking to migrate to openshift/k8s)
- leetrout 10y agoAnd Nomad is an order of magnitude easier than k8s. Might not be as full featured but for basic use cases of give me X resources to run Y it's really great. I was shocked at how easy it is to setup and run. The folks at Hashicorp are doing some great things.