19 ms·
Introducing Empire: A Self-Hosted PaaS Built on Docker and Amazon ECS
- justinsb 11y agoThis looks great: simple yet powerful. I'm working a lot with Kubernetes, and you don't actually need to run an overlay network on AWS (or GCE). On AWS, there's some VPC magic that surprised me when I first saw it! But I believe that's beside the point; it's not about ECS vs Kubernetes, it is about what we can build on top. In particular, I think the idea of embedding a Procfile in a Docker image is really clever; it neatly solves the problem of how to distribute the metadata about how to run an image.
- ejholmes 11y agoExactly! One of our goals was to also make the scheduling backend pluggable, so we're hoping that the community will implement a Kubernetes backend in the future. There's a lot of similar concepts between both, but we ultimately chose ECS for it's ease of operation and the integration with existing AWS services like ELB.
- fosk 11y agoThis is neat. You might want to check out KONG (https://github.com/Mashape/kong https://github.com/Mashape/kong) instead of putting a plain nginx in front of the containers/microservices. It is built on top of nginx too, but it provides all the extra functionality like rate-limiting and authentication via plugins.
- steinnes 11y agoGood thought, I even think the article mentions kong as a future replacement for nginx :-)
- ejholmes 11y agoYep! We're definitely looking into Kong in the future. For now nginx + some static configuration works amazingly well for us. Very maintainable and minimizes external dependencies.
- steinnes 11y agoI can identify with that, we built a custom nginx based load balancing / routing solution at QuizUp while I was there :-)
- moatra 11y agoKONG definitely looks interesting, and I'd love to know more about it. However, there's definitely not a lot written about it yet. For example: I've gone searching through the blog posts, github readme, and KONG documentation, but I still have no idea _why_ it needs Cassandra. What does it store in there?
- jkarneges 11y agoKong uses Cassandra for storing config. This makes it easy to run a Kong cluster. Just add more instances that share the same Cassandra cluster.
- moatra 11y agoIs rate limiting state stored in Cassandra? One of the main graphics on the KONG docs shows a Caching plugin (http://getkong.org/assets/images/homepage/diagram-right.png http://getkong.org/assets/images/homepage/diagram-right.png), but the list of available plugins doesn't include such an entry. Is that because caching is built in? Is the cache state stored in Cassandra? Or is the plugin yet to be built?
- fosk 11y agoAll the data that Kong stores (including rate-limiting data, consumers, etc) is being saved into Cassandra. nginx has a simple in-memory cache, but it can only be shared across workers on the same instance, so in order to scale Kong horizontally by adding more servers there must be a third-party datastore (in this case Cassandra) that stores and serves the data to the cluster. Kong supports a simple caching mechanism that's basically the one that nginx supports. We are planning to add a more complex Caching plugin that will store data into Cassandra as well, and will make the cached items available across the cluster.
- sinzone 11y agoWorth mentioning KongDB, to easily provision a cloud Cassandra instance for free: http://kongdb.org http://kongdb.org
- phantom_oracle 11y agohttp://www.openshift.org/ http://www.openshift.org/ Just putting this out there in case anyone is looking for an alternate open-source PaaS. I've never personally used it before (self-hosted), but it may be something that someone out there is looking for.
- sagivo 11y agopersonally i use dokku (https://github.com/progrium/dokku https://github.com/progrium/dokku). i would be happy to see one standard "heroku-like" paass since i feel too many people trying to tackle the same problem.
- dstroot 11y agoAnd there is dokku-alt. https://github.com/dokku-alt/dokku-alt https://github.com/dokku-alt/dokku-alt
- dubcanada 11y agoThat project is not really maintained anymore, and it is nothing more then dokku with a bunch of plugins preinstalled. I would suggest just using dokku and installing the plugins you need. Dokku has moved forward in their API and stuff a lot since dokku-alt was forked and the documentation for dokku does not really apply to dokku-alt.
- helloiamaperson 11y agohttp://www.openshift.org/ http://www.openshift.org/ or http://cloudfoundry.org/index.html http://cloudfoundry.org/index.html, both used in production by f500 companies
- Veratyr 11y agoCloud Foundry requires a ton of (compute) overhead to get set up. It's very much intended for large projects. Hell, you need a separate VM just to install it. For anyone looking at a Dokku alternative, Cloud Foundry isn't one. Openshift is nice though.
- helloiamaperson 11y ago> Hell, you need a separate VM just to install it. I'm not sure why that's a problem. If you want something that's actually like heroku like in terms of uptime and what not, you need something that can manage the health of the cluster. Dokku's cool, but it doesn't make sense for anything you actually need to depend on. If it doesn't make sense to pay for the overhead of running your own paas, just use Heroku instead.
- stephenr 11y agoDoes this really classify as "self hosted" if it's heavily dependent on AWS?
- mentat 11y agoIt runs inside networks you control at least at the configuration level, so I would say so.
- stephenr 11y agoAs I explained in a sibling thread, I would consider "self host{ed,able}" to mean that you can run it on an arbitrary machine (virtual and/or physical) regardless of provider (e.g. this is AWS specific, so I can't run it on physical hardware I own/rent or even on a competing provider of virtual machines)
- nickpsecurity 11y agoI don't think so. It's their hardware, infrastructure, and engineers hosting it. They control those things. You get the rest. Sounds like an AWS-hosted solution with some advertised advantages over other solutions. Definitely not self-hosted. Note: I think the only 3rd party thing I'd call self-hosted is colocation where I delivered the server, they plugged it in, and the most they do is reboot it for me.
- stephenr 11y agoFrom the point of view of software, I generally consider something self-host{ed,able} if I can run it on a machine I choose, without enforced network/environment requirements.
- nickpsecurity 11y agoIt's a fair viewpoint. I guess my critical point is control: control over the hardware, its software, legal rights to it, and so on. If they're in control, it's theirs. How can it be myself if outsiders control or own it? I guess a combo of philosophical and legal.
- floridaguy01 11y agoaws is silly expensive. Why didnt you build this on top of digitalocean? Digitalocean is so awesome right now. They dont even charge for bandwidth overages.
- virulent 11y agoHaha, except AWS doesn't lock your account randomly or stop droplets for benign abuse reports. Also, OP required Redshift. DO does not offer that.
- davidbanham 11y agoDO is great for a lot of things, but it's not AWS. You can't allocate extra disk to a droplet, for example. AWS is a _much_ more complete offering than DO.
- tracker1 11y agoAgreed... AWS and Azure both offer a lot of services beyond just VPS hosting. Hosted database services, and extended blob/s3 storage are pretty valuable in and of themselves. DO/Linode don't offer the equivalent, which means maintaining your own.. which is fine, but if you're relatively small, or a single person... time you dedicate to operations tasks is time you aren't developing features and/or fixing bugs. One's business is paramount... technology is just a tool to serve that.
- serferfish 11y agoDigitalocean is silly expensive. Why don't you look at Atlantic.net they are so awesome and charge way less than overpriced digitalocean. Why not run it on your laptop which you've already paid for, that would be even cheaper than overpriced atlantic.net!
- dubcanada 11y agoUnless I am missing something Atlantic.net is 1-10 cents cheaper then digitalocean? https://www.vultr.com/pricing/ https://www.vultr.com/pricing/ is 20% cheaper right now at least.
- jordanthoms 11y agoHow do you handle running one-off tasks (consoles, migrations etc) on this setup? This is something most of these systems seem to ignore...
- ejholmes 11y agoWe actually have a relay (https://github.com/remind101/empire/tree/master/relay https://github.com/remind101/empire/tree/master/relay) service that can be run alongside Empire that acts as a proxy to interactive Docker sessions. It's a bit of an experiment right now and something we'd like to solve better in the future, but it allows you to run containers with `emp run <command> -a <app>`.
- jordanthoms 11y agoInteresting! We are in a similar situation to where you were, with a bigish app on Heroku which we are keen to move over to EC2 to join the rest of our infrastructure, definitely keen to see how empire develops
- tracker1 11y agoInteresting... except for being limited to Node.js initially (now JVM too), I would think AWS Lambda would be almost ideal for this.
- rymohr 11y agoThank you, this looks awesome! As someone who still hasn't embraced docker due to all the orchestration / discovery madness I really appreciate such an elegant solution. I love and run everything on AWS so building on top of ECS is just another selling point.
- nickpsecurity 11y agoThis work has plenty about it that was interesting. The best part to me was their answer to "why not feature X?" They said they prefer to build upon the most mature and stable technologies along with naming a few. Too many teams end up losing competitiveness by wasting precious hours debugging the latest and greatest thing that isn't quite reliable yet. Their choice is wiser and might get attention of more risk-conscious users.
- mixmastamyk 11y agoCongrats, not everyone can create a simple elegant platform and write about it in such an accessible manner. I suppose you're standing on the shoulders of giants, but still. This is the level of engineering/communication I always shoot for, and which (somewhat disappointingly) is rare where I've worked.
- ejholmes 11y agoThanks for the kind words. As a shameless plug, we are hiring: https://www.remind.com/careers https://www.remind.com/careers :)
- bgentry 11y agoReally cool stuff. Seems like you found a good way to hand off most of the hard stuff to AWS and only do a few key things yourselves to make the experience better. As such I think Empire has the potential to be a viable option for many companies, which is something I rarely say about a PaaS project :)
- ejholmes 11y agoThanks Blake! I think somebody mentioned that we were standing on the shoulders of giants. I think most of your contributions around this domain qualify for that :)
- scanr 11y agoInstead of nginx, we've had a pretty good experience using vulcand (https://github.com/mailgun/vulcand https://github.com/mailgun/vulcand) as the front-end router for our micro-services.
- LunaSea 11y agoAny reason(s) for switching to Vulcand rather than NginX ?
- scanr 11y agoFirst off, nginx is awesome and you can do all the things we did with vulcand with nginx, so it was just a question of friction. The reason we went with vulcand is that it natively supports what we wanted to do i.e. route to micro-services based on dynamic etcd driven configuration. To do the same thing in nginx (at the time), we would have either had to use confd or custom lua.
- loki77 11y agoWe actually looked at VulcanD in an older version of Empire. When we decided to use a routing layer with this version of Empire, rather than just letting Empire/ELB expose each service (mostly because it is a lot easier for us to later shut off public access to each service) we threw together nginx because it was so simple. I think at this point everytime we move a service we add like 5 lines to an nginx config, re-deploy the router in Empire, and the service is exposed. The internal 'service discovery' makes this a lot easier, since we just have to tell nginx to route to http://<app_name> http://<app_name> - no domain, no port, nothing more than the app_name thanks to DNS/resolv.conf search path & ELB stuff.
- UserRights 11y agoHow to autoscale with this?
- loki77 11y agoSo at this point Empire itself doesn't deal with things like Autoscaling. That said, the Demo Cloudformation template (and the bootstrap script that kicks it off easily for you) make use of Autoscaling groups for the instances that containers are being run on. So, in theory you could autoscale just like you always would. Monitor stats for a host, if a bunch of them start to run low on resources, kick off an autoscaling event. That said, there's been quite a bit of talk about integrating Empire with Autoscaling, so that when, say, ECS couldn't find any instances with resources free for a task, Empire could kick off the autoscaling events for you. Could be pretty awesome :)
- smanuel 11y ago> We tried Deis briefly but ultimately decided that it was more complicated than we felt it needed to be. That kind of reminds me of https://xkcd.com/927/ https://xkcd.com/927/ Sorry if that's not the case. I've also played briefly with Flynn and Deis and I haven't found anything that complicated that would need a whole rewrite and changing the entire approach. Moreover with Deis I can easily change providers (DO, AWS, Azure, etc.) and with Emprire I'm bound to ECS. At least that was my first impression, I have to read more.
- athrun 11y agoIMHO, it's best to see Empire/Deis/Dokku/etc. as a mean to an end, not the end itself. While Empire itself may be tied to AWS, your app is still a portable, 12-factor, Heroku-compatible app. You can run it elsewhere.
- ejholmes 11y agoOur approach was to re-use as much existing technology as possible, which is not the case for most others. That's the "complication" I'm referring to here. Empire grew from the need for a production grade platform that was going to be stable. Empire doesn't actually lock you into ECS. The scheduling backend is pluggable and could support Kubernetes/Swarm in the future.
- whalesalad 11y agoI'm interested in hearing more about how you use this in terms of development lifecycle. Does a container image get created for every release of your app? I've always wondered about the more correct approach to this. This is how I currently use Docker: 1) Custom base image with all the things my company needs like supervisord, libpq, etc.. 2) Custom per-service base images like ones with Java for our Clojure services or Python for our research services which are built off of the base. 3) A release consists of pulling the latest version of the base image, example, acme-python, and then injecting the latest project code into it. My concern here essentially boils down to the image repo. Github needs to add container storage because while I admire Docker Hub's efforts, I don't trust it.
- ejholmes 11y agoWe have a setup that has been working out well for us: 1. We build docker images on every commit, in CI, and tag it with the git commit sha and branch (we don't actually use the branch tag anywhere, but we still tag it). This is essentially our "build" phase in the 12factor build/release/run. Every git commit has an associated docker image. 2. Our tooling for deploying is heavily based around the GitHub Deployments API. We have a project called Tugboat (https://github.com/remind101/tugboat https://github.com/remind101/tugboat) that receives deployment requests and fulfills them using the "/deploys" API of Empire. Tugboat simply deploys a docker image matching the GitHub repo, tagged with the git commit sha that is being requested for deployment (e.g. "remind101/acme-inc:<git sha>"). We originally started maintaining our own base images based on alpine, but it ended up not being worth the effort. Now, we just use the official base images for each language we use (Mostly Go, Ruby and Node here). We only run a single process inside each container. We treat our docker images much like portable Go binaries.
- juliangregorian 11y agoWhy not just add a Dockerfile to your git repo and add a `docker build` step to your deploy?