8 ms·
How we deploy faster with warm Docker containers
- Aeolun 4y agoWhile a great improvement, 45 seconds to change my serverless code still feels like a lot? At least I’m comparing it to Cloudflare Workers, which deploys fast enough (<1 sec) that I can actually use it as a development environment.
- moralestapia 4y agoCF Workers run straight on top of V8, IIRC? Which I think it was a great design decision. But I wonder how they achieve that w/ the other backends like rust.
- winstonewert 4y agoThey run Rust compiled to WebAssembly.
- moralestapia 4y agoSure, but what's the stack? Wasmer or something like that? I haven't bothered to look it up, for sure they've written about it.
- syrusakbary 4y agoCloudflare Workers uses v8 under the hood afaik (which makes sense, since they are js-first). Although they recently announced that Wasmer was being used in their 1.1.1.1 DNS (however, that has little to do with the Workers product I believe)
- oofbey 4y agoIf I had to wait 45 seconds to see a single line of code change run, I would be looking for a new dev environment. Reminds me of compile times in the 1990’s. Serverless is “cool” and all but productivity is what really matters.
- shalabhc 4y ago(author here) I agree it would be fantastic to have sub second deploys! Doing this for user specified Python environments is challenging in different ways than doing it for a JS SDK like Workers. Note that just provisioning a GitHub runner takes about 10s, before any deploy code even starts up. In theory we could rsync the code directly to the code server and reload it. But any of these options require bigger architectural changes. Also you can already have a local environment setup for faster iteration and use this for the shared environments with other developers, where the speedup is still great to have.
- MuffinFlavored 4y ago> Doing this for user specified Python environments Who is using this/asking for this/why?
- shalabhc 4y agodagster.io is an open source Python library for building data pipelines. When using Dagster your project may depend on other Python data analysis and warehouse libraries like pandas, snowflake, pyspark and so on. Each org's requirements for their data orchestration will be different. When executing the pipelines in Dagster Cloud, we need to replicate their Python environment.
- faizshah 4y agoIs this approach not just Capistrano but using containers and specific to python? You could just use firecracker microVMs and ansistrano to implement this same workflow but get better isolation from firecracker.
- shalabhc 4y ago(author here) Firecracker is definitely very interesting. Would require more ops work for us to run bare metal EC2 (we currently use Fargate). IIUC reusing pre-existing environments would require us to share ext4 filesystems across the VMs. Not sure if antistrano helps here but will look into it.
- est 4y agoI remember once there was a method to use BitTorrent to deploy images/packages in intranet.
- hotpotamus 4y agoI remembered one called "Murder" which looks like it was a Twitter thing back in the day[0]. Looks like Facebook did it too. I guess there was a phase that tech went through in 2010ish. [0]https://blog.twitter.com/engineering/en_us/a/2010/murder-fast-datacenter-code-deploys-using-bittorrent https://blog.twitter.com/engineering/en_us/a/2010/murder-fas...
- notpeter 4y agoUber did it too. Kraken is docker registry that gossips containers via bittorrent. https://github.com/uber/kraken https://github.com/uber/kraken
- wolfgang42 4y agoIt does make sense to avoid having servers that saturate their link with a zillion identical downloads on every deploy; but I wonder why they went for BitTorrent, rather than using multicast the way e.g. Windows enterprise installs do it.
- adql 4y agoCoz you need working multicast across entire network or server-per-datacenter to multicast (and data still needs to get to that server). It's much simpler to just point a server to torrent tracker and let them exchange stuff between themselves
- rezonant 4y agoWhoa, this seems to take for granted that devs targetting serverless infrastructure _must_ deploy up into the actual serverless infrastructure during development, unless I'm misreading? Why is there not a way to simulate the final infrastructure locally, so that development is possible without extremely inefficient pushes into a "dev" serverless deployment?
- naikrovek 4y agothere are solutions for this. localstack comes to mind, but it is more expensive per user than GitHub enterprise and copilot combined. the localstack devs are very proud of localstack.
- dominicst 4y agoLocalstack is brilliantly done but no great options for other cloud providers for non-AWS companies. It also can be a bit of a pain doing the initial set up, once done it definitely improves the development process/experience. Majority of the time for me a docker compose with my dependencies does the job well enough and is much faster to get up and running with
- miohtama 4y agoIt's called localless (:
- danmur 4y agolol
- rozap 4y ago[the year is 2028] Dear Valued Customers, As CEO, I wanted to take a moment to update you on a change to our pricing model. Moving forward, there will be a cost increase for each new function that you write and deploy, as well as an additional fee for each variable declared in that function. We believe this change will enable us to continue to provide you with the best services while also ensuring our business remains sustainable. Our team will provide a detailed quote for any new functions you request, outlining the costs and estimated timeframes. We are committed to providing our services to everyone, regardless of their background or circumstances. This change will not affect the quality of the services that we provide, and we remain dedicated to providing you with the best possible experience. Thank you for your continued support, and please do not hesitate to reach out to us with any questions or concerns.
- deleted 4y ago[deleted]
- dsunds 4y agoAWS has a few projects to reduce launch times for example https://aws.amazon.com/blogs/containers/reducing-aws-fargate-startup-times-with-zstd-compressed-container-images/ https://aws.amazon.com/blogs/containers/reducing-aws-fargate... https://aws.amazon.com/about-aws/whats-new/2022/09/introducing-seekable-oci-lazy-loading-container-images/ https://aws.amazon.com/about-aws/whats-new/2022/09/introduci...
- nathants 4y agolambdas update-code api takes less than a second. using a container instead of a zip for lambda has advantages, but speed is not one of them. i auto rebuild my go zip and patch aws on every file save. it’s done before i alt tab, up arrow, and curl. script: https://github.com/nathants/aws-gocljs/blob/master/bin/dev.sh https://github.com/nathants/aws-gocljs/blob/master/bin/dev.s...
- satyanash 4y ago> Consider git – it only ships the diffs, yet it produces whole and consistent repositories. IIRC git does _not_ ship diffs. It copies whole files even for the tiniest change. The compression layer handles the de-duping, which is a different layer.
- alfons_foobar 4y agoYup, you are correct - git stores snapshots of files, not diffs.
- yencabulator 4y agogit typically stores deltas of snapshots of files.
- slavik81 4y agoI would interpret 'ships' as 'pushes', because git does send delta packfiles on the wire.
- 8organicbits 4y agoI'm curious about the remaining 20s to start the code on the container. What's the bottleneck there? It seems like you'd need to identify the container, tell it which S3 object has the new source.pex, the container would download it, and then when you run `PEX_PATH=deps.pex ./source.pex` you're up. All that feels like it should take less than 20s. If picking the container takes long, you could probably start that process as sources.pex is built.
- shalabhc 4y agoGetting the message to the right container is one bottleneck. Currently this is routed through a couple of hops and includes some polling. This could all be optimized (if we had a direct line to the container) but the same messaging model is used in other contexts and would need architectural changes. Another bottleneck is running `source.pex` itself takes a few seconds to start up because it analyzes user code (and in some cases may do expensive computation.) But you're right: if `source.pex` as a hello world program, just downloading and running it should be pretty fast - I'd expect around 1s.
- nitwit005 4y agoIt seems strange to use a highly managed deployment environment like Fargate, but then build another deployment tool on top of it to do things in a simpler way. It feels like EC2 is being reconstructed on a platform meant to hide it.
- nijave 4y agoOr (dare I say) look into EKS. Kubernetes can spin up containers faster than ECS in my experience (as of ~1 year ago). Seems like the ECS control plane just has more latency (even with EC2 instead of Fargate)
- shalabhc 4y agoIs EKS safe for multi-tenant use? When we looked it appeared unsafe if we want to run our users code next to each other because of possible isolation issues.
- nijave 4y agoI guess that depends on your use case and risk profile. Linux containers are a pretty well established isolation mechanism and you can potentially add some additional safety with per-tenant dedicated nodepools. If pods have added privileges or there is a really low risk tolerance, maybe that's not enough isolation. Sounds like you can change the container runtime with EKS (not sure if that impacts AWS support) so you could use gVisor or runvm https://www.verygoodsecurity.com/blog/posts/secure-compute-part-2 https://www.verygoodsecurity.com/blog/posts/secure-compute-p...
- lewo 4y ago> The key factor behind our decision was the realization that while Docker images are industry standard, moving around 100s of megabytes of images seems unnecessarily heavy-handed when we just need to synchronize a small change. I think the culprit is more the GitHub Actions cache than Docker since it seems to be hard to get a clean cache management. I'm not sure about caching Docker image layers, but caching the Nix store with GitHub Actions is pretty complicated (not even sure it's possible): this means we have to download all required Nix store paths on each run, but i consider this is because of a GitHub Action cache limitation. So, did you consider using another CI, which offers better caching mechanisms? With a CI able to preserve the Nix store (Hydra[1] or Hercules[2] for instance), I think nix2container (author here) could also fit almost all of your requirements ("composability", reproducibility, isolation) and maybe provide better performances because it is able to split your application into several layers [2][3]. Note i'm pretty sure a lot of Docker CI also allows to efficiently build Docker images. [1] https://hercules-ci.com/ https://hercules-ci.com/ [2] https://grahamc.com/blog/nix-and-layered-docker-images https://grahamc.com/blog/nix-and-layered-docker-images [3] https://github.com/nlewo/nix2container/blob/85670cab354f7df69dd4af097c27cf9bc5826cb2/examples/uwsgi/default.nix https://github.com/nlewo/nix2container/blob/85670cab354f7df6...
- FBISurveillance 4y agoThere's been a recent Launch HN of Depot.dev [1] - I've integrated it quickly into my GitHub Actions workflow and it's blazingly fast (13x speedups for me). It also was a drop-in replacement since I was using Docker Bake and Docker Action and Depot mimics that almost fully (except SBOM and provenance bits). It also works with Google Cloud Workload Identity Federation so image pushes to Artifact Registry didn't need any tweaking. [1] https://news.ycombinator.com/item?id=34898253 https://news.ycombinator.com/item?id=34898253 Disclaimer: not affiliated, a happy paying customer.
- shalabhc 4y agoThanks for the interesting links - I'll check them out! We would need not just another CI but also another container platform because launching a docker container is also slow. Irrespective of the CI, I believe all cached Docker layers will need to be downloaded onto the build machine before it can be rebuilt. Still, I believe it is possible to build and deploy faster even with a "docker image only" design and it's something we are still looking at. The question is what is the lower bound here - would be hard to beat "sync a file to a warm container and run it". Pex gives us a pretty good lower bound that is also container platform agnostic.
- polyrand 4y agoQuite cool to see PEX. I've used a similar package, shiv[0], with great results, and I always wondered why these are not used more. I think Python zipapps are really nice for bundling executables. [0]: https://shiv.readthedocs.io/en/latest/index.html https://shiv.readthedocs.io/en/latest/index.html
- IanCal 4y agoI'm confused, this feels like a complicated way of creating a docker layer. Building and hashing the dependencies is exactly what adding the requirements file/etc and building a layer does. > moving around 100s of megabytes of images seems unnecessarily heavy-handed when we just need to synchronize a small change. Consider git – it only ships the diffs, yet it produces whole and consistent repositories. Just ship the layers, that's literally what docker does, right? Is this a whole build process just to get around the ephemeral nature of github actions?
- mschuster91 4y agoThe problem is you can't (easily...) create a layer atop of a Docker image without pulling all of it before. Basically, a conventional Docker CI build process looks like the following in what I use (where <tag> is something like branch-foobar): 1. docker login <repo> 2. docker pull <repo>/<image>:<tag> || true 3. docker pull <repo>/<image>:branch-master || true 3. docker build --pull --cache-from <repo>/<image>:<tag> --cache-from <repo>/<image>:<master> -f Dockerfile.<branch> -t <repo>/<image>:<tag> . 4. docker push <repo>/<image>:<tag> Using --cache-from cuts down dramatically on the build times since (assuming your branch-master tag gets rebuilt nightly) at least the Dockerfile layer steps for downloading the base OS image and installing software can be automatically skipped. But still, even swapping out the last step in the Docker build makes it necessary to pull the entire image first - assuming your average Java application, that's like 500+ MB for OS+Java JRE+OS dependencies. If you're sure that the change is only in the last COPY/ADD step, you still have to pull the image despite it not being needed technically.
- IanCal 4y agoSo that's a problem with the builder being ephemeral then. I am fairly sure a docker layer is just a tarfile. If you're just adding, say, python code and performing no execution I feel like you could create the tarfile and add it to the list of layers. edit - This seemed dismissive of me. It isn't intended like that, I'm genuinely curious as to whether this is the right approach. The post feels intuitively like it's reinventing some things that exist, but perhaps for good reason. I think it'd be nice if everything else was "just docker". Perhaps I'm missing a key point - it feels like just having a build machine with storage would immediately solve the problem, there would be no remote pulling. Is that not a solution?
- jayrwren 4y agoI don't see how this has anything to do with faster Docker containers. Looks more like faster python distribution.