13 ms·
Firecracker – Lightweight Virtualization for Serverless Computing
- colemickens 8y agoThe crosvm and Rust have me intrigued. I was hoping for something like this since I saw the first hints of Rust showing up in ChromeOS in crosvm. A compare/contrast with Kata Containers would also be interesting. Their architectures look similar. (Kata Containers [1] being another solution for running containers in KVM-isolated VMs, that has working integrations with Kubernetes and containerd already. Not affiliated, but I'm tinkering with it in a current project, though I'm also now keen to get `firecracker` working as well.) Obviously, if nothing else, qemu vs crosvm is a big difference, and probably significant since my understanding is that Google chose to also eschew using qemu for Google Cloud. [1]: https://katacontainers.io/ https://katacontainers.io/
- aliguori 8y agoKata Containers is a lot of infrastructure for running containers and it uses QEMU to run the actual VMs. Firecracker just replaces the QEMU part and we're eager to work with folks like the Kata community. I love QEMU, it's an amazing project, but it does a ton and it's very oriented towards running disk images and full operating systems. We wanted to explore something really focused on serverless. So far, I'm really happy with the results and I hope others find it interesting too.
- zaxcellent 8y agoWe felt the same way about QEMU before we started crosvm. Glad to see you all found some use out of it.
- wirelessben 8y agoThe devops training site katacoda.com will be interesting to watch. They spin up and tear down _so_ many VMs, their cloud bill must be monstrous. Firecracker is much leaner, so they would save a lot of cycles by spinning up Firecracker over Kata.
- colemickens 8y agoKatacoda has nothing to do with Kata Containers... I'm not sure how you can make any of the conclusions anyway, unless you know a lot of seemingly private details about how KataCoda is implemented.
- mcrute 8y agoMore discussion here: https://news.ycombinator.com/item?id=18539532 https://news.ycombinator.com/item?id=18539532
- sudhirj 8y agoWhat this is allows, and I'm hoping a full fledged service will be announced on Thursday or Friday, is running containers as Lambdas. i.e. if you application starts fast enough, you can just set a container to start and run as a request comes in. It can also shut down when it's done running. This allows things like per second billing for container runs, serverless containers (there's no container running 24/7, only when there's traffic), etc.
- discodave 8y ago> containers as Lambdas. How similar is AWS Fargate to what you're describing?
- xmly 8y agoPricing model is not serverless. The basic serverless principle is no-use-no-pay.
- talawahtech 8y agoTheoretically, if container start up times are around 125ms it should be possible to achieve this with Fargate + Knative's "scale to zero" functionality[1]. AWS is already working on improving Fargate/Kubernetes integration. [1] https://cloud.google.com/knative/ https://cloud.google.com/knative/
- coder543 8y agoDoesn't a Kubernetes master on Fargate cost a significantly non-zero amount of money? Pretty sure it does.
- talawahdotnet 8y agoSure, but if you already have one then there is no incremental cost. It would be great if they announced that they were gonna remove EKS master costs altogether. Technically Firecracker should make it possible for them to run that infrastructure more efficiently :)
- Tehnix 8y agoI really hope this helps with the cold start times on Lambda. We were currently looking heavily into moving our API from Lambda to EKS, but if this impacts cold start times, I think we will look at how it ends up looking like in practice.
- sudhirj 8y agoMost code start time problems on Lambda I've seen are VPC related - public network Lambdas start in milliseconds, with the main lag being the userspace code startup time. Starting a Lambda inside a VPC involves attaching a high security network adapter individually to each running process, which is likely what takes so long. I assume AWS is working on that, though, they've claimed some speedups unofficially. If your security model allows, try running your Lambdas off-VPC.
- Tehnix 8y agoThe VPC startup times are insane, so we quickly move our lambdas out of that, accepting the trade off. Our normal cold starts are in the 1-2 second range, and the app initialization comes after. Too high for an API facing users :/
- sudhirj 8y agoThat's odd. Are you running the interpreter runtimes (NodeJS/PHP), compiled binaries (Go) or VMs like Java?
- Tehnix 8y agoWe are using Node.js. From what I’ve seen online, Python should be the fastest though. Edit: wanted to add, that from what I’ve gathered from people testing online, bundle size didn’t really matter, but perhaps someone else has some information that points to the contrary?
- sudhirj 8y ago
- krat0sprakhar 8y ago> Firecracker was built in a minimalist fashion. We started with crosvm and set up a minimal device model in order to reduce overhead and to enable secure multi-tenancy. Firecracker is written in Rust, a modern programming language that guarantees thread safety and prevents many types of buffer overrun errors that can lead to security vulnerabilities. This is awesome! Really excited to try this out!
- jhwang5 8y agoThis is great. The line between use cases for VM vs container are getting super blurry.
- blasdel 8y agoThere's a Github Pages FAQ describing why it was made and how it fits with other solutions: https://firecracker-microvm.github.io/ https://firecracker-microvm.github.io/ and a high-level design document about how it works https://github.com/firecracker-microvm/firecracker/blob/master/docs/design.md https://github.com/firecracker-microvm/firecracker/blob/mast...
- espeed 8y agoInteresting name choice. When I clicked on the link and saw the name and design, my first thought was, "Is this a Firebase knockoff...?" [1] ... and then I scrolled to the bottom to see the copyright and saw this project is by Amazon Web Services. [1] https://firebase.google.com https://firebase.google.com
- whalesalad 8y agoI’m very excited to play with this technology in the same way I love playing with Elixir/Erlang and userland concurrency models. I also love the idea of docker (and use it daily) but dislike the ergonomics. My first thought is, particularly with the emphasis on oversubscription, how does the kernel of the host schedule work?
- talawahtech 8y agoThis is huge! It basically removes the VM as the security boundary for something like Fargate [1]. This should lead to a significant reduction in pricing since Fargate will no longer need to over provision in the background because VMs were being used even for tiny Fargate launch types. It should hopefully eliminate the cost disparity between using Fargate vs running your own instances. Should also mean much faster scale out since you containers don't need to wait on an entire VM to boot! Will be interesting to see what kind of collaboration they get on the project. This is a big test of AWS stewardship of an open source project. It seems to be competing directly with Kata Containers [2] so it will be interesting to see which solution is deemed technically superior. [1] https://aws.amazon.com/fargate/ https://aws.amazon.com/fargate/ [2] https://katacontainers.io/ https://katacontainers.io/
- tlrobinson 8y agoIt sounds like it’s already being used in Lambda and Fargate, though I’m not sure how long that’s been the case: > Firecracker has been battled-tested and is already powering multiple high-volume AWS services including AWS Lambda and AWS Fargate
- Aissen 8y agoIndeed, this seems very similar to kata+runv+kvmtool(lkvm). I'm curious why they don't provide a comparison. Here's what I gathered: - it seems to boot faster (how ?) - it does not provide a pluggable container runtime (yet) - a single tool/binary does both the VMM and the API server, in a single language. Can anyone else chime in ?
- justincormack 8y agoFrom memory the original version of Intel Clear Containers had its own kvm based vmm but they moved back to qemu (or a more minimal patched version they maintain). They are working on containerd support so should be similar to Kata soon.
- Aissen 8y ago
- polskibus 8y agoHow does this compare to containers?
- perbu 8y agoContainers share the OS kernel and some services. This is a virtual machine monitor, so it deals with virtual machines. A container can only run Linux containers. Firecracker can likely run other operating systems, such as IncludeOS. You can't run those in containers.
- tlrobinson 8y agoThis looks great, I’m just wondering what Amazon’s motivation for open sourcing it is. It seems like some pretty critical secret sauce for making services like Lambda and Fargate both secure and efficient.
- dm3 8y agoPushing the adoption of "serverless" - benefits Amazon ultimately as it's the largest provider.
- fapjacks 8y agoRight. In the end, AWS saw containerization as an existential threat, and serverless is its response to the commoditization of AWS (and other specialized cloud vendors) by containerization technology. Serverless helps AWS re-couple your application back into the specialized AWS vendor environment, once more requiring you to keep specialized and costly AWS-specific knowledge on-hand in order to build and deploy your application. It gives them plenty of room to advertise and provide all of those edge case services to you once more (and their usage charges) and also helps them prevent you from treating AWS like a rack of networked CPUs to serve as the substrate for your application containers. People (mostly AWS folks -- dig a little deeper into who is writing much of the serverless blog posts out there) keep pushing the "serverless is containers" but that's just a tactical response. Add a layer of abstraction and it's very clear why AWS is betting so hard on serverless. Originally, AWS commoditized the old datacenters by providing the same network/CPU substrate, but at a higher cost because you outsourced the management of those resources to AWS. And AWS slowly dripped out new and convenient services for your application to consume, allowing you to outsource even more of your application needs to this one vendor. And while the services offered by AWS were just a little bit different, they were functionally similar. And that's how you locked yourself into using AWS instead of CoreColoNETBiz or whatever datacenter you were using before. I remember one of the first major outages of us-east-1, which caught most of the internet with its pants down (interestingly, the answer was just to give more money to AWS for multi-region redundancy). AWS had a pretty good thing going: Outsourcing the management of all those resources to AWS is expensive! But that's when containers came along and people at AWS started to take notice. With containers, people could de-couple their applications from Dynamo and Elastic Beanstalk and VPC and all those specialized services that cost so much time/money. Instead, you could just cram all that shit into containers, without needing to set up IAM roles or pore over Dynamo documentation or dump so much time into getting VPC set up just right. And that's the whole point of containerization: Easily build your services in a homogeneous environment with exactly the software you want to use and eliminate that technical debt of vendor lock-in and the enormous cost center of specialized vendor knowledge (e.g. Dynamo, IAM, VPC, etc etc). Treat the cloud -- any cloud -- like a bunch of agnostic resources. Docker commoditized the commoditizers. And serverless is how AWS plans to get you to re-couple your application tightly to their specialized web of knowledge and services. They get to say that you're still using containers, but they need to gloss over the fact that you're locked into the AWS version of containers. You cannot "export" your specialized AWS-only knowledge of Fargate or Lambda or API Gateway or ECS to Google Cloud or Azure or some dirt cheap OVH bare metal. You're tightly re-coupled to AWS, having bought into their "de-commoditization" strategy. Which I need to stress is totally fine, if you're okay with that. It just needs to be made clear what you are trading off.
- solatic 8y ago"Process Jail – The Firecracker process is jailed using cgroups and seccomp BPF, and has access to a small, tightly controlled list of system calls." So basically, a gVisor alternative?
- ec109685 8y agogVisor doesn't use KVM: "Machine-level virtualization, such as KVM and Xen, exposes virtualized hardware to a guest kernel via a Virtual Machine Monitor (VMM). This virtualized hardware is generally enlightened (paravirtualized) and additional mechanisms can be used to improve the visibility between the guest and host (e.g. balloon drivers, paravirtualized spinlocks). Running containers in distinct virtual machines can provide great isolation, compatibility and performance (though nested virtualization may bring challenges in this area), but for containers it often requires additional proxies and agents, and may require a larger resource footprint and slower start-up times."
- solatic 8y agoYeah but one of the main ways in which gVisor provides security is by intercepting system calls and strictly limiting which calls can be made. Firecracker may use KVM instead of running entirely in usermode, but as far as most of us are concerned, that's an implementation detail. The pertinent question is whether the price of security is limiting the possible system calls, which means that Firecracker won't be able to run arbitrary containers, just as gVisor doesn't guarantee that it can run arbitrary code (which may require filtered system calls).
- ec109685 8y agoThat’s not true. Your guest application has access to all Linux system calls in the guest VM. You can see here the security model: https://github.com/firecracker-microvm/firecracker/blob/master/docs/design.md https://github.com/firecracker-microvm/firecracker/blob/mast... The firecracker process itself is limited in the system calls it can make, but kvm allows the guest Linux process the ability to expose a full set of system calls to end user applications.
- tatoalo 8y agoIt’s QEMU without all the legacy stuff, they also open sourced it, interesting.
- sudhirj 8y ago@zackbloom, @kentonv hint hint. Isn't this roughly the same memory footprint as a Worker? CONTAINERS ON ALL THE CLOUDFLARE THINGS!
- sudhirj 8y agoIf you can implement it by tomorrow afternoon before the Andy Jassy keynote you might be able to steal some thunder.
- ec109685 8y agoYou still have a full Linux kernel running inside the vm though?l with Firecracker versus essentially a fiber with cloudflare.
- zackbloom 8y agoHeh. Truthfully, what I'm most excited about right now is being able to start a worker in less time than it takes to make an internet request. When you can do that you get magical autoscaling and it becomes just as cheap to run it in hundreds of places as one. As long has you have to invest ~100ms of CPU to get one of these VMs running I'm not sure it will have quite the same economics.
- sudhirj 8y agoYeah, jokes aside I simply don’t think it makes sense to run full processes on the edge. Not yet, anyway. Script isolates makes a lot of sense with current hardware limitations, but full processes at the edge are coming sooner or later.
- zackbloom 8y agoThat would make me a little sad. I'm not excited about the idea that we figured out the ideal way for a program to be encapsulated in 1965 and it will never change.
- polskibus 8y agoDoes it support Windows?
- chupasaurus 8y agoKVM-based, so no it doesn't.
- perbu 8y agoKVM supports Windows just fine, which is why you can run Windows on GCP and Openstack. And Firecracker seems to support enough of a machine to boot Windows as long as the windows instance has support for libvirt disk devices and a libvirt NIC. However, it seems they boot in a slightly unconventional way. They take a elf64 binary and execute it. This works for Linux and likely some other operating systems that can produce elf64 binaries. Windows supports legacy x86 boot and UEFI, but likely not elf64 "direct boot". So if you can get windows into an elf64 binary and have it run without a GPU you could have it boot. So, likely not. But the reason isn't due to KVM.
- steveklabnik 8y agohttps://firecracker-microvm.github.io/ https://firecracker-microvm.github.io/ says > What operating systems are supported by Firecracker? > > Firecracker supports Linux host and guest operating systems with kernel versions 4.14 and above. The long-term support plan is still under discussion. A leading option is to support Firecracker for the last two Linux stable branch releases.
- xaduha 8y ago> microVMs, which provide enhanced security and workload isolation over traditional VMs, while enabling the speed and resource efficiency of containers. Reminds me of rkt + kvm stage 1 https://github.com/rkt/rkt/blob/master/Documentation/running-kvm-stage1.md https://github.com/rkt/rkt/blob/master/Documentation/running... Too bad it didn't take off.
- sdart 8y agoDoes this provide any multi host cluster management capabilities?
- kraemate 8y agoClear containers (now called kata containers) did this more than three years ago, with similar performance numbers (sub 200 ms boot times). It is frustrating, but not surprising, to see the same regurgitated solution receive this much excitement. The firecracker documentation also does not mention the similarity with prior work, oh well. [Not affiliated with Intel in any way---just a long-time proponent of the clear containers approach.]
- cleansy 8y agoAfter Amazon released its implementation the whole eco system profits, as it creates diversity and buzz around that topic. I think it's great to have (open source) alternatives, especially with the marketing weight of amazons solution entering "playing field". Also: is it clear that kata was first? Three years doesn't sound like they've been miles ahead. [Not affiliated with either side]
- talawahtech 8y agoThe FAQs on the Firecracker website[1] specifically address the difference between Firecracker and Kata Containers. The main thrust being that they have decided not to use QEMU and have instead chosen a much more minimal "cloud-native" oriented approach that deliberately abandons certain features in order to gain greater security, efficiency and agility going forward. They also decided to implement it in Rust. Based on the the responses I have seen from non-Amazon employees with experience in this space[2][3][4], it looks like their approach is solid. It should also be noted that one of the main architects of Firecracker was formerly the project lead for QEMU[5][6] 1.https://firecracker-microvm.github.io/#faq https://firecracker-microvm.github.io/#faq 2.https://twitter.com/bcantrill/status/1067326416121868288 https://twitter.com/bcantrill/status/1067326416121868288 3.https://twitter.com/jessfraz/status/1067286831287418881 https://twitter.com/jessfraz/status/1067286831287418881 4.https://twitter.com/kelseyhightower/status/1067294780948832258 https://twitter.com/kelseyhightower/status/10672947809488322... 5.https://twitter.com/jessfraz/status/1067282499938721792 https://twitter.com/jessfraz/status/1067282499938721792 6.https://twitter.com/anliguori/status/1067293131366785024 https://twitter.com/anliguori/status/1067293131366785024
- 8y ago
- mark212 8y agostill seems much slower than the model used by Cloudflare for what they call "workers."[1] A recent blog post a few weeks back was the subject of considerable discussion here[2], and it seems to me to be doing much the same thing as Firecracker, but still faster because there's less overhead. But maybe I'm missing something. [1] https://blog.cloudflare.com/cloud-computing-without-containers/ https://blog.cloudflare.com/cloud-computing-without-containe... [2] https://news.ycombinator.com/item?id=18415708 https://news.ycombinator.com/item?id=18415708
- tlrobinson 8y ago> But maybe I'm missing something. From the "Disadvantages" section of your first link: "No technology is magical, every transition comes with disadvantages. An Isolate-based system can’t run arbitrary compiled code. Process-level isolation allows your Lambda to spin up any binary it might need. In an Isolate universe you have to either write your code in Javascript (we use a lot of TypeScript), or a language which targets WebAssembly like Go or Rust." "If you can’t recompile your processes, you can’t run them in an Isolate. This might mean Isolate-based Serverless is only for newer, more modern, applications in the immediate future. It also might mean legacy applications get only their most latency-sensitive components moved into an Isolate initially. The community may also find new and better ways to transpile existing applications into WebAssembly, rendering the issue moot."
- tuananh 8y agothe way i see it, firecracker is more flexible but cloudflare workers isolate is faster. amazon can't afford the limitation of Isolate hence this project.
- testbotlo2 8y agoCan someone explain me how does this work? Is it an orchestration service for containers like Kubernetes or is it any different?
- xrd 8y agoMy big question is: is this something only exciting for people doing lambda at massive scale? Qemu is exciting technology and has paved the way for all kinds of interesting layers. So, creating a slimmed down improvement that really makes it faster and provides a new lambda-ish execution context is great. I'm sure Amazon cares about that. I'm sure people doing millions of lambda calls a day care about that. But, if I'm an entrepreneur thinking about building something entirely new, is there something I'm missing about this that would make me want to consider it? Lambda and Firebase Functions are exciting partially because they break services into easy to deploy chunks. And, perhaps more importantly, easy things to reason about. But that's not the big deal: the integration with storage, events, and everything else in AWS (or Firebase) is what really makes it shine. It's all about the integration. When I read this documentation, I'm left wondering whether I want to write something that uses the REST API to manage thousands of micro vms. That seems like extra work that Amazon should do, not me. Am I missing something important here? Surely Amazon will integrate this solution somewhat soon and connect it to all the fun pieces of AWS, but the fact that they didn't consider or mention it makes me think it is something I should not consider now.
- nunez 8y agoI am extremely excited by this. i wonder if this can be used to provision jit kubernetes workers.