13 ms·
The AWS EC2 Windows Secret Sauce
- altcognito 8y agoIt would make sense to apply these optimizations to any kind of startup since presumably anticipating resource allocation ends up as a cost savings and improvement in user experience regardless of the base OS. There are plenty of Linux AMI images that are bloated and slow to start.
- simonh 8y agoIt would only be practical to do this for very commonly used AMIs that AWS create and provision themselves. Commonly used because you don't want unused AMIs languishing in the instance pool long term. Only their own, because these AMIs are heavily manipulated and then customised for the user so you have to be positive the AMI is compatible with the manipulations. I wonder if Azure does something similar.
- brian_herman__ 8y agoVery interesting. Now that 4 minute wait time makes sense now.
- deleted 8y ago[deleted]
- scarface74 8y agoI’ve been a Microsoft developer for over two decades. But once I started architecting solutions on AWS and seeing the Windows Tax first hand - in terms of resource requirements and licensing costs - I started trying to avoid Windows like the plague. Also the true costs of infrastructure became my problem - accounting can see exactly what a solution costs - instead of some amorphous cost in the IT budget.
- briffle 8y agoI really like the better visibility into the true costs. Too many times in the past, a team had a budget for the software they needed to deploy a project, but the Virtualization software, storage array, etc, were all out of the IT budget. that makes the IT budget look bloated, and hard to explain, and ripe for being seen as a 'cost center'.
- mc32 8y agoThat’s too bad and unfortunate of some companies who see IT as a black hole. Imagine having an R&D dept which could incur lots of costs and little result individually but as a group produce good licensable IP. Sometimes you have to realize some things are a cost of doing business.
- scarface74 8y agoIt's not that they don't realize it. The R&D department doesn't have to be accountable for costs. The R&D department isn't incentivized to control costs because it isn't our problem. Proper tagging of resources shines a light on the R&D department and the business knows exactly what our projects' infrastructure cost is.
- 0x0 8y agoDoes AWS have to pay a license fee to maintain a pool of Windows instances? Who foots the cost of the license if a prepared/pooled instance is never allocated to a customer?
- jbigelow76 8y agoThere has to be a custom deal in place, but it could be similar in spirit to a hot standby MS SQL Server instance where you only pay for one license but have a mirror of the machine ready to go if the first one goes down (my knowledge of licensing was only tangential to my dev work and could be several years out of date now).
- bharrison 8y agoThere are MS license agreements for cloud solutions providers that allow license usage up to a certain limit, with monthly reporting of actual usage for billing purposes.
- jak92 8y agoShared windows instances? What could go wrong?
- camtarn 8y agoMy assumption is that they're not shared. Once you request an instance from the pool, if the desired number of pooled instances is still the same, another one would immediately begin the imaging process to take its place.
- rebelde 8y agoAWS might as well apply any Windows Update changes so that they deliver a fully-patched and secure instance, instead of an insecure image from a few months ago. It isn't just AWS, all Windows cloud servers seem to be delivered unpatched.
- gtsteve 8y agoCustomers probably wouldn't appreciate getting a Windows instance with an inconsistent patch level, you want to know that when you start a specific AMI ID you tested in your staging environment that it will be the same AMI you launch in production. Also they ship a new Windows instance approximately every 3-4 weeks with the most recent updates merged in. They don't make a big enough thing of it but you can subscribe to update announcements.
- drinane 8y agoI'm curious why the author does not think AWS will confirm his analysis. He seems to have hard evidence of how the system works and communications from staff.
- bastawhiz 8y agoI could see them declining to comment to avoid creating the expectation that the system works in a particular way. Being a black box means they can make changes without breaking expectations. Documenting internals, even informally, means folks may come to expect that behavior.
- Zombiethrowaway 8y agoI used Windows instances a few years ago. Beyond the slow start, once started, frequently the CPU would stay stuck at very low %, and my tasks would run very slowly. Eventually I would get to 100%, but it could often take 10 minutes. What I learned from those pains is how to use Linux in the Cloud.
- sl1ck731 8y agoI've experienced this lately with a variety of Amazon Windows images. For example I will boot a 2016 image from this year vs one from last year and last year's will be significantly faster on the same hardware.
- tachyonbeam 8y agoAny chance that might have to do with patches for the spectre vulnerability taking a performance toll? https://en.wikipedia.org/wiki/Spectre_(security_vulnerability) https://en.wikipedia.org/wiki/Spectre_(security_vulnerabilit...
- sl1ck731 8y agoIts possible. I didn't run any numbers or look at patch levels. It just went from running AD FS flawlessly on one to being barely usable over RDP on the other. Now I'm interested and might have to dig up which AMIs I've been through.
- vegardx 8y agoSeems like you were using T2 instances which have a low baseline performance and burst credits. I would imagine that you quickly run out of credits on some of the smaller instance types after creation, given how lengthy and costly (in terms of CPU usage) the instance creation and boot process is.
- Zombiethrowaway 8y agoI was typically using c3.xlarge for CPU-intensive tasks (video processing). Boot time was OK. I would log in on the machine with RDP, because sometimes my processes were almost frozen for a while. It felt like my neighbours were stealing my CPU, but I did not know how to prove it. Once I moved to Ubuntu, same instance type, I never experienced this.
- gtsteve 8y agoThis is an interesting thought; I use Windows instances myself but I use a custom AMI built using Packer on our CI server. Presumably Amazon doesn't have a pool of my custom AMI images lying around. So I wonder if the custom AMI is actually stored as a layer on top of the source Windows AMI and applied to the instance from the pool before it is made available to me. Alternatively, it means I'm missing out on an optimisation and I could get a faster start-up time by using a vanilla Windows AMI and installing dependencies in the user data. Does anyone know or have an educated guess?
- maxaf 8y agoA couple of jobs ago I have spent plenty of time optimizing a Packer-based AMI pipeline that was similar in spirit to that which you've described. It was all quite frustrating, but then the job got yanked from underneath me, so I never finished solving the problem. The article contains a screenshot[1] of a response from AWS support indicating that only EC2's own AMIs benefit from this pooling optimization. It follows that custom AMIs must take the startup time hit all the way. [1]: https://maishsk.com/blog/images/20190307_aws_ec2_secret/forum_post_slow.png https://maishsk.com/blog/images/20190307_aws_ec2_secret/foru...
- gtsteve 8y agoAh, I had missed that! Thank you, that's very helpful to know. I always learn something new every day here.
- Twirrim 8y ago> Seriously though - Windows images are big - absolutely massive compared to a Linux image - we are talking 30 times larger (on the best of days) so copying these large images to the hypervisor nodes takes time. No... just no. Disks aren't local to hypervisors. There's no copying going to be taking place. EC2 instances are provisioned using EBS volumes, which aren't going to be local to the instance itself. EBS is likely doing a disk clone operation, and those are relatively cheap in standard filer operations. Even large images you're talking a drop in the ocean in terms of the overall time. The main issues with Windows in a cloud environment comes from that first boot scenario. You could get past some of that by keeping a pool of warm instances around, but it'd require a lot of work on the Windows side to handle the provisioning use case. On Linux the instance boots, init processes kick off, and right at the end cloud-init creates and configures accounts with ssh keys etc. and away you go. Typically anywhere between 30-60 seconds boot time depending on the distribution. Windows isn't that accommodating. Images that are used in cloud environments have to be "generalized". You install it on specific hardware, and then tell it that it's not to give a crap about hardware specifics, but oh you must have these specific drivers etc. in you. It also tells it to clear up after yourself while it is at it. First boot happens, and Windows goes through a mandatory process whereby it identifies hardware, installs and configures drivers etc. etc. etc. You can inject a script to carry out the process of things like setting up a user, and generating a password, but that's fairly minor on the scale of things. This first process requires a reboot. There's no escaping it. That's really the reason why Windows provisioning takes long. There's unavoidable reboots, and unavoidable Windows driver installation and configuration, where the Linux distribution approach to the kernel makes life easier (every driver you're likely to need is a module and available on boot... unless you've got dracut running in the default host-only mode and it has made you a totally stripped down initramfs.) Add on that Windows booting takes longer than Linux even under optimal conditions and you end up with a slower launch.
- pjc50 8y ago> Windows goes through a mandatory process whereby it identifies hardware, installs and configures drivers etc. etc. etc. You can inject a script to carry out the process of things like setting up a user, and generating a password, but that's fairly minor on the scale of things. This first process requires a reboot. There's no escaping it. This is the "OOBE" (out of box experience) phase. I'd have thought they'd just skip it and provision you a pre-warmed image .. but I guess because there's a license/activation dependency they can't do that? I wonder how many MWh could be saved by Microsoft adding a "acquire cloud volume license at boot" mode. Or shoving the licensing/uniquification requirements into the platform TPM.
- osullivj 8y agoAlso note AWS Windows is not plain vanilla. It comes with an x64 Python 2.7 build built in, and the x64 python27.dll is on the standard path. Bit of a gotcha if you're deploying x86 Python 2.7 to AWS Windows!
- spacenick88 8y agoWhy on earth would you want to install 32 bit Python 2.7 if you already have 64 bit Python 2.7?
- lmz 8y ago32 bit C extensions?
- SonOfLilit 8y agoOne example that comes to mind is that you use an extension library that has Windows binaries available out of the box for x86 but not for x64.
- osullivj 8y agoYes - extensions. Particularly 32bit .xlls.
- sudovancity 8y ago>Now that I have got your attention with a catchy title - let me share with some of my thoughts regarding how AWS shines and how much your experience as a customer matters. This is the one guaranteed way to turn me off to whatever you're going to talk about in the article.
- lfx 8y agoI'm pretty sure GCE does the same, just last week I had a need to test something and with 1vCPU instance was ready in about 3-4 mins. Next time will check logs to confirm it. But seems the reasanoble thing to do.
- londons_explore 8y agoI don't think they have a pool of instances at all. It's a generalized image which they boot up for you. Cloning the image, even though it is many gigabytes, takes milliseconds since the underlying storage (EBS) will be some log-based storage. If they really wanted to optimize boot time, they would freeze the fully booted machine (keeping all the RAM contents) and then clone the frozen instance. That should be able to get running in just ~10 seconds (enough time to copy enough of the RAM contents to be able to log you in). They probably won't do that because having every user running from a fork of the same image could have some weird repercussions - for example the kASLR would be the same for all machines, making designing exploits much easier.
- cududa 8y agoWindows licensing mechanism prevents this scenario
- londons_explore 8y agoAmazon are big enough they could get Microsoft to rework how the technicals of licensing work.
- np_tedious 8y agoSafe to say this is only possible with default/official AMIs?