3 ms·
I'm loathed to perpetuate a 3 year old article, but... One of the key contributing factors to this kind of network degradation of AWS (/or other cloud vendor)
by dotBen 13y ago
I'm loathed to perpetuate a 3 year old article, but...
One of the key contributing factors to this kind of network degradation of AWS (/or other cloud vendor) is the abundance of the "bad neighbor test" - where a client performs tests to see if they can achieve a 'preferred' amount of CPU/IO on the host of their new instance.
Resource sharing rules at the host level actually means that if everyone is trying to max out their instance, you would still get the equal share you are entitled to and guaranteed with your instance, so what the bad neighbor test really means is whether you can actually go into your neighbor's CPU allocation due to their under-use.
Well, if everyone does that then the system degrades as someone has to a party using less than their allocation and the amount of instance spots that don't fail the 'bad neighbor test' become non existant.
The overall health of the entire network would actually be better if folks didn't do this practice and instead everyone simply evened out their use across their instances that enjoy additional resources and stuck it out with the %age of their instances that only achieved their guaranteed minimum resource use and no more.
My company uses another "cloud-like" vendor and although we don't perform 'bad neighbor' tests upon new instances, it is fair to say our application benefits from the fact that the majority of the instances on their network are under-utilized and we can push into the max CPU of the host beyond the limits we pay for. Where instances do share 'bad neighbors' (ie we can only get what we paid for and no more - boo hoo, etc) we still keep the instance but simply route around that and distribute less load than on other nodes in our network.
That doesn't become the most cost effective mechanism, but the "savings" of the 'bad neighbor test' are probably negligible and ironically by not doing this we become the "good neighbors".
- jamesaguilar 13y agoActually, if the system is well-isolated and correctly subscribed, at worst you'll get exactly your reservation. It is not necessarily the case that someone has to lose.
- dotBen 13y agoRight. The problem is that a 'bad neighbor' test is often defined as when one can only get that guaranteed minimum amount of resource. That's the part that doesn't scale.
- jamesaguilar 13y agoSeems like the simplest solution is simply to pay for what you need. How many businesses are really compute-cost-bound these days?
- michaelt 13y agoThe amazon documentation doesn't let you pin down exactly what that guaranteed minimum is, as half the measures are things like 'IO Performance: Moderate' or 'Compute Units: 4' or aren't specified at all (like EBS performance) Makes sense from Amazon's perspective, of course - less promises to keep, more flexibility. Can't blame people for measuring the performance empirically, in the absence of hard guarantees. Just that produces results that happen to be wrong.
- nucleardog 13y agoCompute units are actually a quantitative unit. One compute unit is, to the best of my recollection as I'm on mobile, the equivalent work of a specific class of 1.7GHz CPU. Other than the "I/O Performance", I think all of their specs are pretty well defined if you're willing to dig up the appropriate docs.
- michaelt 13y agoI suppose it depends on whether your application depends on ephemeral disk random or sequential I/O, EBS I/O, I/O to the internet at large, cpu cache, ram bandwidth, support for AVX instructions and so on. To be fair it's understandable why Amazon doesn't promise these features will or won't be present - it would make their already-complicated product offering even more complicated. And for a great many applications, customers won't be sensitive to details like CPU cache and disk performance.
- ultimoo 13y agoSince you are well-versed with this issue -- I know that Amazon offers tenancy options while creating virtual machines. Is this option not utilized because the price of single-tenancy is higher than the headache of "bad neighbor test"?
- dotBen 13y agoSo as I mentioned we don't use EC2, we use another cloud service, but the economics and technical issues are the same. But on EC2 Dedicate Instance (ie single tenancy) my guess is that if your application (or, business model) relies on each of your nodes being able to utilize more than your equal share on the host then in fact you would NOT want more than one of your instances to exist on the same physical host, in order to maximize the chances that each instance can grab all of the resources on it's given host. If this is your model, having all your instances on the same physical host would be disastrous. In fact, there's (economic) argument for Amazon offering customers the complete opposite - pay to guarantee that no two instances are ever instantiated on the same physical host.