12 ms·
Ask HN: Azure has run out of compute – anyone else affected?
Last week we at n8n ran into problems getting a new database from Azure. After contacting support, it turns out that we can’t add instances to our k8s cluster either. Azure has told they'll have more capacity in April 2023(!) — but we’ll have to stop accepting new users in ~35 days if we don't get any more. These problems seem only in the German region, but setting up in a new region would be complicated for us.
We never thought our startup would be threatened by the unreliability of a company like Microsoft, or that they wouldn’t proactively inform us about this.
Is anyone else experiencing these problems?
- unreal37 4y agoI remember, in the early days of the pandemic, that Azure Australia ran out of compute too. It happens at the regional level. Are you stuck only to the German region, and can't go to other European regions?
- ksec 4y agoWorth repeating again, AWS, Azure and GCP are all adding capacity, and new Datacenter as fast as they could. We have enough demand to drive the next two generation of leading edge node. That is TSMC N3 and N2. And I assume it will be similar in N1.4 or 14A.
- krmboya 4y agoDamn! It's leaky abstractions again
- wizwit999 4y ago
- autophagian 4y agoClowns are much more solid than clouds, which are famously density-light. Given their traditional proximity to solid ground, too, clowns are a much better choice of foundational substrate than a cloud to build on.
- mianas 4y agoClown-car cluster sounds like it'd be a good name for a compute product.
- api 4y agoMaybe they are doing this to push people into regions with lower energy costs. Of course Northern Virginia or Canada is going to give you much higher ping times.
- sird 4y agoInteresting thought. It would be crazy if turning down business was preferable to just raising prices to reflect increased energy costs. I'm not a cloud expert, but maybe they don't have the infra to price differently in some regions?
- Scoundreller 4y agoThat’s the problem with charging average costs (assuming they do that) but the new user costs are at the margin which can be muuuuch higher.
- semicolon_storm 4y agoAzure definitely has the ability to charge differently per region. They do it pretty frequently.
- danpalmer 4y agoWhy wouldn't they just price higher in those regions? People want/need regions for policy and compliance reasons, not just for ping, particularly with Europe and Germany I'd expect.
- 4ndrewl 4y agoAnd potential data residency issues
- janober 4y agoI do honestly not think there is any bad intent behind it. I am just surprised that this is happening at all (esp. not with a resolution time of multiple months). They must have known for a long time that this would happen, so I would have expected an early heads-up!
- andrewstuart 4y agoYes it’s weird that you have to ask them for instances which some actual physical person looks at your request, thinks about it and says yes or no to. Instead of providing you with a list of the resources they do have, you have to play this weird game where you ask for specific instances in specific regions and then within several hours someone emails back to say yes or no. If it’s no, you have to guess again where you might get the instance you want and email them again and ask. I envisage going to an old shop, and asking the shopkeep for a compute instance in a region. He hobbles out the back, and after a long delay comes back and says “nope, don’t have no more of them, anything else you might want?”. It’s surprising this how it works. Not the auto scaling cloud computing used to bring to mind.
- gtirloni 4y agoIs this a joke comment?
- andrewstuart 4y agoNo, this is my actual experience using azure.
- gtirloni 4y agoI can go on Azure right now and create an instance and nobody will check anything manually and email me back something. Maybe you're confusing Azure with some other small town colocation provider.
- andrewstuart 4y agoNope, I went through this process exchanging more than 30 emails trying to get the instances I wanted.
- tsimionescu 4y agoIf you want 1 instance, you're right. If you want 10 - 20 instances of one type in a region, the other poster's experience matches my own: you have to open a support request to ask for a quota increase, and that is not an automated process.
- jedberg 4y ago> but setting up in a new region would be complicated for us. I've never done K8 on Azure, but my understanding is that Azure is pretty good about coordinating between your own datacenter running windows and Azure. Maybe you can spin up some windows boxes in a cheap datacenter to make it work?
- toomuchtodo 4y agoHetzner has a German presence I believe, and would work for running k8 on bare metal for n8n to burst to temporarily for running their orchestration and/or workflow runners. Might even be cheaper in the long run versus a cloud provider. Just gotta wire up the helm charts, containers, and whatever message bus is pushing their messages around. Can write to blob storage from anywhere if that’s a component of the app.
- zwaps 4y ago> Hetzner has a German presence I believe I sure hope so, as a German company
- Moissanite 4y agoAzure Germany is a separate partition from the rest of Azure - presumably for compliance reasons. This is distinct from AWS, where Frankfurt is just another region, albeit one with high demand.
- Terretta 4y ago> AWS .. Frankfurt is just another region Unlike GCP and Azure, all AWS regions are (were) partitioned by design. This "blast radius" is (was) fantastic for resilience, security, and data sovereignty. It is (was) incredibly easy to be compliant in AWS, not to mention the ruggedness benefits. AWS customers with more money than cloud engineers kept clamoring for cross-region capabilities ("Like GCP has!"), and in last couple years AWS has been adding some. Cloud customers should be careful what they wish for. If you count on it in the data center, and you don't see it in a well-architected cloud service provider, perhaps it's a legacy pattern best left on the datacenter floor. In this case, at some point hard partitioning could become tough to prove to audit and impossible to count on for resilience. UPDATE TO ADD: See my123's link below, first published 2022-11-16, super helpful even if familiar with their approach. PDF: https://docs.aws.amazon.com/pdfs/whitepapers/latest/aws-fault-isolation-boundaries/aws-fault-isolation-boundaries.pdf https://docs.aws.amazon.com/pdfs/whitepapers/latest/aws-faul...
- my123 4y agoCross-region extensibility points are few and far between. See https://docs.aws.amazon.com/whitepapers/latest/aws-fault-isolation-boundaries/abstract-and-introduction.html https://docs.aws.amazon.com/whitepapers/latest/aws-fault-iso... for more details.
- EE84M3i 4y agoAWS has several different levels of region isolation. There are aws region partitions - general, china, us gov cloud (public), us gov secret and us gov top-secret. Inside a partition, there can be some regions that are opt-in - see https://docs.aws.amazon.com/general/latest/gr/rande-manage.html#rande-manage-enable https://docs.aws.amazon.com/general/latest/gr/rande-manage.h... My understanding is that opt-in regions are even more isolated inside a specific partition for partition-global services like IAM and maybe some other stuff.
- xwowsersx 4y agoOof, that sucks and I feel for you. That said... > setting up in a new region would be complicated for us. Sounds to me like you've got a few weeks to get this working. Deprioritize all other work, get everyone working on this little DevOps/Infra project. You should've been multi-region from the outset, if not multi-cloud. When using the public cloud, we do tend to take it all for granted and don't even think about the fact that physical hardware is required for our clusters and that, yes, they can run out. Anyways, however hard getting another region set up may be, it seems you've no choice but to prioritize that work now. May also want to look into other cloud providers as well, depending on how practical or how overkill going multi-cloud may or may not be for your needs. I wish you luck.
- theteapot 4y ago> everyone working on this little DevOps/Infra project. Everyone? That's not going to help.
- deleted 4y ago[deleted]
- janober 4y agoThanks a lot! You are totally right, it is for sure something we will find a solution for. But honestly, do I not want to. As a startup, you have very few resources and deliberately place some exact bets. Deprioritizing everything to work on something for a long time that was not prioritized, just to then end up again where you were before (a working cloud solution) is the last thing any startup should be forced to do. Anyway, it seems like we do not have much choice here.
- xwowsersx 4y agoI hear you. It's not a fun position to be in. And sometimes you're correct to take calculated risks and maybe the expected value was positive here, despite what ended up happening. Without knowing the details about your services and infrastructure, it's hard for me to know what's involved in going multi-region now. Are you sure it's such a a gargantuan effort? I would've thought one person working full-time on this for a week or two would be enough, but again I don't know the details of your setup. One option would be to pay a consultant who is an expert in Azure/cloud stuff to come in and help. May not be cheap, but could be a lot better and quicker for you and better for the business, especially if none of you are really big experts in Azure. I've been here before (I think)...had to wear many hats and scramble to make sales, build the tech, act as de facto DevOps person even without a lot of experience doing it, etc. That is the way, but stuff happens. Happy to chat about specifics if you want to bounce ideas off of me or go through your particular situation. Can't promise I'll have concrete advice, but happy to talk it through.
- jenscow 4y agoMaybe Microsoft had just got their AWS bill?
- usgroup 4y agoWell I thought that was funny :-)
- ttrrooppeerr 4y ago> We never thought our startup would be threatened by the unreliability of a company like Microsoft You will be threatened by your own unreliability of building something that's dependant on one region or one cloud.
- websap 4y agoThis is an insidious argument to make. When building a startup you should choose 1 reliable cloud provider and use their best practices to support high availability.
- gtirloni 4y agoCross-region architecture will be the first thing you hear about.
- pclmulqdq 4y agoNo matter the provider, their best practices all say to be multi-region.
- websap 4y agoDef not true with AWS, unless you reach a particular scale. Not for product market fit. My technology choices would be fully managed services so I could focus on my actual business.
- Brian_K_White 4y agodef true with everything. what a ridiculous statement.
- bradknowles 4y agoRead the "Well Architected" paper. Go multi-region.
- websap 4y ago
- whalesalad 4y ago> We never thought our startup would be threatened by the unreliability of a company like Microsoft, or that they wouldn’t proactively inform us about this. Yikes, this is totally the first thing you need to come to expect when working with MSFT.
- bri3d 4y agoEvery cloud provider will have these issues with specific instance types in specific regions, although the Azure Germany situation sounds perhaps a bit more dire. At my past (much larger) employers we’ve always run into hardware capacity issues with AWS too - we’re just able to work around them. Building on cloud requires a lot of trade offs, one being a need for very robust cross-region capability and the ability to be flexible with what instance types your infrastructure requires. I’d use this as a driver to either invest in making your software multi regional or cloud agnostic. Multi regional will be easier. If you’re already on k8s you should have a head start here.
- PaulHoule 4y agoThere is a "minimal viable product" of documenting the configuration of your system so you can (1) run development, test, staging instances, (2) jump to another region when necessary, (3) from other disasters. Ideally you have a script that goes from credentials to the service to a complete working instance.
- Innominate 4y agoAs much as this happens, I don't feel it's something to be expected or even okay. The major cloud services are expensive. This extra cost is supposed to provide for cloud services' high level of flexibility. Running out of capacity should be a rare event and treated as a high priority problem to be fixed asap. Without the ability to rapidly and arbitrarily scale, they're just overpriced server farms.
- bagels 4y agoSome problems can't be fixed (eg. chip supply chain problems) even if you have more money.
- Havoc 4y ago>Some problems can't be fixed (eg. chip supply chain problems) even if you have more money. They can't magic chips into existence, but leaving a major region like Germany high & dry for almost half a year sounds like planning went wrong frankly. If it were a matter of chips I would have thought on a 3+ month timescale they can steal a few from another region that has a bit of fat
- pwarner 4y agoAzure, despite being smaller than AWS, I think has more regions. So each one must be smaller, which likely means less spare capacity. I also sort of suspect the spot market is less robust there. Lots of Azure is lift and shift on premises workloads, and those aren't using spot. Without people using spot, it's even harder to have spare capacity...
- pclmulqdq 4y agoAzure uses much smaller datacenters than AWS or GCP. Microsoft wasn't a big compute user before cloud, and it's a lot easier to manage and build for smaller DCs. Amazon and Google both needed huge DCs before being clouds.
- jeffbee 4y agoEC2 us-east-1 is chronically stocked out, too. Black Friday is the worst day of the year for this. At work, we pre-allocated tons of EC2 machines we don't really need, to hedge against EC2 stockout coinciding with some kind of incident. Yes, we are part of the problem.
- Jamie9912 4y agoPart of what problem? I don't remember us-east-1 ever running out of instances
- don-code 4y agoIn a former role, I used EC2 in us-east-1 to host the front door e-commerce site for a consumer electronics company. AWS suggested that we go through the Infrastructure Event Mangaement process (https://aws.amazon.com/premiumsupport/programs/iem/ https://aws.amazon.com/premiumsupport/programs/iem/) for Black Friday and Cyber Monday, so that staff on Amazon's side could guarantee that they'd have capacity to run our system at its forecasted peak. The strategy they helped us arrive at was two-pronged: 1. Pre-launch all needed infrastructure. Yes, for all their "cloud scale", it was actually suggested that we preallocate all of our servers the week before, rather than rely on autoscaling. 2. Order capacity reservations for all of those instances (https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capacity-reservations.html https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capa...). This ensure that, if any of those instances go bad, we'd be able to relaunch them without going to the back of the line, and finding out that there was no more compute capacity available.
- andrewstuart 4y agoMessage to cloud providers: List what you do you have available so we can choose. Do not force users to randomly guess and be refused until eventually finding something available.
- layer8 4y agoWhy would they make any promises, or be upfront about their resources at the risk of becoming less attractive compared to competitors with more resources? It’s not like many people are shunning the cloud for that reason today (although maybe they should).
- bushbaba 4y agoYour price point and the clouds margin is tied to not sitting on lots of unused instances. you want there to be adequate capacity not excessive capacity
- layer8 4y agoIt goes both ways: cloud providers don’t want to make promises about capacity, and cloud users don’t want to make promises about usage. I don’t know about price point. Dedicated servers can be cheaper than cloud in many cases, if you have the appropriate know-how, and the cloud business is very profitable for a reason.
- crmd 4y agoI need big m4n instances with 100gbe for product demos, and spinning them up lately is like trying to get Taylor Swift tickets on Ticketmaster. We end up wasting money running them for days at a time instead of on demand because we’re afraid of losing them. It’s infuriating that AWS doesn’t have an API that returns a list of AZs with available inventory for a given instance type.
- andrewstuart 4y agoWhy not run them elsewhere? There’s lots of providers apart from AWS/Azure/GCP. Or buy a machine and put it in your office. Self hosting can often be cheaper and more available and probably faster than using a cloud.
- deleted 4y ago[deleted]
- z3t4 4y agoI think there is some general rule in business that you should not depend on a provider that if they lost your business it would be less then one percent of their revenue. Or be ready when they drop the guillotine.
- ComputerGuru 4y agoFor anyone getting started, that means no dependencies at all. Even colocating would be out of the question, according to your metric.
- jodrellblank 4y agoUse Azure and AWS so that you're not dependent on either one. (You could depend on another startup with no revenue).
- iso1631 4y agoIn the olden days you use to buy computers from Dell and were well under 1% of their revenue. But if they dropped you as a customer, you bought them from HP instead, no problem.
- z3t4 4y agoIt depends on how long it would take you to find another colocating company. If there is another co-location on the other side of the street you could simply take your computer there - then there is no dependency.
- vaderade 4y agoYes, my company found this out trying to add both a database and a serverless app to our existing infrastructure in Germany West Central in July. They had no ETA for more GWC capacity back then and told us to move to the North and West Europe regions.
- donedealomg 4y ago
- jakear 4y agoPeople saying "shame on you for not being multi-region" are missing the point: This is a German company with German customers subject to German data residency laws. For them to store German data in a region besides Germany requires getting informed consent from the "data subject", who must be "pre-informed about the potential risks involved in cross-border data transfer". [1] This is why Azure has a dedicated German partition, just as it has a dedicated Chinese partition. Now, they could go the GDPR/Cookies route and prompt absolutely every user on pageload, but doing so would annihilate the purpose of the law into monotonous smithereens, just as it did with Cookies. Good on them for defaulting to the "more secure" mode, but yes this is a potential consequence. Happy to hear from any German amigos present if I've got something wrong. (But watch out... you might be putting HN at risk - their servers aren't (likely) in Germany!) [1]: https://incountry.com/blog/which-german-data-privacy-laws-you-need-to-comply-with/ https://incountry.com/blog/which-german-data-privacy-laws-yo...
- pclmulqdq 4y agoTime to sign up for AWS or GCP, then. If you're using kubernetes anyway, you'll be fine with the switch.
- scarface74 4y agoSays someone who has never done a large migration of any type…
- foobiekr 4y agoYou should be good to go except for debugging accounts billing monitoring habits documentation security evaluation …
- dilyevsky 4y agoI’ve had capacity issues before with both gcp and aws in smaller regions so not a panacea
- dehrmann 4y ago
- Jamie9912 4y agoNot the first time this has happened to Azure, they are always under-provisioned. Move to AWS
- arecurrence 4y agoThis is not as rare as public clouds may lead people to believe. I have had to move workloads around since AWS began (even between public clouds on occasion). In particular, GPU availability has been a continuing problem. Unlike interchangeable x64 / arm64 instances with some adjustments based on the new core and ram count... if no GPU instances are available then I simply cannot run the job. AMD's improved support has increasingly provided an alternative in some situations but the problem persists. I recommend doing the work to make the business somewhat cloud agnostic, or at the very least multi-region capable. I realize this is not an option for some services that have no equivalent on other clouds but you mentioned databases and k8s clusters which are both supported elsewhere.
- andrewstuart 4y agoGPUs are better run in your own office. All cloud providers charge much, much more for GPUs than if you run a local machine. Cloud GPUs are also a lot slower than state of the art consumer GPUs. Cloud GPUs: much slower, less available, much more expensive.
- bushbaba 4y agoSay you want 100 GPUs all inter connected to your multi petabyte data lake that’s being fed by your production workload. Sure you could buy all that equipment but I’d wager it’s cheaper, more agile, and greater velocity from it being in the cloud
- cosmic_quanta 4y agoI would argue that the cost profile is different. Local GPUs are a big up-front cost. But assuming that your workload is stable, in the long run I think local GPUs ends up being cheaper per-hour than cloud. For startups, it doesn't make sense to make the up-front purchase, fine. But if you're optimizing for long-term (amortized) costs, I'd be curious if cloud is cost-effective.
- 4y ago
- geonaut 4y agoAny optimisations you can make? Will have the advantage of saving you money across all platforms/regions
- janober 4y agoYes, some are possible and we are already doing that. Sadly will it only delay the time we run out of resources. If we would talk about a few weeks, we could for sure make it, but over 4 months is sadly not possible.
- option 4y agoI’ve heard from a friend who works at Microsoft that due to energy crisis in Europe plus their data locality laws, Microsoft is indeed running short on datacenter capacity there and can’t do anything about it no matter how much they are willing to spend.
- janober 4y agoVery interesting. As mentioned in another post, I am sure they are not trying to screw us or anybody else up. Is after all for sure not in their interest. But not flagging that to users I do not get at all. I would expect to get at least a warning email a week about that, plus warning in the dashboard, but there was literally nothing.
- option 4y agoas I’ve heard it actually affects much more than Azure, also all their cloud-based productivity suite.
- MattGaiser 4y agoCouldn't even put up their own solar panels?
- option 4y agolol, what?
- unixhero 4y agoBig fan of n8n!
- janober 4y agoThanks a lot! That is always great to hear!
- a99c43f2d565504 4y agoWhile you're at it at making your "infrastructure as code" cloud agnostic perhaps take a look at tools like Terraform (the only one I'm familiar with). I've just started the work of defining whatever we need to provision in their notation with the objective that it can be done with a single push of a button in the future.
- scarface74 4y agoThere is nothing “cloud agnostic” about using Terraform. Anyone who says this has no experience actually trying to implement it. Terraform has different providers for each cloud provider and the code is not transferable any more than saying if you use Python to script your infrastructure it will be transferable.
- robertlagrant 4y agoAgreed. I've advised people same before. You can build to Kubernetes cluster-agnostic (mostly), but the stuff that gets you to that point will be very cloud-specific. The reason for Terraform, and it's a good one, is your Terraform-related tooling doesn't have to change, e.g. if you route all your infra change approvals through Terraform Cloud), and you can coordinate multi-service changes, e.g. update Auth0 infra to do X, then AWS to do Y.
- janober 4y agoIt has actually been done that way. For technical reasons is sadly a move to a new data center even with that very complicated and time consuming.
- craigkerstiens 4y agoThis is nothing new, Azure has been having capacity problem for over a year now[1]. Germany is not the only region affected at all, it's the case for a number of instance types in some of their larger US regions as well. In the meantime you can still commit to reserved instances, there is just not a guarantee of getting those instances when you need them. The biggest advice I can give is 1. keep trying and grabbing capacity continuously, then run with more than what you need. 2. Explore migrating to another Azure region that runs less constrained. You mention a new region would be complicated, but it is likely much easier than another cloud. 1. https://www.zdnet.com/article/azures-capacity-limitations-are-continuing-what-can-customers-do/ https://www.zdnet.com/article/azures-capacity-limitations-ar...
- rsynnott 4y ago> In the meantime you can still commit to reserved instances, there is just not a guarantee of getting those instances when you need them. ... wait, what? How are they defining 'reserved'?
- alexeldeib 4y agoRI are a billing concept (discounted rates for long term commitment). Dedicated capacity exists, but it’s different (compute reservation groups or dedicated hosts). You can combine CRG/DH with RI for the desired effect, although IMO it’s a bit confusing. (Azure employee)
- robertlagrant 4y agoIt's a billing mechanism. You pay less if you guarantee use. Sadly, they don't guarantee availability of things to use :)
- rsynnott 4y agoYup, I'm aware of reserved instances (from an AWS PoV) but I always assumed they were, at least theoretically, well, reserved!
- 1-6 4y agoInfinite resources is only marketing and no hyperscaler on the market should ever promise that or give people that impression if they haven't accomplished scaling all throughout the entire supply chain.
- mmcconnell1618 4y agoI used to be a technical seller for Azure. This situation is obviously not great for you as a customer but there are proactive steps you can take to prevent this going forward. Reach out to your sales team and work with them on your roadmap for compute requirements going forward. The sales team has a forecast tool that feeds back into the department that buys and racks the equipment. If you can provide enough lead time, they will make sure you have compute resources available in your subscriptions.
- janober 4y agoThanks a lot, that is very helpful and great to know! Def. something we will do in the future.
- dustedcodes 4y agoWhy work with a human in the Azure sales department and plan cloud resources a year ahead? What’s the point of the cloud at this point? Then it just becomes a 100x more expensive version of hiring an infrastructure person and plan with them your own physical resources a year ahead.
- wstuartcl 4y agoWhat you describe is like the inverse of 90% of the reason companies host in the cloud. What makes needing to forcast and reach out to a sales guy to eventually stock hardware for your needs (while now competing against other customers for those resources) any better than hosting on prem. AWS for sure has had resource constraints in different AZs (especially during flack Friday and holiday loads) but I have never had an issue finding resources to spin up especially if I was willing to be flexible on vm type.
- mmcconnell1618 4y agoUnder most circumstance, this isn't needed unless you have a big ask. Say, you need 1,000 specific cores and GPUs, then this process is the best way to ensure you have them available. The original poster probably has the ability to spin up other instance types in their region. If there is no compute capacity in the entire region, something went wrong operationally. I'm not suggesting you should put in a request for every new resource you need, but if you have a specific instance type or a large number needed, it helps. You're not losing the ability to shut them down the next day if you don't need them, you're just telling the Azure team that you expect to spin some up around a certain time. If you're making a significant request of compute capacity, the team has the ability to reserve those instances for your subscriptions so that you're not competing with others for those cores.
- alexeldeib 4y agoWhat VM sizes? Besides what’s already been said, internal capacity differs HUGELY based on VM SKU. If you need GPUs or something it’ll be tough. But a lot of the newer v4/v5 general compute SKUs (D/Da/E/Ea/etc) have plenty of capacity in many regions. If changing regions sounds like a pain, consider gambling on other VM size availability. (azure employee)
- janober 4y agoActually nothing fancy, for sure no GPUs. Just Standard_E4s_v4.
- deleted 4y ago[deleted]
- alexeldeib 4y agoAh, bummer. If it helps, you can try this to list out VM sizes with comparable capabilities and see if you have better luck with any others (--all not really necessary since it filters by NotAvailableForSubscription and similar): az vm list-skus -l germanynorth -r virtualMachines --all > germanynorth.json jq '.[] | select( any( .capabilities[]; (.name == "vCPUs" and (.value | tonumber) >= 4 )) and any(.capabilities[]; (.name == "MemoryGB" and ( .value | tonumber ) >= 32) ) )' germanynorth.json 4/32 because that's what E4s_v4 would have.
- janober 4y agoThanks a lot! Just checked internally. Apparently are there some instances which we could get but would not work cost wise (have for example a lot of CPUs but we mainly care about RAM). Additionally, is there also still a region-wide CPU limit that would still cause us problems. So sadly not a long-term solution. But thanks a lot!
- fock 4y agolooking at the time you seem to spend on this issue and the fact you're apparently only needing low double digits of those instances. Are you really sure you shouldn't just buy a bunch of machines (500cores/2TiB go for ~60k€), throw them into a colo and then spend that time on actually doing stuff?
- victor106 4y agoI am sorry to say but at this point Azure is so f’ed up I think it should only be considered after AWS and GCP. The documentation is terrible and the Azure portal is so slow and laggy I can’t even believe it. Not to mention how unreliable their stack is.
- jesseryoung 4y agoRan into a similar issue last year in the East US region. We contracted support and they gave a similar response. From my understanding talking to people who use AWS and GCP this isn't uncommon across cloud platforms. While we could've just swapped a deployment parameter to deploy to another region, we opted to just use a different SKU of VMs for a short period and switch back to the VMs when they were available again. We haven't seen issues since.
- wstuartcl 4y agoyeah AWS tends to have capacity issues during high volume periods like black Friday (I think this is now actually because most large users pre reserve a buffer pool of vms sitting unused) -- but I have never had an issue where AWS has told me there would be no capacity for months. Its usually swapping AZ or regions or being slightly flexible on your sku. And if you are sensitive to this and find it happening take a look at your sku loadout you may be choosing a very high demand vm type and shifting just slightly gets way more capacity. ^^ and by capacity I am talking like 10's or 100s of vms being available not 1.
- ethotool 4y agoThere is no such thing as unlimited when it comes to resources and/or scalability in the hosting market. You might want to find a local colocation provider, buy a few network switches and servers as a secondary production and backup environment for your startup. Deploying your own infrastructure gives you full control over your startup. Yes it will raise your overhead and yes it’s not cheap but for a sustainable operation it’s a requirement in my opinion. I currently use Azure but I also have my own deployment with my own IP addresses and ASN which I keep spare capacity and keep some important servers on there incase something happens with Azure. Definitely helps me sleep better at night.
- cfeduke 4y agoI worked briefly in an enterprise facing sales organization that targeted multi-cloud deployments. Azure always had capacity problems. As ridiculous as it sounds, having an enterprise's applications exist on multi-cloud isn't terrible if the application is mission critical - not only does this get around Azure's constant provisioning issues but protects an organization from the rare provider failure. (Though multi-region AWS has never been a problem in my experience, there is a first time for everything.) Data transfer pricing between clouds is prohibitively expensive, especially when you consider the reason why you may want multi-cloud in the first place (e.g., it's easier to provision 1000+ instances on AWS than Azure for an Apache Spark cluster for a few minutes or hours execution - mostly irrelevant if your data lives in Azure Data Lake Storage).
- lmeyerov 4y agoAs part of launching our global GPU edge network, we need to support low-volume regions, which means a small number of T4 gpu in different timezones. Azure ran out last Christmas, or at least refused us capacity, and is only adding the next tier of A10's (~2x+ costlier?). We haven't had as much of a problem getting GPUs of different grade on GCP + AWS. I get a form email every 2w from Azure IT that they are working on it. Not as much of an issue for bigger GPUs. (Also... If into k8s, python, GPUs, graphs, viz, MLOps, working with sec/fraud/supplychain/gov/etc customers on cool deploys, and looking for a remote job, we are hiring for someone to take ownership here!)
- sabujp 4y agothis is due to the energy crisis in europe caused by the war
- robjan 4y agoWe've been having this problem in Singapore for a couple of years now. Can't add any VMs to our k8s cluster and can't provision a number of services which made our multi-region BCP more complicated.
- janober 4y agoYears?!?! Guess I then have to be happy that in our case it is "just" around 4 months.
- robertlagrant 4y agoInteresting semi-confirmed anecdote: when lockdown hit, Azure began to refuse to allocate servers. One of the main reasons was they prioritised servers in this way: 1. Government/health/defence cloud customers 2. Teams, which was exploding in use and they wanted to capitalise on it 3. Regular cloud customers
- isoprophlex 4y agoYeah this was real. I remember this. For a while they selectively deprioritized customers, like you say. I'm not judging, just confirming the observation.
- fuzzy2 4y agoSort-of. I have a Postgres flexible database in the West Germany Central region that can no longer be scaled. It was only created for testing purposes, so no biggie. The backend is basically a managed Compute resource. If you need more reliability, I see only one way out: Go multi-region or even multi-cloud.
- cyptus 4y agowe have the same issue and escalated it through multiple azure teams. our quota has been silently set to 0 while there where still instances running. this worked fine until auto-scale scaled the instances down in the night to 1. at the start of the day auto scale was not able to scale back up to the initial amount which did lead into heavy performance issues and outages. we needed to move the instances as azure support did not help us. after many calls with azure and multiple teams involved we finally did not get the quota approved (even if we did have it already and was not asking for „new“ quotas). also we decided to not be able to host in the German azure region anymore. Even if we could get the quota this is a business risk we don’t want to bear anymore to not be able to scale for unexpected traffic. this is huge for us as our application requires German servers. We are still in research where to host in future.
- cyptus 4y agoInteresting is that you can get instances in dev/test subscriptions without any trouble.
- jcmontx 4y agoWhy not just creating a bigger DB instance in another region for a few months? Sure, you'll take a performance hit, but 99% of users won't notice or care
- janober 4y agoAh yes, that is what we did in the end for the database. But that is not our main issue, rather that we do not get any more instances for our k8s cluster and those we can sadly not just spin up somewhere else.
- lyind 4y agoIf this is a serious problem for your business, you use K8s and require assistance quickly moving your workloads, consider contacting: https://www.giantswarm.io/ https://www.giantswarm.io/ (I work at Giant Swarm.)
- l-p 4y ago> We never thought our startup would be threatened by the unreliability of a company like Microsoft You're new to Azure I guess. I'm glad the outage I had yesterday was only the third major one this year, though the one in august made me lose days of traffic, months of back and forth with their support, and a good chunk of my sanity and patience in face of blatant documented lies and general incompetence. One consumer-grade fiber link is enough to serve my company's traffic and with two months of what we pay MS for their barely working cloud I could buy enough hardware to host our product for a year of two of sustained growth.
- roflyear 4y agoI have DB connection issues at least a few times a week. Annoying.
- janober 4y agoWe actually use Azure for ~2 years now. It worked the most time reasonably well, even though we had also a few issues. But our current issue + ready your and other comments will probably result in looking for a new home.
- marcosdumay 4y agoNew Microsoft customer at all.
- adrr 4y agoAzure has some of the biggest outages like when they went down on Feb29th for the whole day. https://azure.microsoft.com/en-us/blog/summary-of-windows-azure-service-disruption-on-feb-29th-2012/?cdn=disable https://azure.microsoft.com/en-us/blog/summary-of-windows-az...
- Godel_unicode 4y ago10 years ago, has there been something similar recently?
- flippingbits 4y ago
- lars_francke 4y agoWe have had this issue in and since 2018 https://www.opencore.com/blog/2018/6/cloud-has-a-limit/ https://www.opencore.com/blog/2018/6/cloud-has-a-limit/ That said: We also had this issue on GCP last month. We found that all three (AWS) are unreliable in their own ways.
- DannyBee 4y agoMost of Europe expects the winter to be quite painful from a power perspective. It would not be surprising if cloud providers (major power users) are being asked to not increase (or even decrease) power usage. The timeframe they gave would match that kind of ask. I wonder whether you see the same behavior from other cloud providers there (ie if you ask them whether new capacity is available, what do they say)
- arcturus17 4y ago> It would not be surprising if cloud providers (major power users) are being asked to not increase (or even decrease) power usage. I doubt it. It will be easier - and probably safer - to ask citizens and physical industry (eg, factories) to bear the brunt than to risk having problems in critical IT infrastructure. Ask people and factories to turn the heat 3 degrees down and the effects will be more or less predictable. Asking to shut compute power down at random will have unpredictable consequences.
- Nemo_bis 4y agoIt's not about shutting down existing machines. The power grid operator might be less willing to approve upgrades to serve increases in capacity. (No idea whether that's the case.)
- Nemo_bis 4y agoSomeone else confirmed my guess https://news.ycombinator.com/item?id=33744179 https://news.ycombinator.com/item?id=33744179
- arcturus17 4y agoThat makes more sense.
- analyst74 4y agoObviously Azure failed its customer here, but everyone with data centers in Europe is tightening their belts and preparing for the worst. I suspect AWS and GCP just have more headroom in EU.
- usgroup 4y agoPerhaps it’s a per customer limit to ration capacity? If so maybe you can legitimately work around it by creating multiple Azure billing accounts.
- janober 4y agoCould be possible. But as far as I know would two accounts in the same data center not work for us for technical reasons.
- RajT88 4y agoGet in touch with your CSAM. They will be able to get you assigned a capacity manager, if you don't already have one assigned. It is the function of the capacity manager to help you plan ahead based on what the data center capacities look like going into the future. Meet monthly with your capacity manager. Get representation across different technology interests - database, compute, storage, event hubs, etc. Don't ever skip these meetings.
- steelframe 4y ago> Get in touch with your CSAM Well that's an unfortunate acronym collision.
- RajT88 4y agoI thought for sure it was a military term (recalling SAM missiles), until I saw it in the news just today. GODDAMN. Sidebar: MSFT is the king of acronym collisions.
- andrewstuart 4y agoWow. It’s crazy that this could be valid advice, but it is.
- wstuartcl 4y agoNot much better than "meet with Infrastructure in Nov to plan next years capacity and server purchases" for on prep -- has Azure really degraded down to this?
- RajT88 4y agoIt's quite a bit better than that, in fact. They talk to their customers to try and understand all the big deployments coming to understand if there is going to be a crunch at the region/AZ level. I'd be surprised if other cloud providers aren't doing that in some form. I only have experience with Azure (so far).
- rickette 4y agoReading these comments it looks like everyone runs into this all the time. As a counterpoint: never run into this on Azure, scaling up/down 20-30 vm's a day. Hope it stays that way...
- zxcvbn4038 4y agoI’m sure Microsoft is just as surprised as you are. Almost every European facility I ever worked with was constrained by either space or power so you had to be really on top of your capacity management. Facilities in the US seem to have unlimited power and floor space so you never have to deal with either issue.
- Animats 4y agoIs there a secondary market for reselling Azure capacity? Can you bid against other Azure customers?
- kccqzy 4y agoStockouts have happened on both AWS and GCP too. Most of the time the problem is no longer a problem if you build your infrastructure not to rely on a single region or availability zone. On EC2 especially, even if you can't change to a different region, try changing to a new instance type and that might work.
- chunk_waffle 4y agoWho else has heard countless times something like "with company X's cloud platform you don't need to file a ticket and wait weeks for another team to provision a physical server, just spin some more up bro." The reality is you do, you've just outsourced the problem.
- natch 4y ago>We never thought our startup would be threatened by the unreliability of a company like Microsoft Had you never heard about (and this is unfortunately not a joke) Microsoft’s music service they once had, shut down after a few short years leaving customers without the ability to listen to the music they had paid to listen to? The service was called, this was the trademarked name, Microsoft “Plays for Sure.” You cannot make this stuff up.
- userbinator 4y agoThat's also the name of the DRM system it had.
- jobhenri 4y agohelp
- dharmab 4y agoI worked in a top 15 Azure customer. This is not unusual at all, especially in the newer regions. Talk to your TAM before you make attempt major capacity changes in a region. They may have advice on specific SKUs to use or which zones have capacity (e.g. when austrailaeast was being built 80%+ of the capacity was in one zone for many months). If you aren't a big spender you may not have a TAM who can get this info for you. Welcome to Azure.
- dszoboszlay 4y agoGood news is that today is Black Friday, so the e-commerce industry is running at peak capacity. In 30 days it will be Christmas, and by then (the very latest!) everybody will scale back, so you have a good chance to gain access to more compute before you reach the end of your runway.
- ig1 4y agoAsk your VCs/angels for help, this is the kind of thing they can definitely help with. (Speaking from experience - one of our portfolio companies had a similar challenge and we used our network to get to one of the execs of the vendor involved)
- janober 4y agoThanks a lot. Yes that is also something we are trying in parallel.
- omk 4y agoWhile some may immediately run a comparison between Azure, AWS and GCP let it be noted that any cloud platform facing this and making it to headlines is not good for the cloud industry over all.
- prmoustache 4y agoThe thing is: if cloud vendors struggle getting new machines, imagine your small company trying to order and get delivered new on-prem servers quickly. I worked for a company that worked mostly on-prem until 1y ago and last time they had ordered machines availability from Dell was scarce with huge delays.
- teaearlgraycold 4y agoAnd people roll their eyes when I say I’m dedicated to AWS.
- hgsgm 4y ago
- trasz2 4y ago>We never thought our startup would be threatened by the unreliability Daily reminder that cloud services are vastly less reliable than traditional hosting; it’s just that they manipulate the terminology to deflect that, replacing reliability with availability, aka “making impression of working”.
- somenewaccount1 4y agoIf you want help duplicating your k8s cluster workload, hmu. I love K8s and love contract work. $45/hr. Good luck!
- MildlySerious 4y agoThis is a bit tangential, but now might be a good time to experiment with raising the price of your product. It might extend the time you have until you have to stop accepting new users entirely, in case your migration is taking longer than needed.
- deleted 4y ago[deleted]
- Haga 4y ago
- rockylhotka 4y agoMy understanding is that the German region is not run by Microsoft, but a German company. This provides a legal shield required by Germany to try and prevent the US government from accessing data on those servers.
- aftbit 4y ago>These problems seem only in the German region, but setting up in a new region would be complicated for us. This seems like your fundamental problem. If you design an architecture that is limited to a single region of a single cloud provider, you are very likely to encounter issues at some point. Luckily you have a full month to solve this problem before it will prevent you from accepting new users. My suggestion is to start making your app multi-regional or multi-provider ASAP.
- habibur 4y agoLooks like github is down right now. Or is it only me?
- Epa095 4y agoIn Norway East Azure were incapable of provisioning new VMs for several(4-5) days, caused by some IP issue. The only solution was 'try to provision in the night, and don't turn it off if you get one'. Their status page showed green through the whole period though, even though nothing needing compute worked. So that was cool....
- purebscloudoff 4y agoThe Batch Service schedule history monitor sucks. It is inaccurate and doesn't sync the job order correctly. You can call them, they will get on the phone and then say they fixed it. Then you call them again because they didn't and they give you the same answer. Can't blame them, most of them are on H1B's. Nobody wants to be the squeaky wheel in that position. So you will just get the runaround all the time.
- AaronFriel 4y agoSurprised to see no mention of T-Systems, the subsidiary of Deutsche Telekom, that operates Azure Germany.
- calltrak 4y agoI am so glad we made the decision to pull https://Bigger.Bio https://Bigger.Bio off azure a while ago. It was nothing but problems on their platform.
- TexanFeller 4y agoInfinite scaling clouds, they said. In AWS at work we spin up large numbers of EMR nodes and every few days get stuck waiting for availability of certain instance types in our region too. I guess we could reserve more, but that defeats a lot of scale up and down advantages.
- mirekrusin 4y agoServerless runs out of servers.
- plantain 4y agoI'm baffled to read stories that suggest Azure is a viable competitor to GCP/AWS - they're an absolute nightmare on capacity. It took me six months to get approved to start six instances! With multiple escalations including being forcibly changed to invoice billing - for which they never match the invoices automatically, every payment requires we file a ticket.
- orik 4y agoWhat sort of nodes are you using, can you add a node pool with a different SKU?
- jiggawatts 4y agoOne of the biggest benefits of k8s is that you can easily mix in pools of different hardware types without a “rebuild”. Something to try in scenarios like this is to add the “weird and wonderful” VM SKUs that are less popular and may still have capacity remaining. For example, the HPC series like HBv2 or HBv3. Also try Lsv3 or Lasv3. Sure they’re a bit more expensive, but you only have use them until April.
- Kalanos 4y agoit's a european site during the world cup haha
- deathanatos 4y agoI've seen this before. I think it was in us-west1, ran out of VMs of the size we used for CI. Had to move to a different region. (Never moved back…) It is shocking to me that it happened at all. Capacity planning shouldn't be so far behind in a cloud that wants to position it as being on-par with AWS/GCP. (Which Azure absolutely isn't.) To me, having capacity planning be solved is part of what I am paying for in that higher price of the VM. > We never thought our startup would be threatened by the unreliability of a company like Microsoft, or that they wouldn’t proactively inform us about this. Oh my sweet summer child, welcome to Azure. Don't depend on them being proactive about anything; even depending on them to react is a mistake, e.g., they do not reliably post-mortem severe failures. (At least, externally. But as a customer, I want to know what you're doing to prevent $massive_failure from happening again, and time and time again they're just silent on that front.)
- exelib 4y agoM$ just don't want your money. We had experienced this problem many times in Irland and German regions. Never experienced it with Hetzner or AWS.
- HeavyStorm 4y agoI'm having trouble getting a instance with GPU in east US, but that's always a problem.
- rkwasny 4y agoThey basically have far too many small regions and are growing like crazy, multi-region deployments will be a must unfortunately. Maybe you can spin up some part of the infrastructure that are not latency sensitive in the nearby region?
- choward 4y agoHa. I knew something like this would happen eventually. Isn't limitless scalability one of the biggest selling points of using "the cloud"? If you have to buy your own computers anyway why even use the cloud? You could try using different clouds providers but eventually the clouds run out. Which brings me to another important point. If we run out of computers meaning supply can't keep up with demand, then who are the winners? The people who own the computers. Cloud providers and self hosters. Because of the high demand cloud providers can raise their prices and that's directly converted to profit since expenses remain the same, i.e. price gouging. Good job all you cloud loyalists who use the cloud for everything.
- SergeAx 4y ago"There's no such thing as cloud - it is always someone's else computer". Although we may try to rely on the unwillingness of the cloud provider to lose revenue, probability of events like this can never be fully discounted.
- ErnesTechDotCom 4y ago[dead]
- NicoJuicy 4y agoSame issue in France fyi
- Mave83 4y agoJust don't trust in marketing and save yourself a lot of money. On prem for all base or long term (6+ month) resources. Cloud only for peaks. And never use single cloud providers dependent features. Then you will never have such troubles at all.
- vipull 4y agoI’M. having a tough time ALso, with microsofft. They seem to IIgnoRe, then repent.: finally APologgise.:( I think u should switch to a new COMpuute. GCc.-pp.?? When we were running our own compute back in 09: and resources ran out or were unreliable, we cld shOUt at the server maintainer and/OOr install better hardware oUUrselves. NOt-THE.case anymore.:( :(( -Vip
- runamok 4y agoI don't have much knowledge about azure but is it possible to add different instant types and/or sizes? E. g. in the EC2 world if AWS was out of m5.xlarge I would try to add a worker group with m6.xlarge or m4.xlarge. If that did not work I might try to replace my xlarge with 2xlarge...
- YaBa 4y agoSad to ear that, but people have a wrong idea about the cloud, it's just others people hardware and like everything, there's a limit. They cannot warn you because it's very hard to predict how many new customers will come or if the existing ones will create more instances. I know about a bank with the same issue, basically, they've hogged all the resources in a specific region and yet, they need more. Unfortunately this things take time, MS cannot setup a new datacenter in a couple of days. >but setting up in a new region would be complicated for us. Why? it's easy: https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/move-resources-overview https://learn.microsoft.com/en-us/azure/azure-resource-manag... Latency issues from app to DB?
- aliswe 4y agoCreate a new nodepool/scaleset in another region (i think that should be possible)
- foxandmouse 4y agoIs this related to the hardware shortage during the pandemic? I'm assuming they couldn't scale at the rate which they intended pre-pandemic.. This seems like a much larger issue than they're making it seem. The promise of the cloud was unlimited scalability. I never thought of cloud resources as finite.