5 ms·
As someone whose job involves maintaining uptime of a critical system that's dependent on Cosmos DB this sort of thing is scary. Where there's been other reliab
by DishyDev 4y ago
As someone whose job involves maintaining uptime of a critical system that's dependent on Cosmos DB this sort of thing is scary. Where there's been other reliability issues with Cosmos before we've not had an understanding customer base, and it feels very out of my control.
I'm finding a lot of the reliability guarantees of Azure PaaS services are overblown or come with big caveats when you start to work with them in a serious way. For example I've had some bad reliability issues with Azure Functions not firing, or their premium function runtimes becoming unresponsive. And it seems like that's just the start of the outstanding issues with them https://github.com/Azure/azure-functions-host/issues https://github.com/Azure/azure-functions-host/issues
I think people need to look more carefully at these PaaS guarantees and look at what that 99.999% reliability Microsoft are claiming actually means.
- rrdharan 4y agoEven as bad as their reliability issues are, I'd still be more worried about their security issues: https://www.wiz.io/blog/chaosdb-explained-azures-cosmos-db-vulnerability-walkthrough https://www.wiz.io/blog/chaosdb-explained-azures-cosmos-db-v... https://msrc-blog.microsoft.com/2021/08/27/update-on-vulnerability-in-the-azure-cosmos-db-jupyter-notebook-feature/ https://msrc-blog.microsoft.com/2021/08/27/update-on-vulnera...
- xiwenc 4y agoDo you know what blew me off? When azure executes maintenance on for instance postgresql servers, there is no record of that activity in the activity logs or anything to note in service health. The service was unavailable during the maintenance. And stronger yet when the database is unusable due to an incident the cpu is maxed out and it doesnt allow any successful connection, nothing is detected. How can this be a premium iaas/paas? Azure feels like the MS teams of tele conference. Companies buy in because they are already in the MS world. Not because azure is better.
- djbusby 4y agoPremium is a price point not a service level
- xiwenc 4y agoIndeed. Typical case where the purchasing people are not the same people that has to use it day to day.
- nijave 4y agoPostgres on Azure is terrible. Were you using Single Server, Flexible or Hyperscale/Citus?
- xiwenc 4y agoSingle servers now. We are considering to move to flexible server or just database as a service (non-server)
- nijave 4y agoYeah, that's the one we've had a lot of problems with. > And stronger yet when the database is unusable due to an incident the cpu is maxed out and it doesnt allow any successful connection, nothing is detected Apparently Azure's storage system that backs this uses some sort of thread pool and the thread pool can lock up/become exhausted leading to I/O starvation. When this happens, connection attempts fail. When the connection attempts fail, it can lead to a connection storm where all these new connections rolling in exhaust the CPU. The telltale indicator is Postgres checkpoints getting behind. All the while, the DB I/O metrics look like they're completely fine because it's not hitting an I/O limit, it's hitting thread pool exhaustion in the some storage system under the instance, outside of Postgres. You can also get some clues if this is the problem by enabling Performance Insights and checking the Waits tab. If all the top waits are related to I/O activity, that's another dead giveaway the storage system is locked up again. You can just web search the name of the waits to see what causes them. AWS has some nice docs detailing Postgres waits
- xiwenc 4y agoThanks for the detailed explanation! We didnt look into this so detailed yet but what you are describing sounds familiar. Since we have premium support (P1?), we had some internal azure postgresql engineer look at the issue and they pushed the problem back to us. Blaming our app not built correctly. That has been ping-ponging for over a year now. Finally i saw this semi-acknowledgment in their health status yesterday. Do you happen to know a proper solution? Are you waiting for them to fix this issue or moved to a different db service? Perhaps the flexible server is better?
- xiwenc 4y agoThe gods are angry i think. Woke up today and all our pg servers were unavailable. Checked service health and azure shows a global pg incident impacting pg servers. And the funny thing? Status.azure.com is all green. No events in activity overview. No service health within the affected instance. Workaround advised by azure? Upgrade to next plan. We already reached the maximum size. Maybe time for Citus . More $$$ for M$
- aoetalks 4y agoThat feels broken. Did you open a support ticket or a GH issue on the documentation page?
- ethbr0 4y ago> I think people need to look more carefully at these PaaS guarantees and look at what that 99.999% reliability Microsoft are claiming actually means. Hypercloud managed service SLAs: all the fun of novel complex, technical solutions in production + the transparency of cast iron + the pendanticism of being a contract lawyer Which leaves exactly zero people who are excited to be at that intersection.
- nijave 4y agoAfter using AWS for 3+ years and GCP for about 6 months, I can say Azure significantly lags behind them. Their service reliability is astonishingly poor. I think our most recent issue was 67 VM failures in a VMSS (of 55 nodes) backing AKS (Azure Kubernetes Service) in a single month. The health events said there were some kind of "remote storage errors" making the VMs unhealthy That's a couple months after the Ubuntu/systemd incident (Azure's "blessed" Linux image is Ubuntu and it has unatttended-upgrades enabled including on managed infrastructure like AKS (where you can't turn it off without dirty hacks). A bad Ubuntu update caused hosts to lose their DNS from DHCP config rendering massive amounts of machines in partially broken states) https://thenewstack.io/ubuntu-linux-and-azure-dns-problem-gives-azure-fits/ https://thenewstack.io/ubuntu-linux-and-azure-dns-problem-gi...
- wharfjumper 4y agoAsking as someone who knows nothing about CosmosDB other than the frequent negative comments I read on HN, do you know what led to you using it?