4 ms·
A very much needed feature. Had a nightmare scenario in my previous startup where Google Cloud just killed all our servers and yanked out access. We got back ac
by theanirudh 3y ago
A very much needed feature. Had a nightmare scenario in my previous startup where Google Cloud just killed all our servers and yanked out access. We got back access in an hour or so, but we had to recreate all the servers. At that point we were taking Postgres base backups (to Google Cloud Storage) daily at 2:30 AM. The incident happened at around 15:00 so we had to replay the WAL for the period of about 12.5 hours. That was the slowest part and it took about 6-7 hours to get the DB back up. After that incident we started taking base backups every 6 hours.
- tticvs 3y agoDid you have any recourse against Google Cloud? Did you ever find out why they did that?
- theanirudh 3y agoI have forgotten the exact reason but it had something to do with not having a valid payment method. Some change on Google Cloud end triggered it - they were billing initially with the Singapore subsidiary and when they changed it to the India one, something had to be done from our end. Hardly got any notices and also we had around 100k USD in credits at the time. Got it resolved by reaching out to some high level executive contact we got via our investor. Their normal support is pretty useless.
- booi 3y agoi'm surprised the solution here isn't... moving out of google cloud. that is terrible
- idlephysicist 3y ago> Got it resolved by reaching out to some high level executive contact we got via our investor. Oh man that is my nightmare. Nothing says "broken system" like having to circumvent the system to get something done.
- shawabawa3 3y agoI've read about this happening a lot with google cloud If your payments fail for whatever reason google will happily kill your entire account after a few weeks with nothing other than a few email warnings (which obviously routinely get ignored)
- teaearlgraycold 3y ago> Google Cloud just killed all our servers > we were taking Postgres base backups (to Google Cloud Storage) Rule #1 of backups - do not host backups in the same location as the primary
- ddorian43 3y agoThe egrees fees will be bigger than your db cost.
- silon42 3y agoYes, maybe (some kind of diff/sync could maybe help), but this means using such a cloud is a bad IT practice.
- theanirudh 3y agoYes, the egress fees on base backups alone were higher than the cost of the DB VMs. If we replicate the WAL also, it would be way higher. In the post, the example DB was 4.3 GB, but the WAL created was 77 GB.
- sgarland 3y agoThe joys of WAL bloat [0]. UUIDv4s? [0]: https://www.2ndquadrant.com/en/blog/on-the-impact-of-full-page-writes/ https://www.2ndquadrant.com/en/blog/on-the-impact-of-full-pa...
- deleted 3y ago[deleted]
- scoot 3y agoThat's incorrect. You definitely do want backups in the same location as production if possible to enable rapid restore. You just don't want that to be your only copy. The canonical strategy is the 3-2-1 rule: three copies, two different media, one offsite; but there are variations, so I'd consider this the minimum.
- winrid 3y agowow! how big was the WAL? what kinda IOPS/disks are you using?
- theanirudh 3y agoDon't remember the size, but the disk we were using had the highest IOPS available on Google Cloud. That was one of the reason why we had restore from GCS since these disks wouldn't persist if the VM shut down. I think it's called Local SSDs [0]. We were aware of this limitation and had 2 standbys in place, but we didn't ever consider the situation when Google Cloud would lock us out of our account, without any warning. 0 - https://cloud.google.com/compute/docs/disks/local-ssd https://cloud.google.com/compute/docs/disks/local-ssd
- zilti 3y agoWe simply take incremental ZFS snapshots
- WolfOliver 3y agodo you need to stop the db for the backup in order to ensure consistency of the snapshot?
- kevincox 3y agoYou shouldn't because a filesystem snapshot should be equivalent to hard powering off the system. So any crash-safe program should be able to be backed up with just filesystem snapshots. There will likely be some recovery process after restoring / rollback as it is effectively an unclean shutdown but this is unlikely to be much slower than regular backup restoration.
- sgarland 3y agoNope, CoW is wonderful. Postgres will start up in crash recovery mode if you recover from a snapshot, but as long as you don’t have an insane amount of WAL to chew through, it’s fine.