4 ms·
You're absolutely right that the cloud is not magic, but you do get some guarantees with EBS. From their website: "Each storage volume is automatically replica
by jespern 17y ago
You're absolutely right that the cloud is not magic, but you do get some guarantees with EBS. From their website:
"Each storage volume is automatically replicated within the same Availability Zone. This prevents data loss due to failure of any single hardware component."
We don't keep the database on the same EBS, and we have segmented database traffic out to several EBS volumes (for WAL, etc.) That's not the issue.
We take regular snapshot backups. We didn't lose any data. We have everything, we just can't get to it.
Regardless of what might make sense in this situation, it's not working for us. We've moved both our instances and the volumes to different availability zones, to no avail.
I just received a call from AWS engineering, assuring us that we are currently their top priority, and a team of engineers are working to fix the problem. They're seeing the issue on their end, and fortunately for them, it seems rather isolated to our instance.
Could we have taken precautions to prevent this problem? Maybe. We hadn't, cause we didn't anticipate a problem as exotic as this one. The only way to keep persistent data on EC2 is using EBS, and right now, it doesn't work for us, at all. This is not a common problem that could've been solved with backups or snapshots, or whatever.
- keefe 17y ago>The only way to keep persistent data on EC2 is using EBS, and right now, it doesn't work for us, at all. S3 should work too? Unless it was a global EBS failure, you should be able to restore from any backup to a new set of instances and stores, why doesn't that work?
- jespern 17y ago...As data you can access as a filesystem. S3 is great, but pretending it's a filesystem is going to get you awful performance. As I said, our data is not lost, we have snapshots and backups, it's sitting right there on the mount, we're just not getting any sort of acceptable throughput. New instances does not fix the problem. Ironically, we were looking into having S3 as the backend for our data, for scalability/redundancy purposes, but this pretty much puts a stop to that.
- gfodor 17y agoWhen we had the problem we fixed it by snapshotting the screwed up volume, and creating a new volume from that snap. Did you guys try this?
- jespern 17y ago(I don't know why I can't reply to the post below, so I'll reply to myself): Yes, we did try this, and it produced the same problem.
- keefe 17y agoOh, I wasn't suggesting pretending it's a file system - I had been thinking of a place to dump the data for backups, thinking fresh instances + fresh EBS would solve the problem. I think you answered this already in the other post - that you booted a new instance and a new EBS with some backup and the problem remained?? This seems like such a horrendous failure on AWS' part, unless it has something to do with how you are accessing the EBS (too many connections or something). I could understand if a given EBS fails, but if you can restore the data from an independent backup and spin back up with new instances and new EBS this indicates a very concerning systemic problem in EBS!