7 ms·
Announcing RDS/Aurora Snapshot Export to S3
- nitesh_aws 7y agoFull disclosure - I work for AWS and my opinions are my own
- sandGorgon 7y agoHmm..there's no restore? So this is like a one way export. Also Glacier export would be nice for long term cold storage of database backups. But a restore capability is essential.
- zitterbewegung 7y agoIt looks like this is more positioned to perform analysis on snapshot using other tools and not a backup solution .
- aquark 7y agoI’m interested in an easy backup-out-of-aws solution. We use snapshots for regular backups, and have implemented a solutions to share and copy snapshots to a different aws account. But given that our production database is absolutely vital to the existence of the company I’d like to have an easy way to create a secure ‘not affliated with Amazon’ copy regularly in case of the nightmare scenario where an account is compromised/terminated
- rsync 7y ago"I’m interested in an easy backup-out-of-aws solution." rsync.net has 'rclone' built into the environment[1] so it is trivially easy to call rclone, over SSH: ssh user@rsync.net rclone s3:/some/bucket /rsync/account ... and perform arbitrary data movements into, out of, and between S3/glacier/nearline/gdrive resources. Your "out-of-aws solution" contains ZFS snapshots, on any schedule you like, so you can just do a "dumb" sync to rsync.net and handle the retention with the snapshot schedule. Also, the snapshots are immutable (read-only) so your data gains some protection against ransomware and other threats like accidental deletion. You should trust us with your backups because we've been doing this since 2001. [1] https://rsync.net/products/rclone.html https://rsync.net/products/rclone.html
- aquark 7y agoGreat, thanks -- will take a look at this. The challenge with RDS in the past has always been getting the data easily and scriptable into an S3/etc bucket to start with. This might be a good enough solution. Restore would be a pain, but at the point its needed the pain would be acceptable.
- rsync 7y ago"Restore would be a pain, but at the point its needed the pain would be acceptable." I can't speak specifically to this, but I will point out that you can just as easily run those 'rclone' commands in reverse, to move data from an rsync.net snapshot out to (insert cloud): ssh user@rsync.net rclone .zfs/snapshot/1dayago/bucket s3:/some/bucket
- sandGorgon 7y agoAFAIK rsync.net doesnt work with RDS (which is a managed solution with no ssh access).
- rsync 7y agoI can't speak to RDS specifically, but if you can get RDS into either S3 or Glacier, then there is an easy path to rsync.net: rclone. The ssh access is not to Amazon, or S3 or RDS, etc. - you run the rclone command at rsync.net, over ssh: ssh user@rsync.net rclone s3:/some/bucket local/rsync.net/dir
- deleted 7y ago[deleted]
- aeyes 7y agoThe export does type conversion so this is one-way by definition :(. https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_ExportSnapshot.html https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_...
- zten 7y agoThe snapshots already deliver your desired functionality. This dumps it as a different format. In fact, it probably depends on restoring the snapshot to a database in order to implement this format change. I think this is a replacement for rolling your own database exports for analytics applications with Sqoop or Spark.
- jungturk 7y agoExporting to parquet has existed for a bit for AWS RDS (via AWS Data Migration Service), but this should make doing so more straightforward (since it doesn't require managing any DMS compute). AWS DMS also supports incrementals (using the change-data-capture features in the DB).
- zten 7y agoHmm, I wonder why they didn't push people towards DMS instead? This S3 export offering certainly commands a premium price for the privilege.
- anbotero 7y agoI just wish DMS supported real-time data synchronization for utf8mb4 on MySQL... That really destroyed my workflow with errors when I moved to utf8mb4 (requirement) and then had to stop using DMS altogether. It’s pretty much the only thing I used from that services.
- jungturk 7y agoThere are a number of unfortunate shortcomings depending on what your DMS source/target are. Another that bit us was (the lack of) support for JSON types in MySQL sources. Was there no option to add a transformation to your DMS routine to handle that?
- sandGorgon 7y agothere is a ridiculous price difference between vanilla s3 and managed backup storage. Its close to 10x. backup storage on postgresql is billed at $ $0.095 per GiB-month same on Aurora is billed at $0.023 per GB-month If you include Glacier, then it is 100x. I truly believe that AWS charges these costs because they make a crazy amount of premium with that. And that is why these exports are architected to be non-restorable.
- nitesh_aws 7y agoYes, that's correct, one way export (for now). Once exported you can set lifecycle policies in your S3 bucket to automatically transition to cold storage or deletion.
- joncrane 7y agoThe canonical AWS answer for things that you want in Glacier that only support export to S3 is to put a lifecycle policy on the bucket to move to Glacier after 24 hours. One day in S3 = 10 days in Glacier and since Glacier is presumably for years, the cost difference is trivial.
- dpedu 7y agoIt is not designed to be a backup mechanism.
- whalesalad 7y agoThis is a welcomed addition. Would love to see a “restore to existing DB” option. Sucks having to restore a backup to a new instance all the time.
- otterley 7y ago(Disclaimer: I work for AWS.) Can you tell us more about the use case? I'm not sure I understand the need here, or how it might improve your work patterns. Also, note that this is an export feature, not a backup/restore feature. Aurora already has native backup/restore: https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/Aurora.Managing.Backups.html https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide...
- verelo 7y agoI've run into this need. Generally, you have a DB, you just want to bring it back to a point in time. As RDS instances can be part of a bigger setup, it's not always trivial to point your application at a different instance, so re-using an existing instance can be helpful.
- bklyn11201 7y agoAmazon Aurora Backtrack? https://aws.amazon.com/about-aws/whats-new/2018/05/amazon-aurora-backtrack-can-move-a-database-back-in-time/ https://aws.amazon.com/about-aws/whats-new/2018/05/amazon-au...
- verelo 7y agoThats a great feature, if you're on Aurora. In my case these were MySQL RDS instances created ranging from 2012 and beyond. Migration of the app to Aurora is possible (after testing), but it would be a much larger change.
- otterley 7y agoWe recommend you do migrate - lots of great features in Aurora while remaining 100% MySQL compatible. (EDIT: I should have said "wire and protocol compatible" as opposed to "100%". I regret the error.)
- jbverschoor 7y agoNice! More database dumps exposed.
- johnhughesco 7y agoHas being that negative worked out for you IRL?
- bdcravens 7y agoI imagine a significant number of companies rolled their own solutions to dump to s3 already, so this is no less secure.
- scrollaway 7y agoAre there lightweight solutions to reading parquet files in Python? Any time I want to deal with parquet in AWS lambda I have to deal with the entire Pandas suite which is a pain on lambda.
- cavisne 7y agoOne interesting way is to use S3 Select which can read parquet, then you just need a dependency on the AWS sdk
- johnrob 7y agoQuick question (answer might be in docs somewhere): can you choose which tables to export?
- nitesh_aws 7y agoYes, you can filter. More instruction in the docs.
- nishantvyas 7y agoWhat happens to bandwidth saturation on the backup database/RDS host? is it unlimited? Capped? user-defined? i.e. would backup impacts the ongoing transactions/packet transfer?
- Tobani 7y agosnapshots would happen as normal in RDS. This is just the process of moving them "offsite" to s3.
- nishantvyas 7y agogot it. thanks.
- giseir 7y agoExcuse me for technical illiteracy, but would that make sense to export these snapshots daily? Is there any way to consistently get the (every 24h) data from the database in S3?
- giseir 7y agoLet’s say I have a database with 1TB of data. I export it daily to S3 with this Snapshot export. Does it mean I will be adding 1TB every day to my S3 storage?
- HatchedLake721 7y agoYes
- giseir 7y agoThanks, now I understand it. Then it makes sense to delete the earlier snapshot and only keep the latest one to not store redundant data.
- jungturk 7y agoYou can also use the CDC features of AWS Data Migration Service to just get incrementals (also in parquet if preferred) rather than full snapshots. https://aws.amazon.com/blogs/database/aws-dms-now-supports-native-cdc-support/ https://aws.amazon.com/blogs/database/aws-dms-now-supports-n...
- cmclaughlin 7y agoPaying for the full snapshot and not just the filtered data is a bummer
- etaioinshrdlu 7y agoI use rds Aurora but i run a daily cron job to export the entire thing with mysqldump + gzip to s3. I pass some flags to mysqldump to avoid locking the db and otherwise interfering with production. It also dumps from the reader node not the writer. I also clear some tables and restore it daily to a development DB. Sadly this looks like it only supports parquet, a rather usual db dump format. Even if, I'm sure, it's way more efficient to process using some modern tools than a text sql dump. I just like open formats and interoperability and AWS's rds offerings are always just slightly 'off'.
- sudhirj 7y agoParquet is an open format. It’s part of the Apache foundation. Would have preferred CSV as well, though, easier to work with. There’s also a little known feature called S3 Select that allows you to run SQL on S3 files and selectively retrieve data, which is what Parquet is for in the first place.
- etaioinshrdlu 7y agoCan you load that dump using any open source tools outside of AWS and recreate the database identically? Because you can do that with mysqldump.
- derision 7y agoI don't think this is meant for recreating the database, but rather exporting large sets for analysis somewhere else. But as far as tools yes, parquet can be parsed by any tool that knows the format. It's not too uncommon
- deleted 7y ago[deleted]
- samokhvalov 7y agoGreat news. But unfortunately, it is not the original data snapshot as the title says, it's Apache Parquet format. So this is useful for analytical tasks only. For operational tasks, still, the only option (besides RDS cloning) for Postgres is pg_dump or pg_transport, losing physical layout.