6 ms·
Dumb question, and I'm guessing this is a "if you don't know it's not for you" situation, but what is the point of a virtual tape? Isn't the point of a tape tha
by SCdF 4y ago
Dumb question, and I'm guessing this is a "if you don't know it's not for you" situation, but what is the point of a virtual tape? Isn't the point of a tape that it's not virtual? Or is this more replicating tape software apis (WINE/proton style) so you can get rid of physical tapes (because you no longer care about their physicality) without having to change your backup strategy?
- mbutler4 4y agoFrom the perspective of the AWS solution, this is a way of giving on-premises hardware an option versus Iron Mountain or similar for archiving to offsite for DR/BC purposes. Since AWS' storage options are typically pretty insane SLAs, this is acceptable (and certainly not much worse than having old LTO tapes in a storage locker). VTLs have been around for a long, long time. Initially, they were a way of speeding up tape by writing to a much faster media, and allowing for deduplication on the media to reduce storage costs. This was, of course, in the days of on-prem arrays and storage. Over time, LTO (the dominant tape format, I think we're on LTO-9?) became denser and faster, and the speeds at which you could write to a single drive would outpace the line speeds. Given that, methods were introduced to interleave multiple writers to a single drive/cart, and to leverage the drives speed. One thing you don't want on a tape drive is "shoeshining", this is the effect when the drive has no data to write, but the momentum of the reels in the drive mechanism carry the media forward, past the last written segment, past the write head. The drive must then backup and re index to find where it left off while queuing whatever data is incoming. Well, as VTLs were built initially to handle speed of writes and dedup, once the tape drives outpaced those issues, VTL became problematic. It was basically taking a randomly writeable media and turning it into a serial storage mechanism. This really wasn't much use, and VTL started to fade. Many solutions were built, one from NetBackup was OpenStorage; this was an API that a storage vendor could write a plugin for and allow NetBackup to write to their storage as intended: multiple writer, multiple access. Allowing for treatment of a deduplicable storage media as a filesystem versus an emulated tape drive. This API approach allowed for plugins to be constructed to talk to any storage, including cloud. AWS spoke with NetBackup at one point, intending to build a storage mechanism for a very solid solution: allow on-prem customers to export their data to AWS storage, AND to recover that data BACK from AWS storage (think of the data transfer fees you would get out of that!!). When they looked at the options, they decided to build the VTL gateway. I'm not sure of how or where the decision was made, but it seemed a bit odd at the time. As VTLs were fading and the limitations of emulating tape would lead to a caching requirement that might be bad, but with line speeds for interconnect to AWS making the caching requirements even worse. I'm not quire sure how that was done, but I will say, a VTL will always be a stable, solid thing, that tends to simply run and do what was intended; much like a traditional tape drive. As for some of the other discussion here, about reading old tapes and storage longevity; I've seen tapes kept for decades that were readable, Iron Mountain is still in business for a reason. Note that THEY don't have a solid cloud offering, but really, if you had everyones old data sitting around on tapes, and could charge not only for data delivery but media conversion at DR time, wouldn't you? And the issue with old tapes isn't so much the storage in that situation, it's the hardware. LTO only reads 2 generations backward (IIRC), and can only write one gen forwards. So, for your LTO-6 drive, it can read LTO-4 to 6, and can write LTO-6 to 7, again if memory serves. So, finding the Apollo code tapes on a format tape where the only possible drives available to read it live in a museum and may or may not actually work, make the fact that they recovered that source code amazing. Not to mention the software needed to be built to interface to hardware that the last person to build a "driver" for likely is sitting in a nice retirement joint in Sarasota.
- deleted 4y ago[deleted]
- thedougd 4y agoIt's the latter. In a large enterprise, the backup configuration can be wildly complicated, a matrix of systems, schedules, slas, etc. Just reconfigure your backup software to use this new virtual tape device and you're on your way.
- logifail 4y ago> Just reconfigure your backup software to use this new virtual tape device and you're on your way Isn't the whole point of tape that it is a physical thing, and may be taken offsite in a truck / stored in a vault, and all that jazz. "Just reconfiguring" your backup software sounds like a business might just bypass all that without necessarily realising the consequences. Then "we got hacked and our local backups are gone" -> "restore from offsite tape" -> "oh, actually there are no tapes, it's all the cloud now" -> ... ??
- orf 4y agoThen "we got hacked and our local backups are gone" -> "restore from offsite tape" -> "oh, actually all the tapes where stored at the wrong temperature or humidity and now they are fucked" -> ... ??
- kube-system 4y agoAWS's hard disks are a physical thing that is offsite
- logifail 4y ago> AWS's hard disks are a physical thing that is offsite Taking tapes to a secure offsite location means they're air-gapped from the source data, so even in a situation where an entire network is compromised, remotely-stored tapes can't be wiped/encrypted. I'm sure AWS has thought about this issue while designing Tape Gateway, but it's not clear to me how it could know whether a request to retrieve and overwrite a virtual tape was legitimate or not.
- 4y ago
- sithadmin 4y agoThe gateway appliance is fairly compatible with legacy backup solutions, which makes it a great drop-in replacement for physical tape library systems. There are certainly better backup methods available these days (though it's hard to beat the durability and cost of LTO tape for long-term archival), but I've seen AWS's virtual tape used as a good stopgap while other backup/recovery solutions are still a ways out an org's infrastructure roadmap.
- tablespoon 4y ago> though it's hard to beat the durability and cost of LTO tape for long-term archival I'm not entirely comfortable with the trend towards ever more esoteric and seemingly "all or nothing" technologies. Tape (and other physical things) may have their downsides, but I think there's something to be said for something that could (theoretically) be forgotten in a closet and read 50 years later. Likewise, AM radio may not be the best quality, but it seems like you can cover more area with a single installation than any other communication technology, which might come in handy after a serious disaster like a nuclear war (e.g. crank up the nighttime power of some undamaged countryside station running on generators to tell survivors where it's still safe).
- vbezhenar 4y agoIt's very unlikely that you'll be able to read modern tape from closet in 50 years. Tape requires very precise temperature and humidity for storage. Otherwise all bets are off.
- linksnapzz 4y agoOnly one datapoint; but I’ve seen an FPS/Celerity minicomputer boot off of a Qic-20 tape that’d been stored in a non climate controlled Albuquerque storage unit for 19 years…
- brudgers 4y agoEven for a new strategy, using a tape abstraction might simplify the design and/or implementation because backup tools and culture have their roots in tape. There's a sense in which tape is water in which backups swim.
- shrike 4y agoThe VTL was invented because, at the time (2012ish), none of the largest enterprise backup solutions had a good S3 interface and none supported Glacier. After talking with the backup software vendors (Spectrum, Tivoli, Symantec, Commvault, etc) it became clear that adding another backup target wasn't something we (AWS) could get them to prioritize, for perfectly reasonable reasons. We could (and did) apply pressure via our shared customers, even then they estimated it would take years. The fastest way to enable large enterprise access to S3 and Glacier for backups was to meet them where they were. We did this by virtualizing a tape library. Background: I'm one of the original inventors - https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/10013166 https://image-ppubs.uspto.gov/dirsearch-public/print/downloa.... I am no longer with AWS.
- deleted 4y ago[deleted]
- sshagent 4y agoYou did such a great job they still try and steer people this route. I've multiple times mentioned "you know we don't need to emulate tape drives any more"
- free-ideas 4y agoVTL has been around since the 90’s https://en.m.wikipedia.org/wiki/Virtual_tape_library https://en.m.wikipedia.org/wiki/Virtual_tape_library. This is a fantastic example of how broken the patent system is. This is not an invention it’s a good implementation of a well understood problem. It is also priced ridiculously - it is at least an order of magnitude more expensive than operating a physical library and remote physical storage that is beyond cyber threat by virtue of it being disconnected. A real backup is offline and offsite.
- Twirrim 4y agoYou did such a good job, that when I worked on Glacier (2013-2016), our interactions with the team running VTL was already pretty minimal. It was one of the least problematical things that interacted with Glacier. Some of the solutions that external vendors produced were nightmarish and left customers up the creek without a paddle in a disturbing number of situations.