25 ms·
I ended up paying $150 for a single 60GB download from Amazon Glacier
- detaro 11y agoSeems like "precise prediction and execution of Amazon Glacier operations" might be a niche product people would pay for (and probably already exists for enterprise use cases?) That's something that generally keeps me from using AWS and many other cloud services in many cases: the inability to enforce cost limits. For private/side project use I can live with losing performance/uptime due to a cost breaker kicking in. I can't live with accidentally generating massive bills without knowingly raising a limit.
- eropple 11y agoYou can do it pretty easily with the AWS APIs, and in the process scram-switch only the stuff you really want to kill.
- detaro 11y agoHow would you script a protecting against the issue described in the article? Unless you make sure to check the cost before each single request you can't stop it (and if you accidentally send one massive retrieval request even that isn't enough) For other services it is easier, but even then, setting up and managing my own cost control mechanism is a level of complexity (and risk of failure) I'd really want to avoid, esp. since I probably use AWS to avoid management overhead.
- eropple 11y agoYou can't in this particular use case, but I can't envisage AWS providing a cost-control system that would stop this, either. It doesn't make sense--they're not calculating costs as they go. What I'm saying is that what Amazon would provide you is not functionally different from what you can do yourself right now. I would be a lot more worried about a risk of over-charging myself if AWS wasn't incredibly good about refunding accidental overages.
- detaro 11y agoas user re pointed out above, apparently for Glacier that functionality actually exists: https://docs.aws.amazon.com/amazonglacier/latest/dev/data-retrieval-policy.html#data-retrieval-policy-using-console https://docs.aws.amazon.com/amazonglacier/latest/dev/data-re...
- eropple 11y agoHuh, that's...surprising. Glad it exists, though, as Glacier definitely is opaque. Thanks for pointing that post out (to you and to `re).
- deleted 11y ago[deleted]
- Nexxxeh 11y agohttp://docs.aws.amazon.com/amazonglacier/latest/dev/data-retrieval-policy.html#data-retrieval-policy-using-console http://docs.aws.amazon.com/amazonglacier/latest/dev/data-ret... and https://aws.amazon.com/blogs/aws/data-retrieval-policies-audit-logging-glacier/ https://aws.amazon.com/blogs/aws/data-retrieval-policies-aud... were linked elsewhere in the comments. Edit: By user?id=re.
- Ciantic 11y agoI concur. I have not tried variety of AWS services because I have no idea what it would cost me if something went haywire on my server. If I could simply deposit to Amazon a prepaid amount and it would just use this deposit until it's depleted, after which the services I have would grind to halt. This would be a perfect way for me to try it.
- deleted 11y ago[deleted]
- Zekio 11y agoPricing should always be made straight forward, easy to understand, and that pricing plan is dodgy as hell
- URSpider94 11y agoPricing plans FOR CONSUMERS should always be straightforward and easy to understand. This is not a consumer product. Pricing plans FOR B2B should model, as effectively as possible, the underlying costs -- this allows the provider to offer the lowest possible pricing for the services that cost them the least to provide, with expensive services priced accordingly. As others have mentioned on this thread, utilities are really, really good at this -- they come up with extremely complex rate plans for their largest customers that help them achieve whatever economies they are aiming for, for example incentivizing customers to provide level-loading (which is effectively what Amazon is doing in this retrieval scheme).
- jack9 11y ago> This is not a consumer product. That's a manufactured excuse for a fundamentally bad product api and pricing structure. When dropbox is more useful than AWS, amazon has screwed the pooch (which they do pretty often). Segmenting users by arbitrary circuitous logic into "consumers" (can't find a good use for it) and "enterprise" (can find a good use for it) isn't constructive. Both classes should avoid it, because it's not even an inexpensive choice, for what you get.
- manigandham 11y agoHow is Dropbox more useful? Dropbox and AWS Glacier are vastly different products and show exactly the divide between consumer and enterprise that you say doesn't exist.
- boulos 11y agoI think the reality is that most cloud customers are approximately consumers. Some big/sophisticated customers may want precise knobs to tune their spend versus "capability" but many businesses just want you to store their data. Pricing models like Glacier's are scary as hell to any CFO, unless they're shown a plan that says "So we're going to save $XXM on this, 100% certain". Adding on your utility company meme, there are different rate schedules for residential and business customers in most places. In Cloud, I think negotiated contracts with advanced customers probably make more sense than complex, unfriendly pricing models for everyone. Big disclaimer: I work on Compute Engine.
- res0nat0r 11y agoGlacier pricing has to be the most convoluted AWS pricing structure and can really screw you. Google Nearline is a much better option IMO. Seconds of retrieval time and still the same low price, and much easier to calculate your costs when looking into large downloads. https://cloud.google.com/storage/docs/nearline?hl=en https://cloud.google.com/storage/docs/nearline?hl=en
- _g9ex 11y agoThere's also Backblaze B2 (public beta): https://www.backblaze.com/b2/cloud-storage.html https://www.backblaze.com/b2/cloud-storage.html Their pricing is great (0.5c/month) but I'm a little worried about their single DC.
- andmarios 11y ago3 days ago they sent a newsletter about their new (alpha quality) b2sync tool which essentially is a “rsync to backblaze” utility. This makes their offer very interesting.
- josephagoss 11y agoCan this tool offer deduplication (so that changing a folder name does not re-upload thousands of files) or is that something I would have to code into my own backup solution?
- andmarios 11y agoNo, it is even less than rsync. It can only upload a folder recursively and skip files with common modification time on local and remote. Considering though that before that you couldn't even upload a directory, only files, this is a huge step. :)
- BorisMelnik 11y agoI currently need some help whipping up a solution for this.
- lazyant 11y agoYou don't have a backup until you test its restore.
- anonfunction 11y agoI've never heard that and I'm stealing it.
- ekimekim 11y agoCome on, don't downvote someone for learning something (you think is) well known for the first time. Especially when their response is "that's great and I'm going to use it". https://xkcd.com/1053/ https://xkcd.com/1053/
- lmm 11y agoThe comment is noise - no different from "+1" or "me too". If you found a comment helpful and want to thank someone for it, the way to do that is an upvote.
- anonfunction 11y agoIt's a bit different, I added that I had never heard of the phrase and that I would be using it.
- slyall 11y agoI've had some big Glacier bills in the past, even the upload pricing has gotchas[1] These days the infrequent access storage method is probably better for most people. It is about 50% more than Glacier (but still 40% of normal S3 cost) but is a lot closer in pricing structure to standard S3. Only use glacier if you spend a lot of time working out your numbers and are really sure your use case won't change. [1] - 5 cents per 1000 requests adds with with a lot of little files.
- wallacoloo 11y agoI use Infrequent Access Storage for backups, through a tool called duplicity (or more aptly, I use a GUI front-end for that tool called Deja Dup). Instead of uploading every individual file, it gathers them into 25 mb .tar files & uploads those along with an index describing where each actual file is. That has made the requests negligible for me.
- slyall 11y agoWe combined the files too. And then a customer wanted all their files rather than just one or two. Although that was billed back to the them.
- languagehacker 11y agoThis is a simple case of spending more than you should have because you didn't understand the service you were using. It's impacted a little worse by how silly the whole endeavor is, given the preponderance of music streaming services.
- joshmanders 11y agoUh no, actually this is really a case of a bug in OFFICIAL SDK that caused a higher bill than expected.
- detaro 11y agoNo, requesting all his data at once alone determined the rate. The retries didn't cost extra.
- natch 11y agoThis is why I break my large files uploaded to Glacier into 100MB chunks before uploading. If I ever need them, I have the option of getting them in a slow trickle.
- Nexxxeh 11y agoThis is no longer necessary, they now allow you to specify ranges of partial files. So you can split a single large file into multiple requests to keep you within allowance or within budget.
- natch 11y agoNice. I also upload par2 checksum files for each chunk so changing things now would involve a bit of script rewriting, but that's good to know for the future.
- thedufer 11y agoIs that necessary? They report a checksum for every file when you request an inventory.
- natch 11y agoNo, it's not necessary at all, I hope. I probably should have called them just "Par2" files not Par2 checksum files. They aren't just checksums. They allow reconstruction of the original file even after fairly extensive damage. Which isn't likely, so, yes, not strictly necessary, but who knows, could be useful in the what-if scenario. The checksums you mention only help with checking, not with repair. https://en.wikipedia.org/wiki/Parchive https://en.wikipedia.org/wiki/Parchive
- Pxtl 11y agoConsidering the fact that bugs in the official APIs resulted in multiple retry attempts, he should demand some of his money back.
- pmx 11y agoI have a strong feeling that he would get a refund if he contacted Amazon support, considering it was caused by a bug in the official SDK and he didn't ACTUALLY use the capacity he's being asked to pay for.
- revelation 11y agoI read it as the SDK had a bug, but that wasn't what caused the costs.
- hueving 11y agoMultiple requests didn't cause the issue. It was asking for all of his data to be queued in such a small window of time. The same thing would have happened without the bug.
- cdcarter 11y agoThat said, this is Amazon, who will refund you for a product if you ship it to the wrong address accidentally...I'm sure OP could get a refund.
- random3 11y agoSo depending on how the "average monthly storage" is computed you could get 20x more data in one month and then retrieve the 5% (previously 100%) that you care about for free, and then delete the additional data?
- cldellow 11y agoThere's a 90 day minimum storage cost for data.
- Spooky23 11y agoThe only use case I would be willing to commit to glacier would be legal-hold or similar compliance requirement. The idea would be that the data would either never be restored or you could compel someone else to foot the bill or using cost sharing as a negotiation lever. (Oh, you want all of our email for the last 10 years? Sure, you pick up the $X retrieval and processing costs) Few if any individuals have any business using the service. Nerds should use standard object storage or something like rsync.net. Normal people should use Backblaze/etc and be done with it.
- timv 11y agoBack when I worked in banking we had requirements like that (though we didn't use glacier) We had a legal requirement to be able to product up to 7 years worth of bank statements upon receipt of a subpoena. Not "reproduce the statements from your transactions records" but "give us a copy of the statement that you sent to this person 6.5 years ago" We had operational data stores that could generate a new statement for that time period, but if we received the subpoena then we needed to be able to produce the original, that included the (printed) address that we sent it to, etc. We had (online) records of "for account 12345, on 27th October 2011, we sent out a statement with id XYZ", we'd just need a way to pull up statement XYZ. There's no way(^) we'd ever get subpoenaed for more than 5% of our total statement records in a single month, so something like Glacier would have been a great fit. We had other imaging+workflow processes where we'd receive a fax/letter from a client requesting certain work be undertaken (e.g a change of address form). 90 days after the task was completed, you could be pretty sure that you wouldn't need to look at the imaged form again, but not 100% sure. We could have used glacier for that. We use case that would have cost us (rare, but we needed to plan for it) was "We just found that employee ABC was committing fraud. Pull up the original copies of all the work they did for the 3 years they worked here, and have someone check that they performed the actions as requested." Depending on circumstances & volume that might trigger some retrieval costs, but the net saving would almost certainly still be worth it. (^) Unless there was some sort of class action against us, but that's not a scenario we optimised for.
- bigiain 11y agoI'm happy enough to use it as a "third copy" - for never-expected-to-be-used recovery if both my local and remote backups fail. I know it'll take either a lot of time or money to restore from Glacier, but if my home and work backups have both gone I'll either not care about my data any more, or I'll be perfectly happy to throw a grand or so at Amazon to get my stuff back (or, more likely, be happy to wait up to 20 months for the final bits of my music and photo collections to come back to my own drives).
- profsnuggles 11y agoEven with the large data retrieval bill he still saves ~$100 vs the price of keeping that data in S3 over the same time period. Reading this honestly makes me think glacier could be great for a catastrophic failure backup.
- workitout 11y agoCompare to Google Drive at 100 Gigs for2 bucks and no bandwidth charges that I'm aware of.
- profsnuggles 11y agoI looked up the Google drive pricing and the costs increase quickly over that 100GB level. I'm going to use 3TB of data for my example because that is approximately the amount of data I would be backing up at work. The cost for Google Drive would be the $99 a month 10TB plan, Amazon Glacier is $21.51 a month. This is before you get into things like having the enterprise AWS ecosystem with IAM versus a single user Gmail account. Remember I am only talking about retrieval in the case of a catastrophic failure, the data is already backed up elsewhere. As long as I can manage to go a year without destroying all my backups Glacier comes out on top over Drive even taking into account the retrieval fees. In the best case scenario I never ever retrieve that data.
- sneak 11y agoThis article claims that glacier uses custom low-RPM hard disks, kept offline, to store data. Does s/he substantiate this claim in any way? AFAIK glacier's precise functioning is a trade secret and has never been publicly confirmed.
- olavgg 11y agoI think this is a good guess, http://storagemojo.com/2014/04/25/amazons-glacier-secret-bdxl/ http://storagemojo.com/2014/04/25/amazons-glacier-secret-bdx...
- olavgg 11y agoOh, there also a hacker news thread about this https://news.ycombinator.com/item?id=4416065 https://news.ycombinator.com/item?id=4416065
- elktea 11y agoI briefly used Glacier for daily backups as a failsafe if our internal tape backups failed when we needed them. The 4 hour inventory retrieval when I went to test the strategy and the bizarre pricing quickly make me look at other options.
- sathackr 11y agoThis sounds a lot like demand-billing [1] [2] that's common with electric utilities, particularly commercial, and increasingly, people with grid-tied solar installations. [citation needed] You pay a lower per-kilowatt-hour rate, but your demand rate for the entire month is based on the highest 15-minute average in the entire month, then applied to the entire month. You can easily double or triple your electric bill with only 15 minutes of full-power usage. I once got a demand bill from the power company that indicated a load that was 3 times the capacity of my circuit (1800 amps on a 600 amp service). It took me several days to get through to a representative that understood why that was not possible. [1] http://www.stem.com/resources/learning http://www.stem.com/resources/learning [2] http://www.askoncor.com/EN/Pages/FAQs/Billing-and-Rates-8.aspx http://www.askoncor.com/EN/Pages/FAQs/Billing-and-Rates-8.as...
- harryjo 11y agoImpressive that Amazon can choose to serve a request at 2x the bandwidth you need, with no advance notice, and charge you double the price for the privilege.
- thedufer 11y agoCan you explain what you mean here? I'm fairly certain it isn't true.
- re 11y agoGlacier's pricing structure is complicated, but fortunately it's now fairly straightforward to set up a policy to cap your data retrieval rate and limit your costs. This was only introduced a year ago, so if like Marko you started using Glacier before that it could be easy to miss, but it's probably something that anyone using Glacier should do. http://docs.aws.amazon.com/amazonglacier/latest/dev/data-retrieval-policy.html#data-retrieval-policy-using-console http://docs.aws.amazon.com/amazonglacier/latest/dev/data-ret... https://aws.amazon.com/blogs/aws/data-retrieval-policies-audit-logging-glacier/ https://aws.amazon.com/blogs/aws/data-retrieval-policies-aud...
- vonklaus 11y ago> fortunately it is fairly strait forward to cap your data retreival rates. Amazon has done a great job with this feature. By doing a poor job implementing something for an extremely narrow usecase, in a technology that is outdated and then providing the most complicated pricing structure surrounding every aspect of the product one can't helpbut use the feature: any other provider or service. Like, wtf would be the usecase for amazon glacier in 2016? I dont think I would put hubdreds of petabytes of sata into 20 year cold storage, and the author of this post certainly wouldnt use it again. The fact that i need to read 2 pages of pricing docs and then the 2 pages you linked to control them because I cant estimate them myself, is a sure sign this is absurd
- zcdziura 11y agoTape storage is still the most optimal form of long-term storage. If you need to store things for an exceptionally long time, such as financial data, scientific data, etc, then you're going to get the most bang for your buck on tape.
- barkingcat 11y agoYes but if you were in this target market you'd likely want your information on your own tape with a reading infrastructure you control.
- tomp 11y agoWouldn't it be better for the OP to simply upload 20 * 60GB (= 1.8TB) of random data, wait a month (paying less than 20 USD), and then download the initial 60GB within his 5% monthly limit?
- anonfunction 11y agoI don't like how the title and article reads like a hit piece on Amazon Glacier. It's great at what it is intended for. In addition it seems he still saved money because over 3 years because the $9 a month savings added up to more than the $150 bill for retrieval. I'm surprised that this aspect has not been mentioned here in the comments yet: > I was initiating the same 150 retrievals, over and over again, in the same order. This was the actual problem that resulted in the large cost. At my old job we would get a lot of complaints about overage charges based on usage to our paid API. It wasn't as complicated of pricing as a lot of AWS services, just x req / month and $0.0x per req after that, but every billing cycle someone would complain that we overcharged them. We would then look through our logs to confirm they had indeed made the requests and provide the client with these logs.
- ricardobeat 11y ago> I was initiating the same 150 retrievals, over and over again, in the same order. >> This was the actual problem that resulted in the large cost. That is not true, according to the article itself. The first request for the full 60.8gb already results in a $154.25 bill, regardless of the ones that follow. From that point on he can continue to retrieve 15.2GB/hour for the rest of the month without incurring further costs.
- anonfunction 11y agoThanks for clarifying, I guess I missed that part.
- detaro 11y ago> > I was initiating the same 150 retrievals, over and over again, in the same order. This was the actual problem that resulted in the large cost. Except that it wasn't. The repeated requests were free, because he already set the maximum rate with the first wave of requests. Surprising? Also, it really is not a hit piece. It's an honest report of what he did (and what he did wrong) and that he thinks the docs aren't as clear as they could be.
- 11y ago
- otakucode 11y agoI'm surprised that the author had 150GB of Creative Commons audio CDs to begin with!
- detaro 11y agoDon't assume your local laws apply everywhere in the world. (Hint: they probably don't in Finland, which allows private copies)
- sundvor 11y agoAre you trying to say that the OP was not allowed to back up their own CDs? I have several boxes of space wasting CDs of my own that have all been ripped in lossless, and feel somewhat insulted by that notion. Note: I live in Australia, a somewhat less insane country for DRM than the US (AFAIK).
- JupiterMoon 11y agoCDs don't even have DRM in the general case. However, local laws can still forbid ripping (e.g. here in the UK where after a brief period of legality it is once again illegal to rip a CD). EDIT removed material that might have derailed things further.
- sundvor 11y agoRoger that; I used the DRM term incorrectly, meaning copyright law. Most of my collection was bought in Norway, but I did the ripping a decade ago plus minus a few years in Australia. For the time being we have much better consumer protection here. I don't even use my collection now that I have Spotify Premium. The only music I've bought lately is some 24bit high bitrate stuff.
- nness 11y agoAlso in Australia, and its not that black-and-white in my understanding: This is known as "Format Shifting" — taking one copyrighted medium and converting it to another. In Australia, you are explicitly not allowed to do this with CDs, DVDs and Blu-rays. You are only allowed to keep a digital copy if you continue to retain the original — a backup. If the original is lost or destroyed, your digital copies must be discarded. For example, you can rip a CD and put it on your iPod, or computer, as long as you continue to own the CD. The issue here is that in both cases you also control the device you are copying it to. You don't control but rather lease space on Amazon's servers — so it introduces a grey area on whether you are allowed to backup to such places and whether putting data on those servers constitutes distribution of the copyright material. Realistically, none of this is black-and-white and Amazon could flag it as infringing content and remove it just to cover themselves against DMCA complaints anyway. This is true in both Australia and the US, regardless of the differences in copyright law (Australian copyright law offers far fewer protections than the US, incidentally) because both having similar DMCA laws.
- kennu 11y agoGlacier is more comfortable to use through S3, where you upload and download files with the regular S3 console, and just set their storage class to Glacier with a lifecycle rule. I've used the instructions in here to do it: https://aws.amazon.com/blogs/aws/archive-s3-to-glacier/ https://aws.amazon.com/blogs/aws/archive-s3-to-glacier/
- takeda 11y agoThat can bite you if you have a lot of small files[1], which with automatic archiving of S3 can happen. [1] https://therub.org/2015/11/18/glacier-costlier-than-s3-for-small-files/ https://therub.org/2015/11/18/glacier-costlier-than-s3-for-s...
- prirun 11y agoAmazon makes it very easy to transition data from S3 to Glacier. But, if you want to get it back, you will still have to go through all of the byzantine Glacier rules and fees, including the 4-hour wait. And, there is no "transition from Glacier to S3": if you want to do that, you have to: 1. restore it to S3 (and incur the fees and 4-hour wait) 2. copy the restored S3 object to a new S3 object 3. delete the restored object (or wait for it to timeout)
- Nexxxeh 11y agoThe post is a useful cautionary tale, and he's not alone in getting burned by Glacier pricing. Unfortunately it was OP not reading the docs properly. Yes, the docs are imperfect (and were likely worse back in the day). And it was compounded by the bug, apparently. But it's what everyone on HN has learned in one way or another... RTFM. Was it mentioned in the article that the retrieval pricing is spread over four hours, and you can request partial chunks of a file? Heck, you can retrieve always all your data from Glacier for free if you're willing to wait long enough. And if it's a LOT of data, you can even pay and they'll ship it on a hardware storage device (Amazon Snowball). Anyone can screw up, I'm sure we all have done, goodness knows I have. But at the very least, pay attention to the pricing section, especially if it links to an FAQ.
- magicmu 11y agoLike you said, a good cautionary tale. Even this incident, although definitely rough, isn't that bad compared to what can happen if you scale a full infrastructure improperly, or just overlook something involving a software-based business. Just a couple of months ago, a test EC2 instance that had been forgotten about got spun up by a very archaic piece of code, and ended up costing the guys I'm working with quite a bit. Always RTFM (especially AWS pricing docs) carefully, or feel the burn.
- FireBeyond 11y agoI don't think he really was "burned". Paying 87c a month for a couple of years, then 52c a month for a few more years, to back up 60+GB, and then getting a one off fee of $150 still averages out at around $2/mo. Hardly getting shafted.
- Nexxxeh 11y agoDepends on your budget. If you don't have $150 disposable income to blow on a screwup, it's an expensive mistake. It's like getting a parking ticket for parking somewhere you've unknowingly been parking illegally for years. Yeah, if you average it out, it's cheap. But having to stump up the cash still hurts.
- newsburn 11y agoAnd I paid $5 for 540GB worth of downloads from naughtyamerica.com
- cm2187 11y agoPerhaps a naive question but why would glacier try to discourage bulk retrieval? Is it because the data is fragmented physically?
- syncsynchalt 11y agoWe don't actually know how glacier data is stored (though there are several theories from regular s3 to tape robots to experimental optical media).
- ufmace 11y agoI suppose part of the neat trick of it is that because we don't have to know, Amazon can switch it out for something else anytime it's convenient for them, or some new tech comes up. Or split thing up among several methods and compare costs. As long as they structure their operations such that nothing is ever lost and everything can be retrieved with a few hours notice at any time, they can try anything they want.
- cpayne 11y agoI'm sure that was a commercial decision, not a technical one. They could be running Glacier storage at cost (or even a slight loss). But they make their profit when you try to get your data out. I'm not implying anything nefarious - more along the lines that Amazon (could have!) looked at the market and compared demand to what they were offering. Then found something that scratched an itch...
- forgotpwtomain 11y ago> I’d need more than one drive, preferably not using HFS+, and a maintenance regimen to keep them in working order. I'm really doubting the need for a maintenance regimen on a drive which is almost entirely unused. Could have spent $50 on a magnetic-disk-drive and saved yourself hours worth of trouble.
- jaytagdamian 11y agoThe problem with physical drives is that they can (and do!) fail. The authors point is surely that the backup would need to be checked periodically, and drive failures dealt with. Does magnetic media like this (especially spinning disk) suffer from bit-rot? What about the possibility of mechanical failure? I'd never rely on mechanical disks as the one and only backup of any data critical to me - a two tier approach of mechanical for fast retrieval, and cloud/online backup seems to be the safest bet.
- lmm 11y agoYou absolutely need a maintenance regimen if you want to be able to reliably retrieve files 4 years later. If you just unplug the drive and stick it in your garage, sure it'll probably work, but I'd say there's at least a 5% chance of failure.
- Jedd 11y agoAbout a year ago NetApp bought Riverbed's old SteelStore (nee Whitewater) product -- it's an enterprise-grade front-end to using Glacier (and other nearline storage systems). It provide a nice cached index via a web GUI that let you queue up restores in a fairly painless way. It even had smarts in there to let you throttle your restores to stay under the magical 5% free retrieval quota. It's not a cheap product, and obviously overkill for a one-off throw of 60GB of non-critical data ... but point being there are some good interfaces to Glacier, and roll-your-own shell scripts probably aren't. As noted by others here, if you treat glacier as a restore-of-absolute-last-resort, you'll have a happier time of it. Perhaps I'm being churlish, but I railed at a few things in this article: If you're concerned about music quality / longevity / (future) portability - why convert your audio collection AAC? Assuming ~650MB per CD, and the 150 CD's quoted, and ~50% reduction using FLAC, I get just shy of 50GB total storage requirements -- compared to the 63GB 'apple lossless' quoted. (Again, why the appeal of proprietary formats for long term storage and future re-encoding?) I know 2012 was an awfully long time ago, but were external mag disks really that onerous back then, in terms of price and management of redundant copies? How was the OP's other critical data being stored (presumably not on glacier). F.e. my photo collection has been larger than 60GB since way before 2012. Why not just keep the box of CD's in the garage / under the bed / in the attic? SPOF, understood. But world+dog is ditching their physical CD's, so replacements are now easy and inexpensive to re-acquire. If you can't tell the difference between high-quality audio and originals now - why would you think your hearing is going to improve over the next decade such that you can discern a difference? And if you're going to buy a service, why forego exploring and understanding the costs of using same?
- stordoff 11y ago> Assuming ~650MB per CD, and the 150 CD's quoted, and ~50% reduction using FLAC, I get just shy of 50GB total storage requirements -- compared to the 63GB 'apple lossless' quoted. (Again, why the appeal of proprietary formats for long term storage and future re-encoding?) I did a comparison between FLAC and ALAC (a.k.a. Apple Lossless) on my CD library a few years ago (plus a few 48kHz tracks taken from DVDs), and the difference in total filesize was less than 10% so I doubt that is a major factor. I personally went for ALAC, as it has equal (EAC, VLC) or better support (OS X Finder, iTunes, Windows Explorer, Windows 10 media player, some tagging scripts, iOS) in stuff I currently use. Providing I keep a decoder with the files, its proprietary nature doesn't really bother me - I can always convert to xLAC if desired.
- stuaxo 11y ago"The problem turned out to be a bug in AWS official Python SDK, boto." My only experience of using boto was not good. Between point versions they would move the API all over the place, and being amazon some requests take ages to complete. After that worked with google APIs which were a better, but still not what I'd describe as fantastic (hopefully things are better over last 2 years).
- astrostl 11y agoArq has a fantastic Glacier restore mechanism. You select a transfer rate with a slider, and it informs you how much it will cost and how long it will take to retrieve. It optimizes this with an every-four-hours sequencing as well. See https://www.arqbackup.com/documentation/pages/restoring_from_glacier.html https://www.arqbackup.com/documentation/pages/restoring_from... for reference.
- copperx 11y agoIt's unfortunate that Arq forces you to archive in their proprietary format. That locks you in to the tool.
- TkTech 11y agoNo it doesn't. They've had an open-source CLI tool since at least 2013. https://github.com/sreitshamer/arq_restore https://github.com/sreitshamer/arq_restore
- asymptotic 11y agoI was also concerned about this when I was looking into Arq, so I wrote a cross-platform restoration tool that'll also work on Windows and Linux (not just Mac): https://github.com/asimihsan/arqinator https://github.com/asimihsan/arqinator This is purely based on the author's excellent description of the format in his arq_restore tool: https://www.arqbackup.com/s3_data_format.txt https://www.arqbackup.com/s3_data_format.txt
- deleted 11y ago[deleted]
- arcdigital 11y agoI just started using Arq, but it looks like the format is documented here https://www.arqbackup.com/s3_data_format.txt https://www.arqbackup.com/s3_data_format.txt and there's some open source tools that can work with the format. https://github.com/asimihsan/arqinator https://github.com/asimihsan/arqinator and https://godoc.org/code.google.com/p/rsc/arq/arqfs https://godoc.org/code.google.com/p/rsc/arq/arqfs
- LukeHoersten 11y agoDoes anyone have a success story for this type of backup and retrieval on another service?
- JimmaDaRustla 11y agoWow, thanks for this! I currently have 100gb of photos on Glacier. I am going to be finding another hosting provider now.
- deleted 11y ago[deleted]
- pfarnsworth 11y agoIf there's a bug in Amazon's libraries, can't you ask for a refund?
- roywiggins 11y agoOnce he was charged the $150, it didn't cost him any more to try again to download it, because of how the pricing works. So, he would have been charged that much no matter what if he'd downloaded the data at that speed, and there was nothing to refund once he successfully got his data out.
- atarian 11y agoThey should rename this service to Amazon Iceberg
- markonen 11y agoOP here. Some updates and clarifications are in order! First of all, I just woke up (it’s morning here in Helsinki) and found a nice email from Amazon letting me know that they had refunded the retrieval cost to my account. They also acknowledged the need to clarify the charges on their product pages. This obviously makes me happy, but I would caution against taking this as a signal that Amazon will bail you out in case you mess up like I did. It continues to be up to us to fully understand the products and associated liabilities we sign up for. I didn't request a refund because I frankly didn't think I had a case. The only angle I considered pursuing was the boto bug. Even though it didn't increase my bill, it stopped me from getting my files quickly. And getting them quickly was what I was paying the huge premium for. That said, here are some comments on specific issues raised in this thread: - Using Arq or S3's lifecycle policies would have made a huge difference in my retrieval experience. Unfortunately for me, those options didn't exist when I first uploaded the archives, and switching to them would have involved the same sort of retrieval process I described in the post. - During my investigation and even my visits to the AWS console, I saw plenty of tools and options for limiting retrieval rates and costs. The problem was that since my mental model had the maximum cost at less than a dollar, I didn't pay attention. I imagined that the tools were there for people with terabytes or petabytes of archives, not for me with just 60GB. - I continue to believe that “starting at $0.011 per gigabyte” is not a honest way of describing the data retrieval costs of Glacier, especially when the actual cost is detailed, of all things, as an answer to a FAQ question. I hammer on this point because I don't think other AWS products have this problem. - I obviously don't think it's against the law here in Finland to migrate content off your legally bought CDs and then throw the CDs out. Selling the originals, or even giving them away to friend, might have been a different story. But as pointed out in the thread, your mileage will vary. - I am a very happy AWS customer, and my business will continue to spend tens of thousands a year on AWS services. That goes to something boulos said in the thread: "I think the reality is that most cloud customers are approximately consumers". You'd hope my due diligence is better on the business side of things, as a 185X mistake there would easily bankrupt the whole company. But the consumer me and the business owner me are, at the end, the same person.
- dirktheman 11y agoI'm the first one to admit that Glacier pricing is neither clear nor competetive regarding retreival fees. I do think that a lot of people use it the wrong way: as a cheap backup. I use: 1. My Time Machine backup (primary backup) 2. BackBlaze (secondary, offsite backup) 3. Amazon Glacier (tertiary, Amazon Ireland region) I only store stuff that I can't afford to miss on Glacier: photos, family videos and some important documents. Glacier isn't my backup, it's the backup of my backup of my backup: it's my end-of-the-world-scenario backup. When my physical harddrive fails AND my backblaze account is compromised for some reason, only then will I need to retrieve files from Glacier. I chose the Ireland region so my most important files aren't even on the same physical contintent. When things get so dire that I need to retrieve stuff from Glacier, I'd be happy to pony up 150 dollars. For the rest of it, the 90 cents a month fee is just a cheap insurance.
- mentando 11y agoI like your commitment to backups. Keep it up!
- danieldk 11y agoI have a similar tiered setup, but didn't add Glacier or Nearline yet. 1. Synchronization across multiple machines using Bittorrent Sync. 30-day archive on local machines and one remote (encrypted-only). One local machine has the archive set to non-deleting. 2. Time machine backups of my primary machines. 3. Encrypted backups through Arq to OneDrive. I'll probably soon add encrypted Arq backups to B2 or Nearline. OneDrive is practically unbeatable price-wise. I work at a university, so have the academic discount. That's 65.97 Euro for four years, or ~1.37 per month for 1TB of storage space (though I only use a fraction of that).
- deleted 11y ago[deleted]
- tehbeard 11y agoWith the time machine backup and his backblaze gone as well? There is no perfect backup solution, and I'd be surprised if this one failed in my lifetime.
- jedisct1 11y agoI don't get Glacier. It's painfully slow, painful to use and insanely expensive. https://hubic.com/en https://hubic.com/en is $5/month for 10 Tb, with unmetered bandwidth. A far better option for backups.
- breakingcups 11y agoI'm interested in Hubic, but where do you read that bandwith is unmetered?
- deleted 11y ago[deleted]
- captain_jamira 11y agoIf one can download a percentage for free each month - 5% in this case, and the price of storage is dirt-cheap, then couldn't one just dump empty blocks in until the amount desired for retrieval falls under the 5% limit? In this case, if one wants to retrieve 63.3 GB, uploading 1202.7 GB more for a total of 1266 GB, 63.3 GB of which represents just under 5%. There's no cost for data transfer in and the monthly cost at $0.007/GB would be just $8.87. And that's just for the one month because everything wanted would be coming out the same month. Has anyone tried this or know of a gotcha that would exclude this? And I realize that for the OP's situation, it wouldn't have mattered since he thought he was going to get charged a fraction of this.
- wallacoloo 11y agoI believe that objects cannot be transitioned into the Glacier Storage class until 60 days after their upload date. So at the very least, you will have 60 days of latency (and thus ~$20 of monthly fees) before you can extract your other data. Additionally, I wouldn't be surprised if the 5% is also based on a storage measurement that is pro-rated for the month. So I would let the 1200 GB of data sit in Glacier Storage for another month before extracting anything, just to be (more) safe.
- rakoo 11y agoThat's possibly a way to trick the system, but storing 63.3 GB in S3 would cost OP less than 2$ for standard availability and less than 1$ for reduced availability, not counting request costs (which are not as surprising as the hidden costs in question here). At this scale you should just store it in S3 and be done with it.
- captain_jamira 11y agosure, I wasn't intending to suggest this as a good premeditated maneuver, but there are probably other individuals out there who have found themselves in a similar position to the OP and considering their predicament. If you've gotten into Glacier for the wrong reason, you may already be in the trap, and you can quickly rip yourself free and take a bunch of skin, spend almost 2 years ever so gently prying yourself free, or maybe a third way. That's my angle here. Also, traps don't have to be laid for someone to feel like he's in one, so I'm not putting that on AWS. The cheapest way out seems to be to just grab 5%/month over 20 months, but that's a lot of sustained effort and contact with the service. So I could see a trick like this as a potential middle ground, at three months and ~$30 according to previous comment's details.
- jaimebuelta 11y agoYou will ALWAYS pay more that you expect when you use AWS (and probably other cloud services). This case is quite extreme, but the way costs are assigned, is quite complicated not to miss something at some point...
- prohor 11y agoFor cheap storage there is also Oracle Archive Storage with 0.1c/GB ($0.001/GB). They have horrible cloud management system though. https://cloud.oracle.com/en_US/storage?tabID=1406491833493 https://cloud.oracle.com/en_US/storage?tabID=1406491833493
- alkonaut 11y agoCurious: if you use a "general storage provider" (like glacier) for backup, rather than a "pure backup provider" (like Backblaze, CrashPlan) why is that?
- natch 11y agoI control the encryption algorithms, the compression scheme, the file chunking strategy, the encryption keys, the encryption of archive/file names, the file naming scheme, the addition of Par2 files, everything. And I don't pay the overhead of an add-on service. Also I back up stuff that's not on my hard drive (only on external USB drives) and I'm not sure how the services handle that. If the services give me some of these points, that's not sufficient; they would have to give me all of these points. Only then would I consider them. All things being equal I'd be willing to pay for some convenience but my current solution is all scripted so it's pretty darn convenient.
- alkonaut 11y ago> I control the encryption algorithms, the compression scheme, the file chunking strategy, the encryption keys, the encryption of archive/file names, the file naming scheme, the addition of Par2 files, everything. I can see why detailed control would be one reason, but you could still just have a very controlled backup to your own storage location(s) as a first step and just let a backup service bulk store your already named and encrypted files? It's only the last-resort you need to go to so if it's a huge blob of encrypted data that shouldn't matter too much -- you only need to access that in case of a total disaster where you lost all your own backup endpoints first. > And I don't pay the overhead of an add-on service. The reason I'm asking is because I was under the impression that backup services are much cheaper than pure storage, while still offering some conveniences such as versioning/backup apps. Glacier charges $0.007 per GB per month, that's $7/month just for a single 1TB machine, just for a single version (If my math is correct, it's early)! If you have dozens of versions it quickly adds up. I do 10 machines at around 1TB on average, unlimited storage in unlimited versions, at $1.25 per machine per month (flat rate, regardless of storage volume). I have tried building my own machines, tried looking at storage providers etc., but can't get near. Even if I did only 1-2 machines, the cost in Glacier would break the backup service cost already at a couple of TB total storage.
- dalanmiller 11y agoSo, what's the most cost effective way to download all your files from Glacier then?
- benmanns 11y agoSpread the download over time. Over 20 months would be free. Over 1 month would be the quoted per-GB charge based on your peak download rate.
- z3t4 11y agoI was looking at Glacier for my backups, but it seemed to complicated ... glad I didn't use it. I ended up using some cheap VPS, two of them located in two different countries. And it's still cheaper then say Dropbox.
- KaiserPro 11y agoGlacier is not a cheap/viable backup its even less suited to disaster recovery (unless you have insurance) Think about it. For a primary backup, you need speed and easy of retrieval. Local media is best suited to that. Unless you have a internet pipe big enough for your dataset (at a very minimum 100meg per terabyte.) 4/8hour time for recovery is pretty poor for small company, so you'll need something quicker for primary backup. Then we get into the realms of disaster recovery. However getting your data out is neither fast nor cheap. at ~$2000 per terabyte for just retrieval, plus the inherent lack of speed, its really not compelling. Previous $work had two tape robots. one was 2.5 pb, the other 7(ish). They cost about $200-400k each. Yes they were reasonably slow at random access, but once you got the tapes you wanted (about 15 minutes for all 24 drives) you could stream data in or out as 2400 megabytes a second. Yes there is the cost of power and cooling, but its fairly cold, and unless you are on full tilt. We had a reciprocal arrangement where we hosted another company's robot in exchange for hosting ours. we then had DWDM fibre to get a 40 gig link between the two server rooms
- NicoJuicy 11y agoDoes anyone have a backup script for backblaze or a similar windows app like SimpleGlacier Uploader?
- joosteto 11y agoIf downloading more than 5% of stored data is so expensive, wouldn't it have been cheaper to upload a file 19 times the size of the stored data (containing /dev/urandom)? After that, downloading just 5% of total data would have been free.
- Herald_MJ 11y agoAt that point, your monthly fee would be greater than that of a regular S3 bucket, which has fewer barriers to retrieval.
- rasz_pl 11y agoReminds me of my advice to netflix after peering problems emerged, just push garbage upstream from every client to equalize your peering traffic.
- yetanotherjosh 11y agoIt still wouldn't have been free. The free download allowance is spread daily across the entire month. That is, you can download 5% of your data per month, and only 0.16% of it per day. For your optimization to work, you'd have to retrieve your data over 30 days for it to be free.
- kozukumi 11y agoFor unique data you want super robust storage options, both local and remote. But for something as generic as ripped CDs? Why bother? Just use an external drive or two if you are super worried about one dying. Even if you lose both drives the data on them isn't impossible to replace.
- limeyx 11y agoSo ... why not just upload and additional 60GB / 0.05 and then download the entire 60GB which is now 5% of the total storage for free ?