13 ms·
We have always tried to squeeze costs out of storage. That required a ton of tech/ops work. Bandwidth it turns out isn't that expensive - seems to just typicall
by budmang 9y ago
We have always tried to squeeze costs out of storage. That required a ton of tech/ops work. Bandwidth it turns out isn't that expensive - seems to just typically be overpriced. Hopefully lowering the price will enable people to do more with their data.
I'd love to hear what you haven't done in the past due to bandwidth/download pricing that you'd have wanted to...?
Gleb @ Backblaze
- nbarbettini 9y agoNo feedback from me at the moment, other than to say that what y'all are doing with B2 is awesome. It's great to see more competition in this space, and I love the open-source ethos you've cultivated at Backblaze.
- budmang 9y agoThanks for the kudos!
- unixhero 9y agoGreat news! Maybe now I can consider it for my 10tb offline backups.
- tonyedgecombe 9y agoIf only you had a Linux client ...
- nerdponx 9y agoPretty sure any decent B2 CLI client supports Linux. But for regular backup, this bit hard. I ended up switching to Crashplan for not only Linux support but also support for external hard drives. Then again the official Crashplan client is GUI-only and AFAIK closed-source, so you can't run it on a headless server
- dublinben 9y agoAs a not incredibly impressed customer of Crashplan, I switched to rclone/duplicati and Amazon Cloud Drive for these and other reasons. The price is the same, and you can use whatever client you want.
- nerdponx 9y agoAre rclone and duplicati yet more deduplicating backup programs? we now have: - rclone - duplicati - rdiff-backup - duplicity - attic - borg - bup - zbackup - ...? I'd love to see a comparison of these because right now I have zero idea which one(s) I should use for any given purpose.
- dawnerd 9y agoCan't speak for the others but I use duplicati and send the backups to my unlimited google drive. I like that it encrypts and gives you the key vs some app saying they encrypt it.
- cookiecaper 9y agoI don't know all of them, but I know some. * zbackup: Appears to be unmaintained, superseded by attic/borg. * attic: deduplicating backup repository. Doesn't seem actively maintained anymore either. * borg: deduplicating backup repository. I used this recently. It's cool, but the performance is disappointing. It's single-threaded. I got into the GitHub a little bit and they seem committed to their current way of doing things, which makes it hard for an external multithreaded compressor to take over (everything is processed in 2 MB chunks). * rclone: Meant to provide an rsync-like interface to any cloud backend * duplicity: Automatically performs encrypted, compressed delta-enabled backups, but requires some upfront configuration. Been using this for years. * rdiff-backup: Haven't used, but it seems more like an ad-hoc backup utility for "I want to keep an old version of this folder". Duplicity seems like a better choice for disk-scale backups.
- godzillabrennus 9y agoGlen, thanks for posting! I'm a backup customer of Backblaze (have been for years) and I backup five of my families computers. I'd really like it if your online restore feature would let me buy multiple backups on a single hard drive as an option. I've considered buying a hard drive with my backups from your company multiple times but I wouldn't pay the premium you ask without being able to have all five computers on the drive.
- geostyx 9y agoFWIW, you can return the drive for a full refund.
- infogulch 9y agoI feel like gp wants to keep a hdd as a third copy of the backup somewhere safe under their physical control. This is prohibitively expensive to do for multiple computers at once, and gp would prefer a single hdd that has a copy of the backup for multiple computers, which would reduce the cost.
- budmang 9y agoThanks for the suggestion (and for being a customer.) We've contemplated adding functionality to send the Backblaze laptop/desktop restore to your Backblaze B2 account. Then you could put multiple restores in a 'snapshot' and get that delivered on a drive. Useful?
- adventured 9y agoI'd like to permanently, easily offload all image storage & delivery (consumer oriented services) - without being mangled by the obscene bandwidth costs of AWS. I don't want to roll my own solution to get cheap bandwidth, I want a dependable service that hosts files (images in this case) & can serve them up, and do so at very cheap bandwidth prices. B2 references Vintage Aerial as a success case for storing their large image collection. So I checked out their site, pretty cool service. So I right click -> view image, on the large & thumbnail images on their site and they're on AWS of course (archival storage on B2 I assume). Hey Backblaze, AWS's bandwidth margin is your opportunity.
- brianwawok 9y agoIts hard to be highly available, redundant, and cheap. AWS picked the first two. B2 picked the second two. I don't think they would make a great CDN host.
- brianwski 9y agoBrian from Backblaze here. > AWS picked highly available, redundant > Backblaze picked redundant, and cheap Both AWS and Backblaze have had (and will have more) outages. No single vendor is a good choice for a service that absolutely positively must stay up no matter what. We liked this article that shows how to configure Amazon S3, Microsoft Azure, or Backblaze B2 so you are not susceptible to Amazon outages: https://cloudrail.com/avoid-next-aws-s3-outage-failover-scenario/ https://cloudrail.com/avoid-next-aws-s3-outage-failover-scen...
- truetraveller 9y agoHi Gleb, I want to use your cloud storage instead of S3. My main reason for using S3 is the DURABILITY (not reliability). S3 offers 99.999999%, which means it is replicated numerous times. Which means I DON'T HAVE TO WORRY ABOUT taking extra customer backups, it is "baked" into S3. How is the durability for B2?
- jamroom 9y agoFrom the B2 page: https://www.backblaze.com/b2/cloud-storage.html https://www.backblaze.com/b2/cloud-storage.html "Reliability: 99.999999% durability."
- dastbe 9y agothe OP misquoted S3s durability guarantee. Its 11 9s.
- kcorbitt 9y ago...but in a single day datacenter, which realistically drops several "9"s from that guarantee when you consider tail events. I like b2 and am using it for backup in some projects, but unfortunately it's still not an appropriate choice as a primary data store. The common counterargument I've heard is that you should have multiple redundant storage providers that you manage at the application layer anyway. Problem with that is, there are a near-infinite number of things I "should" do, but very finite resources to do them. S3 backs up across multiple availability zones by default, so using it means one less thing I have to spend brain cycles worrying about.
- cookiecaper 9y agoHi Gleb! The thing I haven't done due to the download pricing is back up very much data into B2. Since the B2 API doesn't support renames, you have to download and then re-upload a file to rename it. I would be storing a lot more there if renames were available! [Note: they may be available now; I haven't checked for a few months. Please tell me they are.] I was an early adopter of B2 and uploaded a bunch of files with naming that's incompatible with the "virtual folders" of the web UI (wrote my own script to handle this since Duplicity et al did not yet have support). Now I'm stuck and the only solution offered is just "pay us to download everything again". Also, the web interface does not have any pagination. Since virtual folders are based on naming, and since rename support is not available, this can flood the browser. I have a bucket that crashes the tab every time I try to open because it's not paged. I could write a little Greasemonkey injector to page it myself, but meh, I don't really care enough to do that. Instead, I just don't really use B2. When I discussed these issues with support and suggested that at least pagination on the web interface was a basic thing that engineering should be able to accomplish quickly, they did not seem to agree.
- budmang 9y agoInteresting. "rename" is generally not supported by object storage, but an API to "copy with a new name & delete" often is, which allows the same result. I presume that would work? (We're planning to add that in the future.) Pagination is on the list as well...
- cookiecaper 9y agoBoth of these things seem like such weird things to avoid adding, especially pagination. I can't be the only user to have a bad experience with the web UI because of this. How hard is it to add a simple pager? It doesn't have to be fancy.
- budmang 9y agoNeither is hard. Want to do both. Just have a list of hundreds of requests and have to prioritize ;-)
- Veratyr 9y agoI'd be interested to know, have you looked into reducing/eliminating bandwidth prices further for clients close to your network? It seems like you're with Unwired, which has heavy inbound traffic and peers at a number of facilities in the Bay Area. It'd be really interesting to know if you could, for example, offer free traffic to Linode, DigitalOcean or Vultr, all of which have offerings in the area and could help balance out Unwired's traffic flow (I think that's a thing ISPs care about). The big thing missing from B2 for me now, relative to the other big players, is the ability to process my data once it's stored with you. If I store my images in S3 for example, I can use EC2 to thumbnail them without paying anything for retrieval. At the moment I can't do that with B2 and it seems unlikely to make sense for you to build your own compute cloud but perhaps you can integrate with others.
- cookiecaper 9y agoB2 is a competitor of long-term storage solutions like Glacier or TarSnap, not virtualized storage like S3. It's meant to be almost 100% write-only. Basically, you pull your backups down in case of disaster.
- budmang 9y agoBackblaze has a backup service for laptops/desktops, but our B2 service is actually more like S3 - the data is available without delays or penalties for access. (Unlike Glacier.)
- cookiecaper 9y agoYeah, it's similar to Google Nearline or TarSnap in that unlike Glacier, the data can be read out instantly instead of waiting for Amazon to restore it off of a tape or whatever they do. The "penalty for access" is the download cost. However, my understanding is that B2 was intended for long-term storage and not everyday cloud storage. Is this understanding incorrect? Backblaze does take the position that B2 should be considered as a competitor for S3? Up to this point, my experience with B2 has been that they don't want to be the authoritative storage source for the data. I seem to recall reading that B2 should not be used as primary storage (though that was in beta days). In support discussions, a representative expressed surprise when it was implied that I may not have local copies of files that already exist on B2 (I do indeed have local copies, but it was beside the point) before confirming that yes, I would need to redownload and reupload to complete a rename.
- sfaruque 9y agoHi Gleb, I'm about to sign up for the B2 Cloud Storage, just noticed a tiny typo in the "Keep costs down" section on https://www.backblaze.com/b2/cloud-storage.html https://www.backblaze.com/b2/cloud-storage.html: "...the fist 1 GB of downloads per day are free."
- aschampion 9y agoJust read it in a kiwi accent and all is fine.
- budmang 9y agoThanks. Will get that fixed today.
- atYevP 9y agoYev from Backblaze here -> We got it fist! ;-)
- truetraveller 9y agoAn OT question about the list() function, which also applies to S3 and Google. Why does it return such a verbose, heavy response. Here's an example: https://www.backblaze.com/b2/docs/b2_list_file_names.html https://www.backblaze.com/b2/docs/b2_list_file_names.html. If I have a 10000 files, this will be a LOT of bandwidth. I JUST want the list of file names, NOT the other redundant fat. Solution: 1) let me specify which fields I want OR 2) compress the response
- brianwski 9y agoBrian from Backblaze here. Backblaze does not charge anything (zero) for the bandwidth of the b2_list_file_names response. What we mean for the bandwidth is "file contents you download". So compressing or not compressing b2_list_file_names won't make any difference in your bill. We do change a TINY "per transaction" charge on api calls like b2_list_file_names. The reason is to break even on the farm of servers (we call them "API Servers" internally) that are needed to support the HTTPS load generated from those calls. But that is really really tiny, like you get 1,000 calls for 4/10ths of one penny.
- truetraveller 9y agoThanks Brian. I like B2. And thanks for "absorbing" the cost of the list() function. My sincere question to you is: Sending this much redundant data in the response can cost YOU a lot, it is not "trivial" when you add millions of requests. How hard is it to allow the user to specify which fields he wants in the response? Also, please answer my other question about the authorize() token headache.
- brianwski 9y agoCatching up on answers: > How hard is it to allow the user to specify which fields he wants in the response? Pretty easy, but like all things we prioritize features based on customer requests. Seriously, you could probably game it by creating 20 different gmail addresses and emailing us once every three days from a new email address requesting the same feature. :-) If we hear from something like 20 different customers the same request we prioritize it higher and get it done. The OTHER way is if a "potentially large" customer approaches us and says they require some feature or they cannot integrate or won't use B2. I won't name the customer for privacy reasons, but we added the ability to put the SHA1 at the END of the upload instead of at the start in the HTTP headers because one "important" customer required it. It also makes sense, the headers come FIRST and some streaming applications don't have the SHA1 until all the data streams through them, thus at the end of the stream. > redundant data in the response can cost YOU a lot Much to my surprise (totally being honest), bandwidth only costs Backblaze about 2% or 3% of our total operating budget (including salaries and datacenter). Modern bandwidth is CHEAP! But even more interesting is that most of our Online Backup customers and B2 customers pretty much only UPLOAD data, but we are forced to purchase symmetric bandwidth (same upload as download). So hilariously enough, outbound bandwidth is utterly free to us until it comes up to match the inbound bandwidth. Free. I've even tried to come up with a scheme where we make it free to our customers up until the moment we actually have to increase our capacity to support outbound, but I can't figure out a scheme to make that happen. So we are pricing it for the situation when we exceed the "shadow", even though it is nowhere near (and probably never will) exceed the shadow. > please answer my other question about the authorize() token headache I looked back and found one question you asked about 24 hours vs weeks vs forever. If that is the question, the 24 hours is a "guideline" because so many customers kept asking what to expect so we made up a number. In reality, you have to follow the protocol which is: if a call is ever rejected for whatever reason with a "expired_auth_token" or "bad_auth_token" then your code should re-authorize. It's source code that you have to write and it doesn't matter how often it is triggered, after you write it (correctly) then it will always occur when it needs to occur, and never occur when it doesn't need to occur. To give you an insight as to why, if we happen to reboot a particular server (or it crashes, or the power supply dies), then all the auth codes will be incorrect and all currently connected customers that used that authentication server need to re-authorize which will come from one of the OTHER authentication servers. We do weekly code updates on Thursdays at 2pm-ish so if you get an authentication token on Friday it will probably work for 6 days, but if you get one Thursday morning it might fail once in the middle of the day. This is all about cost reduction for Backblaze - we simply don't purchase ANY high end load balancers, it is all cheap computers that are expected to fail and break, but because there are hundreds (or thousands of them), even if half of them have lost power or died there are still plenty up and running (now we have some in a different datacenter!) to respond to your software's requests. We don't own a single load balancer (like F5), and we pass the savings on to you. Unfortunately we also pass along a small amount of additional complexity. So it is really straight-forward code and you write it once and it just executes every time it needs to and never executes when it doesn't need to. Your code should not have ANY assumptions of time in it. Unfortunately it isn't something where you can write really bad software and just ignore it, you MUST handle the return value EVERY time it is returned, whenever one of our authentication servers has a hardware failure (which is at unknown times) or whenever we restart Apache Tomcat on the authentication servers (Thursdays at 2:15pm-ish-kinda).
- natch 9y agoI care less about the download pricing for backups, and more about the storage pricing. I wish you had a comparison for your storage versus Glacier storage, because all of the claims that it is "cheaper than S3" just seem non-credible when I factor Glacier into the picture. It is true that Glacier is a different service from classic S3, but you can't just ignore it. Or at least some of us can't. It does exist. I wonder how it compares. If you don't have a chart, can you tell us here? I understand download pricing can get complicated. But again I'm not asking about that. I realize download pricing also is important, but right now I'm just asking about storage pricing. Not sure if you are answering questions here but if yes, please share what you know.
- budmang 9y agoSure. A bit hard to do as a table here, but here it goes: B2 Glacier Storage: $0.005/GB/mo === $0.004-$0.005/GB/mo (by region) Upload: Free === Free + $0.05/1000 transactions Download: $0.02/GB === $0.05-$0.09/GB Retrieval Fee: Free === $0.003-$0.03/GB + $0.003-$0.03/1000 requests Retrieval Time: Milliseconds === 1 min - 12 hours In short, unless you only upload very large individual files, do nothing with them, and don't access any of them for at least 10 years...B2 is lower cost. (And, of course, doesn't required you to wait to get them out.)
- natch 9y agoNice, thanks! Edit: BTW it's really a breath of fresh air to see someone whose product is covered hanging around for a while answering questions. So many people in your position either don't show up on their HN stories at all, or do so only for an hour or so and then leave a ton of questions unanswered. Probably my number one wish list feature for hacker news would be if they could somehow coax people at companies that get linked to return more, including beyond the first hour or first day. I hope you check back on this story later and answer more questions that come in from different time zones.
- budmang 9y ago
- finnh 9y agoNot quite what you are asking, but: what colo facilities have max bandwidth to you, and what is my likely MB/s at the top-performing location? S3 can deliver 500 MB/s to a single EC2 node pretty consistently (same region, multiple streams) which keeps me there.
- brianwski 9y agoBrian from Backblaze here. > what colo facilities have max bandwidth to you Our primary colo is SunGard in Rancho Cordova, California. I believe we are currently provisioned for about 220 Gbits/sec (in this one facility alone). I'm not completely sure what this would top out at if we asked for more, but it is a colocation facility with a bunch of network providers such as Cogent, GTT, Century Link, etc. I kind of expect they can provide anything we ask for? > what is my likely MB/s at the top-performing location I have three answers, you can choose one. :-) 1) If you want to download a file once after a very long time of not accessing it (meaning it is not in the "B2 cache"), and with one thread, it must be fetched from the vaults to the "caching layer". Spooling a single file off of the vaults with one thread is relatively slow at about 15 Mbits/sec. 2) If the file is already in the caching layer, then it is fetched from local SSDs already inside the caching servers. Each individual caching server has a 10 Gbit/sec network card so if you only have one thread then you are absolutely limited to 10 Gbit/sec (and realistically the caching servers are not completely idle so a realistic number might be only 3 or 4 Gbit/sec). 3) If you want to fetch many files with many threads, fetch one file at a time with each thread. I'm not sure how many HTTPS threads an EC2 node can handle, maybe 200 safely? If it isn't in the cache then you get 200 * 15 Mbits/sec = 3,000 Mbits/sec = 375 MBytes/sec. If it is coming out the cache even faster? Hope that helps!
- finnh 9y agoIt does - thanks!
- truetraveller 9y agoAnother IMPORTANT aside. Very annoying. Backblaze requires an authorize() call to obtain a token before calling any API function. Please see https://www.backblaze.com/b2/docs/b2_authorize_account.html https://www.backblaze.com/b2/docs/b2_authorize_account.html. This is unlike S3, which simply needs a STATIC API key, and boom, you can access the API. I don't know why you did this, but it is bad developer UX. Is it because tokens are "cool". I see no need to have this concept of tokens for your API. Why I am annoyed: 1) An API operation now requires one ADDITIONAL call. B2 already requires 2 calls to upload a file, and this now means 3 calls! 2) Related to above, this means bad latency. The roundtrip can be significant, especially since you only have 1 data center in California. My users will NOT be happy. 3) My code is much more complex now. 4) It is more work FOR YOU. This means more processing. More resources wasted FOR YOU. More cost. Why would you do this? You are supposed to be lean. Imagine millions of extra calls EVERY day. Adds up. Is there a solution to this? Can the token be cached? If so, for how long. Also, would love to see a reasonable rationale for this? If there is no rationale, then please provide a STATIC approach to access your API. I REALLY like B2, and I wish they could think of developer UX a little more.
- d4l3k 9y agoIf you look at their API you'll see that the token can be cached for up to a week. I doubt that having a check if the token is up to date will add much complexity.
- truetraveller 9y agoIt really does add complexity: -I have to store the cached token somewhere: -I have to do date arithmetic, what if my server clock is incorrect, -if an API call failed, I have to check for an additional error: a invalid token I don't mind complexity when it is needed, but I do not prefer complexity for no good reason. That being said, can it really be cached up to a week? I see "at most 24 hours", which means it could be even less than 24 hours, so there's no guarantee. Please see: https://www.backblaze.com/b2/docs/b2_authorize_account.html https://www.backblaze.com/b2/docs/b2_authorize_account.html
- monort 9y agoI'd like to check checksums for my data, but it's still too expensive to do regularly. Can you provide an api for doing it on your side?
- brianwski 9y agoBrian from Backblaze here. If you use b2_list_file_versions (or any of the b2_list functions) there are checksums included in that list (SHA1). When you ask for that file list, the results are actually coming out of a large Cassandra cluster (database with many computers and fast SSDs). Ok, so even if you do not request a file list of any kind, Backblaze itself automatically is always walking over the entire data farm checking every single last file's SHA1, plus checking if the files match the database's SHA1. A full sweep completes in less than a month then begins again. If any bit error is detected (which happens semi-regularly since we have over 72,000 hard drives and bits get flipped) then we rebuild the files from the other copies (Reed-Solomon encoded across many separate computers in many separate locations in our datacenter). We love and take care of your data so you don't have to.
- monort 9y agoThanks for the information! Though I'd like to be able to initiate the sweeping process for my files more often, maybe you can charge for such requests?
- jclimac 9y ago@Gleb: I'm from Brazil and if you don't know already bandwidth here is priced like gold, so I just can't ignore the price you are offering. The service I provide has, currently, an average of 60k http requests daily, approximately 10Gb. Those requests are images stored on Google Cloud Storage. I'm going to offer a new service for photographer to create online galleries, by some calculations I had made the average daily requests will increase around 10x, growing up to 600K daily. My main questions about B2 are: 1) Is there any use case you know of someone uploading files to B2 directly from the browser? I know this is possible, but as per I know I would have to open tokens and url uploads on JS code, which, obviously, is not cool (and I won't do that). I can do this on S3 and GCS, their API offers the possibility to create a key on the server and use it on the webpage, but I didn't see anything like that on B2. 2) About http requests, is it ok to have this amount of requests (600k) in a daily basis? I mean I know B2 is excellent for backups but I'm not sure if it will be good serving websites that have a big number of images. Thanks!