12 ms·
I worked on the design of Dropbox's exabyte-scale storage system, and from that experience I can say that these numbers are all extremely optimistic, even with
by kmod 7y ago
I worked on the design of Dropbox's exabyte-scale storage system, and from that experience I can say that these numbers are all extremely optimistic, even with their "you can do it cheaper if you only target 95% uptime" caveat. Networking is much more expensive, labor is much more expensive, space is much more expensive, depreciation is faster than they say, etc etc. I don't think the authors have ever done any actual hardware provisioning before.
I didn't read all their math but I expect their final result to be off by a factor of 2-5x. Hard drives are a surprisingly low percentage of the cost of a storage system.
- deleted 7y ago[deleted]
- Taek 7y agoAuthor here. A lot of these numbers are drawn from experience in the mining world, where people realized that when cost is the ultimate bottom line, a lot of corners can be cut. Sia systems don't need a ton of networking. I ran the networking buildout costs by some networking people, and again it comes down to cutting corners. If you only need 10 gbps per rack, if you don't mind having extra milliseconds added, etc, you can get away with very scrappy setups. The whole point is that it's not a highly reliable facility.
- anonymousiam 7y agoThe third sentence of your Medium article says; "Despite this, the Sia network is able to achieve 99.9999% uptime for files." How do you achieve this in a "not a highly reliable facility"?
- prepend 7y agoRedundancy, I assume. Lots of unreliable facilities and nodes can be very durable in aggregate.
- deleted 7y ago[deleted]
- jakewalker 7y agoSounds kinda like the people who thought that bundling a bunch of bad mortgage debt together in slices could get it rated AAA.
- __jal 7y agoReminds me of a saying from the first dot.com crash. (Or at least I heard it then first.) Tying two bricks together doesn't make them float.
- kybernetikos 7y agoIt's a great saying, but aren't all large ships these days made out of components that don't float individually?
- mcdevilkiller 7y agoNo, they are mostly made of air, which floats (Partial sarcasm, as that is what makes them float)
- feross 7y agoYour mind is going to be blown when you learn how TCP/IP works.
- allendoerfer 7y agoOr computers for that matter. Deep down it is not really about 1s and 0s, more like thresholds in between 1 and 0. To me a big part of computer science is abstracting away unreliable details to make them seem reliable.
- deleted 7y ago[deleted]
- arcticbull 7y agoHere's the issue. We know that due to economy of scale and domain experience, AWS will always have the lowest cost (to Amazon) for storage -- whether that's totally-reliable storage, or sorta-reliable. If there was a demand for sorta-reliable, they'd build a sorta-reliable S3 and undercut you. Then, blockchain adds inefficiency. Therefore, it's basically impossible for any blockchain solution to have a lower total cost to provide storage.
- throwaway9878 7y agoEverybody with more money than you can always undercut you in anything you ever do so why bother ever trying to do anything
- nemonemo 7y agoI think the point is that the cost should not be the main motivator. I would agree that there needs to be other differentiators in addition to cost, which could provide sufficient moat against other competitors, big or small.
- jessriedel 7y agoIt's possible that there could be non-obvious innovations in how to save money with a low reliability threshold, which Amazon might not be able to effortlessly copy.
- res0nat0r 7y agoIt should also be noted that you can get S3 storage for $1/TB/month already if you use the Glacier Deep Archive storage class.
- paulryanrogers 7y agoDid that include network IO to retrieve your data?
- kmod 7y agoSure, let's dig into networking. Who pays for rereplication traffic? If you do 64-of-96 RS encoding, that means for every failure you need to transfer 64x the lost storage capacity. If you're targeting a "low individual uptime but high aggregate uptime" model this means you need to be storing data in multiple sites -- and dedicated cross-geo bandwidth is expensive. I agree that in the happy case you can use low-bandwidth cheap equipment, but to get good reliability you need to provision for larger clustered failures such as rack- and row-level outages.
- DuskStar 7y agoI think his point is that he's targeting low aggregate uptime, too.
- Taek 7y agoDefinitely not, aggregate uptime is extremely high. We've never seen downtime do to network outages, only software bugs. And even then, only some users were impacted by the bugs, we've never in 5 years had a broad outage.
- DuskStar 7y agoGotcha, and your reply to the GP clarified a lot for me!
- Taek 7y agoSure! First up, we don't do repairs every time one host goes down. Standard practice on the network is to wait to do a repair until a full 25% of the redundancy is missing (in 64-of-96, that would be 8 hosts offline). Then you repair all 8 at once, significantly reducing the total amount of repair traffic. But secondly, offline doesn't usually mean dead and gone, with unstable datacenters like this they are usually back online before the user has lost a full 25% of their redundancy. Row level and rack level outages are handled by data randomization. The entire Sia system heavily depends on probabilistic techniques, both on the renting and hosting side. Row level failures will take out some of your data, but nobody should be disproportionately impacted by a cluster failure. On Sia, each piece is at a different site. So 64-of-96 implies that each chunk of data (96 pieces to a chunk) is located in 96 different places. This doesn't help with the geo-bandwidth, but as discussed above there are other techniques to handle that. Surprisingly, bandwidth pricing on the Sia network is even cheaper than storage pricing relative to centralized competition. That's a lot harder to model at scale though, so we aren't as confident the Sia bandwidth pricing will hold up at $1 / TB in the long term. And technically, most of this stuff is customizable per-customer. If your particular use case has a different optimal parameterization, it's fairly easy to tune your client to suit your particular needs.
- kllrnohj 7y agoThe buildout in the article doesn't work. You can't plug in that 4-lane SAS SFF-8087 splitter cable into that motherboard. You're only getting 8 hard drives per motherboard with that setup, not 32. That puts the cost of 192 TB at more like $6240, not $4945. Could be less if you find a good deal on mini-SAS PCIE cards, but still going to be substantially higher than $4945.
- notyourday 7y agoYou are going to get a 4U server that takes 60 disks and populate it with SAS cards driving them using SAS-2-SATA cables.
- ADefenestrator 7y agoEven if it had slots for the splitter cable, Intel and AMD onboard SATA explicitly doesn't support port multipliers as far as I know. You can buy PCIE SAS cards that do for relatively cheap, but then you have to find board with enough PCIe slots. Easy enough on the "gamer" boards but if you want ECC (and you probably do, for storage) and IPMI (you probably do, if you have more than a few dozen servers) your options get much more limited. Other than 1 or 2 Asrock Rack boards, you pretty much have to move into Epyc 7000-series or Xeon Silver or above. Often dual-socket on the Xeons to get a board with lots of PCIe. In theory something like an Epyc 3000-series with lots of PCIe or onboard SATA that supports port multipliers would work great, but I don't think anyone actually makes that.
- samstave 7y agoWhy the heck would one want to store data in a non highly reliable anything run by a team who is building around cutting corners??? Wtf am i missing here? —- So, I assume you’re quite young. ~31 or so... I say this because youre not speaking with a level of professionalism one would ostensibly have when talking about storing (and providing access to) another’s data... And if your actual business model is focused on cutting corners to save costs in exchange for reliability - then the obvious is that you seek to cut corners and costs in all other aspects of your business and are telling people “STAY THE FUCK AWAY” from your business... unless your customers are of similar ilk, and if not - then their data isn't even important enough to persist, your systems or theirs...
- justinmeiners 7y agoYou're taking "cutting corners" out of context. He is describing how individual nodes don't need to run at the standards of regular data servers and are hence cheaper, but in aggregate provide a reliable service. > So, I assume you’re quite young. So you do you have a technical criticism?
- toolslive 7y agoOTOH, your object storage can be even cheaper if you don't do Reed-Solomon erasure coding, but use rateless erasure codes. For example, online codes[1] have been used by Amplidata to have more reliable storage with lower overhead. There are downsides however (no partial reads, no mutability, ....) [1] https://en.wikipedia.org/wiki/Online_codes https://en.wikipedia.org/wiki/Online_codes
- vikiomega9 7y ago> exabyte-scale storage system Somewhat of a random question, can you point me to some state of the art research?
- ghostpepper 7y agoI am not the op nor an expert on this, but I think https://ceph.io/ https://ceph.io/ targets clusters of this scale.
- Mave83 7y agoYes, with croit.io you can manage Ceph clusters with ease, reducing Laber costs and increase reliability. We build clusters of low PB scale that have a TCO with everything from labor to hardware, from financing to electric, from routers to cables and can be run below 3€/TB. For that you can store data in Block(rbd, iscsi), Objekt(S3, swift) or Filestorage(CephFS, NFS, SMB), high available on Supermicro hardware in any datacenter worldwide. Feel free to contact us or use our free community edition to start your own cluster.
- sydney6 7y agoStorage Systems of this scale are thesedays almost exclusively built around object based storage rather than "legacy" block or file backed solutions. I guess today, min.io is would be the way to go. (to "go".. little pun on the end:)
- londons_explore 7y agoGoogle's clusters are all block and file based (depending on what layer you want to use them at)
- statictype 7y agoWhat does object-based mean? How is an object different from a file (which I presumed was a collection of blocks)?
- 7y ago
- notyourday 7y agoWe have done this calculation and even if you put your gear into Equinix/Digital Realty in the most expensive places and use Backblaze-type setup ( which is not optimized and buying retail) bringing 10Gbit to every 4U the price for double-writes at 5TB disks are $10/year per TB.
- chx 7y ago> I didn't read all their math but I expect their final result to be off by a factor of 2-5x. Can't be more than 2.5 because Backblaze B2 already gives you $5/TB/Mo.
- deleted 7y ago[deleted]
- throwaway3157 7y ago> Can't be more than 2.5 because Backblaze B2 already gives you $5/TB/Mo. Well it can be, if they have a lot of inefficiencies. Backblaze could have more experienced engineers who overcame these. I assure you, I can accidentally design a very expensive storage system as I’m not that smart ;)
- pepemon 7y agoWho are you?
- TrueDuality 7y agoBackblaze is operating at their own economy of scale with dedicated deals with suppliers, custom bare bones hardware, optimized processes, etc. It's always more expensive starting out as a little company for raw hardware and processes until you're big enough and mature enough to get the deals and processes in place. That is also not the only service that Backblaze offers and wasn't their first. It could be that B2 is simply a way for them to offset their cost for extra capacity and are running it effectively near-cost for them.
- Aeolun 7y agoIf all their other customers on the $5/month unlimited backup plan store less than 1TB, but since it’s really aggressive about backing up everything (no, please don’t back up my steam games) on your computer I think they go over that.
- 7y ago
- z3t4 7y agoStorage can be far cheaper when decentralized. Sending data over the Atlantic is super expensive compared to LAN networking. Almost all content providers peer with ISP's with onsite hardware. But why stop there, put the "racks" in ppl's basements. Data storage is very compact now a days, you can probably fit 100 TB is a shoe-box.
- jan6 7y agothat'd need trust that they won't run off with your drives, their basement won't get flooded, there aren't power outages often, they have good protections against surges and whatnot... which at ANY datacenter, is automatically included ;)
- Dylan16807 7y agoYou need to trust that not too many of them will have that happen at the same time. Which is not that hard to do. And all those protections in a datacenter are far from free.
- winfred 7y ago>I didn't read all their math but I expect their final result to be off by a factor of 2-5x. I looked at their parts list and it's obvious they aren't serious. CPU is missing, memory is missing, SAS to SATA cables, but no SAS controller, no mounting for the system board. Low effort at best.
- prostheticvamp 7y agoIt seems you’re getting heavily downvoted. I think it’s because CPU and memory are not missing.
- winfred 7y agoThat wasn't on that page when I opened it, you think I'd miss that? They were reading the comments here and updated the page after I wrote the comment. The price was $4700 something, now it's $4945. They simply smooth over it with this: "So we will be using a rig cost of $4500 in our spreadsheet." That way their overall math doesn't change. Should get banned for doing things like that. These guys are up to no good.
- late2part 7y agoI agree with you. Not 2-5x but they are rounding down on costs and optimistic on risks.
- FalconSensei 7y ago> Hard drives are a surprisingly low percentage of the cost of a storage system THIS! There's no use of a 2TB storage if you can't upload/download this amount each month
- tlb 7y agoMany businesses have TB of data they upload and keep for several years, looking at it once or twice over that time if ever. For these you could provision monthly bandwidth as low as 1/30th of the total storage.
- jan6 7y agohave you ever heard of backups and archiving? ;) there are lots of scenarios when you'd want to write some amount, but not read it back for a while, if ever...
- Dylan16807 7y agoIt says right there in the article that bandwidth is charged for separately. So you just buy a line of appropriate size, based on actual usage measurements. And each one of these rigs only needs a gigabit connection to upload/download its full capacity each month, which means the network equipment costs can be minimal.