4 ms·
It's always great to see storage hardware being opened up, thanks Backblaze. I have a question after reading the Vault overview here: https://www.backblaze.co
by nocarrier 10y ago
It's always great to see storage hardware being opened up, thanks Backblaze.
I have a question after reading the Vault overview here:
https://www.backblaze.com/blog/vault-cloud-storage-architecture/ https://www.backblaze.com/blog/vault-cloud-storage-architect...
I'm curious what the bandwidth demand is when an entire host fails and you have to drop a new replacement host in, since you would have 60 Tomes in the Vault that need to be rebuilt at once. You have at least two other parity copies when a host dies (assuming no other drives in the affected host's Tomes are dead on other hosts in the Vault set), so I'm guessing you can afford to wait the handful of days it would take to rebuild the host. I'm still curious to know what rebuild speed you plan for, since I'm guessing you'd be looking at 40Gbps NICs eventually. I was surprised to see only 2x10Gbps.
- scurvy 10y agoFYI, 25gbps ethernet is gaining in popularity much quicker than 40gbps ethernet. Mostly from cloud/scale-out stuff.
- pavs 10y agoI always get confused with these, but are sfp/sfp+ ports also called ethernet ports? For whatever reason I associate ethernet with cat5/6 cable ports.
- dsp1234 10y agosfp/sfp+/qsfp are just the connector types, and can terminate cables that are made of copper or fiber. This is a separate issue from what protocol (ethernet/infiniband/etc) is running over that cable. For an example, take a look at Mellanox's cable page which lists a large number of combinations of connectors and cable types. So in general, to call a random sfp cable/port an ethernet port would be confusing unless both parties already knew it was passing ethernet traffic. http://www.mellanox.com/page/cables http://www.mellanox.com/page/cables
- pavs 10y agoThanks for the explanation.
- ecnahc515 10y agoI'm mostly mentioning this because it's something I just recently re-learned about. The actual connector and cable are mostly independent of each other. The most common connector for an ethernet cable is RJ-45 (what you're referring to as "ethernet ports"). But it's just as possible to have SFP, coaxial, or USB even.
- nocarrier 10y agoHmm, that's interesting. We never did 25 and went to 40 and 100. Well, 10 for normal hosts, 40 for top of rack switches and heavy i/o hosts, and 100 and beyond for future TOR and aggregate switches.
- scurvy 10y agoIt's mainly cost savings due to the simplicity of 25gbps ethernet vs 40gbps. This article has a good run-down: http://searchnetworking.techtarget.com/feature/25-Gigabit-Ethernet-Why-that-why-now-and-whats-next http://searchnetworking.techtarget.com/feature/25-Gigabit-Et...
- brianwski 10y agoBrian from Backblaze here. If we do a full pod swap for a blank, we are disk / CPU limited in the rebuild, not 10 Gbit/sec limited. We have the rebuilds up to taking about 3/4ths of the 10 Gbit/sec link but we continue to try to change the software and buffering to get this higher. It's MUCH better to lose only one drive, because we distribute the rebuild CPU task across all 20 CPUs in the tome and we can get it synced back up in a few hours. All the chattering back and forth still doesn't come anywhere close to maxing out the 10 Gbit. For uploads, each individual pod has a 10 Gbit/sec connection, so ACTUALLY the vault has an aggregate of 200 Gbit/sec facing towards the internet. However, every byte that a pod accepts it has to ALSO retransmit that internally so that cuts the actual theoretical upload rate to 100 Gbit/sec for a properly parallelized application.
- vessenes 10y agoThanks for the blogging and open designs, it's inspiring! Spelling error: check for 'utlizing' -> 'utilizing'.
- brianwski 10y ago> 'utlizing' -> 'utilizing' Dang it! We will fix, thanks.
- nocarrier 10y agoThanks for the reply Brian. I'm a big fan of sharding across hosts and doing your own RS. I'm mildly surprised that you're disk and CPU limited, but there's lots of factors that can come into play with that and it sounds like you are happy with the tradeoffs which is all that really matters. I'm a big fan of what Backblaze has done. If you ever want to talk about replication strategies or similar stuff, I've built large scale storage systems in the past--please feel free to reach out to the email in my profile if you'd like to talk shop and trade notes.
- scurvy 10y ago> If we do a full pod swap for a blank, we are disk / CPU limited in the rebuild, not 10 Gbit/sec limited Wow. This is pretty shocking TBH. My basic ceph nodes will crush a 40gbps link doing a rebuild/rebalance. What's your bottleneck?