36 ms·
You can do a lot more than 60, and you can get a lot bigger than 6T. > It's disingenuous to call this not "state of the art" when it's really quite close, and
by jamwt 11y ago
You can do a lot more than 60, and you can get a lot bigger than 6T.
> It's disingenuous to call this not "state of the art" when it's really quite close, and obviously is meeting a backblaze design goal.
So, I never said it didn't meet their design goal; it looks like quite a nice system, and it's better than their previous gen. Looks like they're doing things right, looks like solid engineering.
But I stand by the statement that, in the absolute sense, this is not quite the current benchmark on density or $/GB in the industry. That was the only point I was trying to make to the HN audience in general, and it's not intended as a critique of Backblaze specifically.
But there are very different amounts of resources (I'm assuming) invested in achieving those benchmarks than Backblaze can likely commit.
Ergo, Backblaze probably completely made the right decision for Backblaze. Both of these things are possible.
- codemac 11y ago> you can get a lot bigger than 6T Come on - most companies reasonably can't. There's a WD 10T smr/helium drive, and seagate is selling that crap 8T video drive. Unless you're talking ssd 3d/TLC density (more $), or the vapor-ware 20T drives that some storage vendors hint at and don't let me file a PO for, I literally don't know what you're referring to (now I'm getting where I can't name names) The reason I'm complaining is because of you don't state: - what the current state of the art is that they're not meeting - what they could do better with their design goals given that information. You merely state that it's just not state of the art, and to come work at dropbox. This is what I mean by being disingenuous.
- jamwt 11y ago> You merely state that it's just not state of the art, and to come work at dropbox. That's about all I can do. However... > This is what I mean by being disingenuous. I'm don't think disingenuous is what I'm being. Maybe I'm not being as forthcoming as you'd like. Maybe my comment wasn't useful without details I'm not providing. But I'm not lying or misrepresenting anything. So let me try to be helpful without using proprietary information. Using information people have unearthed in this thread, we can do some math to prove I'm not completely bonkers: 96 unit chassis * 10T SMR = 960T, which is in the > 5x the density in 4U SMR $/GB is quite good if you work with the right vendor and make the right tradeoffs. If you could manage that 960T with the same compute resources, you'd save a lot on compute, since that's currently a ~30% of their cost. With custom enclosures and hardware, you can do more still. You can have deep relationships with vendors and have them customize firmware for your particular use case. You can take their discards that didn't quite meet spec. You can buy crazy cheap stock with high failure rates b/c your distributed system is insanely good at repairs. You get access to stuff that's not being sold yet. However, let me state again--this would take a lot of work. You need to mold your software around your hardware profile with very small tolerances. You might need to move some I/O stuff in-kernel, and/or move the IP stack out of kernel. You might need to reinforce your DC's raised floors to handle the weight (spoiler: you definitely do unless you planned for these types of machines ahead of time). You need great operations, and really thorough hardware quals b/c your failure domain is huge. Network traffic will go through the roof if a box fails, since you're repairing so much data. Dozens of people on a handfuls of teams will spend time attacking this problem from different angles. So, again--I don't necessarily think Backblaze did anything wrong here. Doing all this might be a crazy idea for them and not worth it. My statement was: it's not state of the art in terms of $/GB or density. But for some shops, they spend so much on storage they can afford invest really, really deeply on optimizing storage cost. That's where the state of the art is.
- deleted 11y ago[deleted]
- nickpsecurity 11y agoInteresting and thanks for the enlightening comments. I proposed similar things to a few people esp the partnership with hardware development, working around faulty (cheap) stuff, custom firmware, and kernel I/O modifications. I expected each of these in a high-end solution. My architecture didn't have an IP stack at all: guards or front-end acted as I.P. gateway while internally using much simpler protocol that still worked on Ethernet. Inspired by Active Messages and FLIR protocols of past. TCP/UDP/IP is unnecessarily complicated with simpler message passing easy in software and FPGA hardware for offloading onto cheap FPGA's. A side note was using one or more I/O coprocessors like mainframe Channel I/O (see Wikipedia) to get huge throughput might be cost-effective depending on hardware selected. Esp if using embedded SOC for I/O offloading. Note: Octeon II's can do many Gbps of line-rate processing for three digits and much simpler internally than a lot of stuff. The SASD project for secure, hard disks also suggested putting logic in the HD or SATA controllers rather than the system. Made me wonder about putting I/O mgmt and most storage logic in custom PCI cards with lots of SATA connectors. Basically, embedded systems enough CPU to manage whatever you throw at them, just running what's good for that layer, and less watts/money than whole servers. Seems like at some point one could port that layer to straight hardware for great performance/watt although high, initial NRE. eASIC can do that relatively cheap, though, with the Nextreme's if a design is FPGA-proven. That was just exploration on paper so I'm not sure how practical it would be. My focus was high security systems and storage which necessitate better isolation. Hence, PCI cards rather than servers. Interesting that there was a lot of overlap, though, between what I thought made sense and what you all ended up doing. Very interesting. Guess I'll keep proven tactics in mind in future designs.
- jamwt 11y agoWow, very cool! So, candidate for offloading that's really commonly considered is "scrubbing". Essentially, protecting against passive bit flips by checking your data against your own checksums. If you have some N TB of data under management, it's really expensive to be striping these large serial reads all the way to an application's userland. So finding ways to keep these as low-priority "background" reads (that always yield to interactive reads) in the scheduling sense that a.) notify the daemon in userland if they fail and b.) without requiring that daemon to be in the data plane, is high-reward. Ideally, the daemon can stay in the control plane so it can manage/report accounting on scrub scheduling and time since last-scrub per disk region, or whatever. You can offload it to a kernel module to avoid `copy_to_user` (if you can't DMA) and/or context switching. Or even offload it to hardware--some custom host adapter, possibly, using custom ATA/SCSI commands to control it and query it. (and the `sg` driver).
- garblegarble 11y ago>seagate is selling that crap 8T video drive Slightly off-topic but could you go into some detail on what you think the drawbacks are for the Seagate 8TB SMR drive? I've been thinking about getting a few for large files while keeping them online-accessible, I've found it difficult to get the thoughts of grizzled IT folks on the matter :-)
- beagle3 11y agoI would like that info as well. I couldn't find any reliable data, but I did find a lot of reviews that said the 8TB drives tend to die within a few months; I'm holding on to the 6TBs for now.