3 ms·
> You merely state that it's just not state of the art, and to come work at dropbox. That's about all I can do. However... > This is what I mean by being dis
by jamwt 11y ago
> You merely state that it's just not state of the art, and to come work at dropbox.
That's about all I can do. However...
> This is what I mean by being disingenuous.
I'm don't think disingenuous is what I'm being. Maybe I'm not being as forthcoming as you'd like. Maybe my comment wasn't useful without details I'm not providing. But I'm not lying or misrepresenting anything.
So let me try to be helpful without using proprietary information. Using information people have unearthed in this thread, we can do some math to prove I'm not completely bonkers:
96 unit chassis * 10T SMR = 960T, which is in the > 5x the density in 4U
SMR $/GB is quite good if you work with the right vendor and make the right tradeoffs.
If you could manage that 960T with the same compute resources, you'd save a lot on compute, since that's currently a ~30% of their cost.
With custom enclosures and hardware, you can do more still. You can have deep relationships with vendors and have them customize firmware for your particular use case. You can take their discards that didn't quite meet spec. You can buy crazy cheap stock with high failure rates b/c your distributed system is insanely good at repairs. You get access to stuff that's not being sold yet.
However, let me state again--this would take a lot of work. You need to mold your software around your hardware profile with very small tolerances. You might need to move some I/O stuff in-kernel, and/or move the IP stack out of kernel. You might need to reinforce your DC's raised floors to handle the weight (spoiler: you definitely do unless you planned for these types of machines ahead of time). You need great operations, and really thorough hardware quals b/c your failure domain is huge. Network traffic will go through the roof if a box fails, since you're repairing so much data.
Dozens of people on a handfuls of teams will spend time attacking this problem from different angles.
So, again--I don't necessarily think Backblaze did anything wrong here. Doing all this might be a crazy idea for them and not worth it. My statement was: it's not state of the art in terms of $/GB or density.
But for some shops, they spend so much on storage they can afford invest really, really deeply on optimizing storage cost. That's where the state of the art is.
- deleted 11y ago[deleted]
- nickpsecurity 11y agoInteresting and thanks for the enlightening comments. I proposed similar things to a few people esp the partnership with hardware development, working around faulty (cheap) stuff, custom firmware, and kernel I/O modifications. I expected each of these in a high-end solution. My architecture didn't have an IP stack at all: guards or front-end acted as I.P. gateway while internally using much simpler protocol that still worked on Ethernet. Inspired by Active Messages and FLIR protocols of past. TCP/UDP/IP is unnecessarily complicated with simpler message passing easy in software and FPGA hardware for offloading onto cheap FPGA's. A side note was using one or more I/O coprocessors like mainframe Channel I/O (see Wikipedia) to get huge throughput might be cost-effective depending on hardware selected. Esp if using embedded SOC for I/O offloading. Note: Octeon II's can do many Gbps of line-rate processing for three digits and much simpler internally than a lot of stuff. The SASD project for secure, hard disks also suggested putting logic in the HD or SATA controllers rather than the system. Made me wonder about putting I/O mgmt and most storage logic in custom PCI cards with lots of SATA connectors. Basically, embedded systems enough CPU to manage whatever you throw at them, just running what's good for that layer, and less watts/money than whole servers. Seems like at some point one could port that layer to straight hardware for great performance/watt although high, initial NRE. eASIC can do that relatively cheap, though, with the Nextreme's if a design is FPGA-proven. That was just exploration on paper so I'm not sure how practical it would be. My focus was high security systems and storage which necessitate better isolation. Hence, PCI cards rather than servers. Interesting that there was a lot of overlap, though, between what I thought made sense and what you all ended up doing. Very interesting. Guess I'll keep proven tactics in mind in future designs.
- jamwt 11y agoWow, very cool! So, candidate for offloading that's really commonly considered is "scrubbing". Essentially, protecting against passive bit flips by checking your data against your own checksums. If you have some N TB of data under management, it's really expensive to be striping these large serial reads all the way to an application's userland. So finding ways to keep these as low-priority "background" reads (that always yield to interactive reads) in the scheduling sense that a.) notify the daemon in userland if they fail and b.) without requiring that daemon to be in the data plane, is high-reward. Ideally, the daemon can stay in the control plane so it can manage/report accounting on scrub scheduling and time since last-scrub per disk region, or whatever. You can offload it to a kernel module to avoid `copy_to_user` (if you can't DMA) and/or context switching. Or even offload it to hardware--some custom host adapter, possibly, using custom ATA/SCSI commands to control it and query it. (and the `sg` driver).
- PuffinBlue 11y agoI, for one, believe what you are saying. There have been murmurs for several years that big outfits like Dropbox/Amazon/Google have moved to use some very fancy storage tech. Just to ponder on the available confirmed info - lets add Samsungs 16TB SSD to a 48 disk 1U rack and times that by 4... I'd say just with a lot of cash and stuff you can buy now you can pass that order of magnitude pretty easily. If you make your own...I can easily see the densities knocking Blackblaze's pod out of the water. As for costs, well massive ordering, preferred rates, custom designs and symbiotic relationships with manufacturers, coupled with a software architecture that complements all these hardware gains could conceivably bring costs down to levels unheard of outside of the big cloud service providers. The conditions and technology available right now make what you are saying within the realms of the 'adjacent possible' if you have enough money to throw at the problem. With the kind of clout Dropbox et al have I can well believe you're getting incredible storage densities but at incredible knock down prices.
- scurvy 11y agoSo you are using SMR drives. Wow, good luck with that. All the flash front-end caching in the world won't improve the miserable performance of those drives. I guess that's the only way to justify moving out of AWS? I'd really like to see some usage stats on your SMR clusters because so far the universal verdict is that they're terrible for anything other than WORM.
- olavgg 11y agoSMR drives may have poor write performance, but still good enough). Read performance is good. Anyway I believe the network is the bottleneck anyway, not the drives.
- scurvy 11y agoThat fits the WORM profile that I suggested. During normal operation, the disk IOPS is the bottleneck. The network only becomes the bottleneck during repair/failure states. Even scrubbing falls under disk IOPS and not network. At least that's what I've found in EC clusters. Replicated clusters might be different.
- jamwt 11y agoHost Managed (or Host Aware) SMR disks are fine. You need to do your own zone management using ZAC/ZBC, and throw out your filesystem. The disks all have a conventional section near LBA 0 for metadata management (e.g. indices). Usual tradeoff--more work in software, but far less $.
- codemac 11y agoAh good, now we're talking specifics. Does REPORT ZONES not point to this conventional section always? I cannot wait for a FORMAT ALL ZONES or something command to control the number/amount of conventional sections at the sacrifice of capacity, though it looks like it's not coming. The power guarantees in the zbc spec (as of r03) seem to not be held by some drives. Writing less than a full zone seems to be a bad idea.
- basch 11y agoup to an order of magnitude more 5x the density so half an order of magnitude, which is much more believable?
- jamwt 11y agoNope, that was just for this example; custom enclosures and larger disks can push quite a bit higher than this.
- PhantomGremlin 11y agoan order of magnitude Many years ago I liked to say that, and a friend would call me out. So I would promptly change my statement to "a binary order of magnitude". 5x the density Maybe he should be saying "a quinary order of magnitude"? :)
- DanWaterworth 11y ago> 5x the density > so half an order of magnitude Almost 70% of an order of magnitude.
- deleted 11y ago[deleted]