3 ms·
Wow, very cool! So, candidate for offloading that's really commonly considered is "scrubbing". Essentially, protecting against passive bit flips by checking yo
by jamwt 11y ago
Wow, very cool!
So, candidate for offloading that's really commonly considered is "scrubbing". Essentially, protecting against passive bit flips by checking your data against your own checksums.
If you have some N TB of data under management, it's really expensive to be striping these large serial reads all the way to an application's userland. So finding ways to keep these as low-priority "background" reads (that always yield to interactive reads) in the scheduling sense that a.) notify the daemon in userland if they fail and b.) without requiring that daemon to be in the data plane, is high-reward. Ideally, the daemon can stay in the control plane so it can manage/report accounting on scrub scheduling and time since last-scrub per disk region, or whatever.
You can offload it to a kernel module to avoid `copy_to_user` (if you can't DMA) and/or context switching. Or even offload it to hardware--some custom host adapter, possibly, using custom ATA/SCSI commands to control it and query it. (and the `sg` driver).
- nickpsecurity 11y ago"So, candidate for offloading that's really commonly considered is "scrubbing". Essentially, protecting against passive bit flips by checking your data against your own checksums." "Or even offload it to hardware--some custom host adapter, possibly, using custom ATA/SCSI commands to control it and query it. (and the `sg` driver)." Now you're thinking. Chips like Octeon already do hardware acceleration of checksums and possibly other I/O boosters. So did mainframe I/O. There's many for FPGA's and ASIC's. Cray's ChipKill technology, implemented in Oracle SPARC CPU's, does something similar for RAM in combination with ECC RAM. So, an ASIC with several SATA connectors and logic for this should be able to do it. Might handle, partially or totally, the protocol for processing the data, too. Wanted to see if I was really thinking ahead by Googling for academic papers. If it's a good idea, there's usually a few people that have done it for years already. Turned up some hits, including on your Reed-Solomon. FPGA-accelerated decision-tolerant coding for reliable distributed storage https://dl.acm.org/citation.cfm?id=1763276 https://dl.acm.org/citation.cfm?id=1763276 FPGA-accelerated, flash store for analytics http://people.csail.mit.edu/wjun/papers/fpga2014-wjun.pptx http://people.csail.mit.edu/wjun/papers/fpga2014-wjun.pptx DB acceleration http://www.cse.buffalo.edu/~vipin/papers/2010/todd1.pdf http://www.cse.buffalo.edu/~vipin/papers/2010/todd1.pdf Filesystem in GPU. Shows GPU's or DSP's can be used for this. Probably Adapteva's chips, too, esp since they might negotiate for volume deals. https://people.csail.mit.edu/idish/ftp/silberstein13asplos-gpufs-slides.pdf https://people.csail.mit.edu/idish/ftp/silberstein13asplos-g... Tilera's were used in a 100Gbps NIDS that I recall. I expect your protocol and algorithms to be simpler. ;) http://www.tilera.com/products/ http://www.tilera.com/products/ So, goes to show either custom, FPGA, or better-suited chips can handle that part at great performance per watt. Cost varies with implementation strategy. Nonetheless, simple functions with streaming data model greatly benefit from this strategy.
- sofaofthedamned 11y agoIf you've got custom drives with custom firmware, why does the scrub need to leave the drive? Couldn't they have a command that kicks off the scrub in-drive, with only a list of failed blocks returned?
- nickpsecurity 11y agoAre we talking about firmware as drivers executing on host or the microcontrollers on the drives? My proposal works on adapter in between the two or could replace HD microcontroller if you had willing vendor. Doing it on host sucks host CPU requiring more powerful and energy consuming processors given No of drives and data amount. Idea is to move straight-forward, possibly parallelizable algorithms off general-purpose, legacy CPU's onto cheaper, efficient ones more suited to algorithm. It wasn't uncommon for FPGA's to get an eightfold boost in performance on stuff like this in various experiments. In my field, high assurance INFOSEC, the main advantage is the use in inline, media encryptors. I really just knocked off specs and ideas of NSA's Type 1 IME given it was the best: http://m.nsa.gov/ia/programs/inline_media_encryptor/ http://m.nsa.gov/ia/programs/inline_media_encryptor/ Also, Orange Book required full mediation and even SCOMP system mediated hardware I/O. Copying these lessons from the past made my design immune to DMA, firmware, and SMM attacks that were later discovered. Like above, offloading gets performance boost cuz inline I/O handles interrupt handling, prevents many cache flushes, and supports hardware accelerators. Also, keygen and storage stay off main server. Fun times.