3 ms·
This is why I’m looking for more info. This is talking about how they are already sending these out to Hyperscale customers, and they will eventually be in indi
by techdragon 3y ago
This is why I’m looking for more info. This is talking about how they are already sending these out to Hyperscale customers, and they will eventually be in individual customers hands… so it would be good to know what file systems and cache setups are needed in order to benefit from this.
I used to just rely on FreeNAS but it’s not as straightforward anymore. I’m having to consider Linux and looking at bCache FS and ZFS vs BTRFS and how all this compares on an all PCIe (m.2) flash drive setup… where a drive (or two) at this sort of size (~50TB) would make a great second backup copy of the flash array that can be started and stopped to periodically make the backups to save power… but then you have to think about copy efficiency since I don’t want to have them wasting read bandwidth from the flash drives… and the best is usually like to like file system copy ZFS -> ZFS and BTRFS -> BTRFS … so it adds another complication into a mix that is already far from simple.
So it’s become something I’m keeping an eye out for… hopefully before it becomes something I may purchase, someone will have already done a good writeup.
- wtallis 3y agoThe simple answer is that the software stack necessary for consumers to effectively use these storage devices does not yet exist. Hyperscalers are pretty much defined as the ones who are big enough to do their own software stack. For the rest of us, we have some good components to work with but also some major gaps that won't be filled anytime soon. ZFS is a great improvement over traditional hardware RAID systems, but is still in many ways clearly a descendant of them. BTRFS has a slightly different mix of features from ZFS that make it a bit more flexible and a better choice for consumers who don't buy drives by the dozen. Neither has a great solution for caching/tiering with SSDs and hard drives. Ceph has a lot of features that would otherwise be almost exclusive to the hyperscalers, but is too complicated for something like a turnkey NAS. bcachefs aspires to eventually have most of the features you would want for a non-clustered storage system. Zoned storage is something many of the above already have some degree of support for, but as a paradigm it has not even started showing up in the consumer computing ecosystem so many of the issues with adopting zoned storage for consumer systems aren't even being worked on.
- XorNot 3y agoZFS has the concept of a separate intent log device though, I wonder if that would be enough to make these types of disks work with it? I have no problem putting a couple of terabytes of flash in front of a ZFS array if it means I can have 40TB of redundant storage in 3 disks.
- wtallis 3y agoThe ZFS intent log is an example of a supremely disappointing, underwhelming way to integrate SSD storage into your system. Fundamentally, it's just a workaround for the fact that most of the write caches preexisting in the storage stack are volatile caches and thus not safe. Putting the ZIL on an SSD allows you to get the safety without sacrificing the performance benefits of the write caching (performance that everyone has come to rely on). But the ZIL doesn't help with read performance, and the data you write to it is never even read unless you have a crash or power failure. And then for the read caching, ZFS has L2ARC as an entirely separate feature. So ZFS does technically have ways to take advantage of SSDs to improve a hard drive storage array—but I don't think anyone would consider it to be an ideal solution, more of a pragmatic minimum viable feature.
- deleted 3y ago[deleted]
- wmf 3y agoY'all keep saying "these" but SMR and HAMR are unrelated.
- wtallis 3y agoIn theory, yes, but the bottom of the article explains that Seagate currently plans for their upcoming 24TB drive to be the last new non-SMR drive and all larger drives (28+ TB) are planned to be SMR. So in practice SMR and HAMR will be going hand in hand, at least from this vendor.
- wmf 3y ago
- klodolph 3y ago> … so it would be good to know what file systems and cache setups are needed in order to benefit from this. I’ve worked on these systems. You might as well be asking F1 drivers for tips to help you commute to work. The hyperscale stuff is not built on top of ordinary filesystems. It’s all clusters of machines, and error correction is handled at the level of clusters. If you’re evaluating systems like ZFS and BTRFS, then you’re already working with a radically different tech stack. At scale, your file metadata is stored in a distributed database of some kind, and the file contents are stored with forward error correction across multiple machines in a cluster. Or something similar, but stored across multiple clusters. The pricing models for cloud storage are designed around the usage patterns—like, if you know that some data is going to stick around for 90 days, then SMR is a win. If a file could get deleted at any time, then SMR is a loss.
- sh34r 3y ago> You might as well be asking F1 drivers for tips to help you commute to work. Boston drivers: "hold my beer."