4 ms·
he is wrong about a few things: 1: checksuming is worthwhile, I have had silent data corruption on both ssd/hdd. 2: compression is also worth it: home dir 74G/
by cosmin800 5y ago
he is wrong about a few things:
1: checksuming is worthwhile, I have had silent data corruption on both ssd/hdd.
2: compression is also worth it: home dir 74G/103G, virtual machines dir 1.3T/2.0T
3: zfs was never supposed to be fast, data integrity is the target.
4: zfs does not need a manual repair tool, is automatic and data at rest is always consistent.
5: in the future x, y, z - yeah sure.
- ziml77 5y agoBy default, ZFS does need scrubs to be performed manually. Ask Linus Sebastian about that one https://www.youtube.com/watch?v=Npu7jkJk5nM https://www.youtube.com/watch?v=Npu7jkJk5nM TL;DW: Data in their massive 1PB server was suffering from bitrot because there was no scheduled scrub to repair the bad data. And they couldn't tell how bad the situation was because without any scrubs happening, the stats on data integrity were inaccurate.
- cosmin800 5y agoI agree, the scrubs are performed manually, but enabled by default in debian/ubuntu, via crontab every two weeks (mdadm consitency check is also triggered from cron) about the data loss in the video, mistakes easy to spot: 1st: using seagate. 2nd: installed by us and never updated 3rd: insufficient reading of docs before going all in on zfs. 4th: buying more seagate drives ;)) I think they had a way higher chance of losing their data going the usual stack mdadm/lvm/ext4/luks/btrfs, I think mastering those is harder than mastering zfs.
- fredoralive 5y agoI think for Linus Media Group the main "meta" issue is that they don't have a dedicated member of staff to handle boring day-to-day IT / sysadmin tasks that you don't make videos about. A video about building a crazy storage server is content for a video so gets done, but somebody needs to make sure its still working / updated, and you don't make videos about routine maintenance that so its forgotten. Although everyone else can learn the important lesson that RAID / ZFS isn't magic and you need to have stuff setup correctly and monitored. The fact that RAID isn't a backup as well[1]. [1] Although if the LMG servers affected are just for data hoarding raw footage that is unlikely to be needed again, it's possible the risk / cost balance pushes away from backups and just relying on RAID, but that's a niche case (and they lost the gamble...).
- Dylan16807 5y agoWhat also makes it niche is getting the drives for free. When you're paying upwards of $30k for a petabyte of cold storage, tape is pretty tempting.
- ziml77 5y agoYes, their setup likely would have been configured, monitored, and maintained properly if they had an IT guy. But they made the (easy to make) mistake of thinking that having enough tech knowledge means you don't need a proper IT/systems department. I'm certain at least that Linus knows that RAID isn't backup. And I'm Linus is going to try had to get the data back, but it seems to me that this isn't some devastating failure for him.
- gjvc 5y agoI bet you they also bought all the same make, model, and batch/vintage drives. If you are building a storage array, do not do this. Ensure that you are using a variety of drive types (obviously same size and interface technology). Doing so guards against the danger of too many drives going wrong at the same time (within the same time window) causing a failure from which it is impossible to recover.