5 ms·
At one of my previous employers, they built a massive "cloud" storage system. The underlying file system was ZFS, which was configured to put its write logs ont
by _ak 11y ago
At one of my previous employers, they built a massive "cloud" storage system. The underlying file system was ZFS, which was configured to put its write logs onto an SSD. With the write load on the system, the servers burnt through an SSD in about a year, i.e. most SSDs started failing after about a year. The hardware vendors knew how far you could push SSDs, and thus refused to give any warranty. All the major SSD vendors told us SSDs are strictly considered wearing parts. That was back in 2012 or 2013.
- justin66 11y agoIf your SSDs were wearing out after a year, and were warrantied for a year, I'm guessing you weren't using "enterprise" SSDs?
- skrause 11y agoEven enterprise SSDs like Samsung's have a guarantee like "10 years or up to x TB of written data". So if you write a lot you can lose the guarantee after a year even with enterprise SSDs.
- _ak 11y agoThey were enterprise models (i.e. not cheap), but they had no warranty in the first place. Every single hardware supplier simply refused to give any. I _guess_ because of the expected wear and tear.
- justin66 11y agoThat's interesting. I don't even know how to buy these things without a warranty. Were they direct from the manufacturer?
- rsync 11y agoJust a note ... we use SSDs as write cache in ZFS at rsync.net and although you should be able to withstand a SLOG failure, we don't want to deal with it so we mirror them. My personal insight, and I think this should be a best practice, is that if you mirror something like an SLOG, you should source two entirely different SSD models - either the newest intel and the newest samsung, or perhaps previous generation intel and current generation intel. The point is, if you put the two SSDs into operation at the exact same time, they will experience the exact same lifecycle and (in my opinion) could potentially fail exactly simultaneously. There's no "jitter" - they're not failing for physical reasons, they are failing for logical reasons ... and the logic could be identical for both members of the mirror...
- dboreham 11y agoThis is good advice, and fwiw the problem it addresses can happen in spinning drives too. We had a particular kind of WD drive that had a firmware bug where the drive would reset after 2^N seconds of power-up time (where that duration was some number of months).
- watersb 11y agoFWIW I bricked a very cheap consumer SSD by using it as write log for my ZFS array. This was my experiment machine, not a production server. Fortunately I had followed accepted practice of mirroring the write cache. (I'd also used dedicated, separate host controllers for each of these write-cache SSDs, but for this cheap experiment that probably didn't help.) So yes this really happens.
- leonroy 11y agoWe ship voice recording and conferencing appliances based on Supermicro hardware, a RAID controller and 4x disks on RAID 10. We tried to mitigate the failure interval on the drives by mixing brands. Our Supermicro distributor tried to really dissuade us from using mixed batches and brands of SAS drives in our servers. Really had to dig in our heels to get them to listen. Even when you buy a NAS fully loaded like a Synology it comes with the same brand, model and batch of drives. In one case we saw 6 drive failures in two months for the same Synology NAS. Wonder whether NetApp or EMC try mixing brands or at least batches on the appliances they ship?
- baruch 11y agoI can tell you that EMC and IBM both use the same drives from the same batch in an entire system of tens to hundreds of drives and while I don't know about all cases completely I did oversee a large number of systems and drives and there was never a double disk failure we had that completely took two drives. With a proper background media scan procedure you also reduce the risk of a media problem in two different drives. Ofcourse, the SSDs we use are properly vetted for design issues and bugs in the firmware actually get fixed for us in a relatively timely manner. You get that level of service with the associated large volume.
- andrepd 11y agoIn a sense, so is all storage hardware. It's just a matter of how long it takes before it fails. This goes for SSDs, HDDs, Flash cards, etc.
- bluedino 11y agoLogging to flash storage is just asking for issues. We bought a recent model of LARGE_FIREWALL_VENDOR's SoHo product, and enabling the logging features will destroy the small (16GB?) flash storage module in a few weeks (!). The day before we requested the 3rd RMA, the vendor put a notice on their support site that using certain features would cause drastically shortened life of the storage drive, and patched the OS in attempt to reduce the amount of writes.
- TD-Linux 11y agoLogging to poorly specified Flash storage is the real problem. They were likely using a cheap eMMC flash, which are notorious for having extremely poor write leveling. Unfortunately the jump to good flash is quite expensive, and often hard to find in the eMMC form factor which is dominated by low cost parts.