3 ms·
Don't all physical SSDs have reliability issues? There's a good reason we replicate data across devices.
by why_only_15 4y ago
Don't all physical SSDs have reliability issues? There's a good reason we replicate data across devices.
- shrubble 4y agoWell NVME drives shouldn't fail at a high rate, but if they don't have good local storage capabilities, then yes they have to build something else that is different.
- deelowe 4y agoHow does NVME make local storage more reliable?
- withinboredom 4y agoI think they were saying it shouldn’t be failing, not that it’s reliable. From personal experience, I got a server with two local drives. Both drives ended up failing within 15 minutes of each other. That was an annoying day…
- jon-wood 4y agoCascading failures in drives is a fairly common occurrence unless you're actively trying to avoid it. People will typically just buy several of the same drive, from the same source, because it's easy and seems to make sense. What you've likely done is bought several of the same drive, from the same manufacturing batch, and then brought them all online at the same time. If there's a manufacturing issue, or just an expected lifetime for those drives, you're going to hit at almost the same moment across all your drives. The ideal here is to buy similarly specced drives from multiple manufacturers to reduce the risk. At the very least buy from multiple suppliers to reduce your risk of getting drives from the same batch if this is something you're going to care about.
- sokoloff 4y agoHere’s one fairly recent example: https://www.zdnet.com/article/hpe-says-firmware-bug-will-brick-some-ssds-starting-october-this-year/ https://www.zdnet.com/article/hpe-says-firmware-bug-will-bri... [2020]
- deelowe 4y ago"Well NVME drives shouldn't fail at a high rate" What is it about NVME that means they shouldn't fail at a high rate? I don't understand how the protocol should matter much.
- withinboredom 4y agoI, personally, think a protocol should be resilient and never fail unless there is absolutely no other choice (like TCP where failures are so common that recovering is part of the protocol).