3 ms·
I don't get where you got the notion that you need to understand the failure model to rely on it. It's certainly good to, and it will prevent over provisioning,
by bluecmd 10y ago
I don't get where you got the notion that you need to understand the failure model to rely on it. It's certainly good to, and it will prevent over provisioning, but a lot of storage solutions today are so complex that understanding the failure models is practically impossible. S3? Super-black box.
Basically all you need to define is the expected output. "99.9pct of my reads are within 1ms" + "The read response is correct 99.9999% of the time" and then verify. Eject the disk if that doesn't hold, restore from replica. The rate of replacements is a factor of how close you are to the true failure model.