3 ms·
> I worry about the pathological cases. Imagine you have an append-only log, and you write and fsync() one byte at a time. Each time you write a byte, the entir
by throwawaylinux 3y ago
> I worry about the pathological cases. Imagine you have an append-only log, and you write and fsync() one byte at a time. Each time you write a byte, the entire flash block (are they still 4KB these days?) has to be erased. So you end up chewing through 4000 durability cycles on your SSD, whereas if you had waited to write an exact 4KB block, then you'd use only 1.
NAND pages are what, something around 16-64kB these days, and block sizes are 10s of MB. In SSDs the logical size of those things may end up being larger at the FTL level if they are ganged together, but that's about your minimum.
Those are program and erase units respectively. With NAND, you do not erase a block to write. You can program pages in a block incrementally, and then you have to erase the entire block before reprogramming any.
Flash translation layers have to make this look like a disk. To do that they will do something like gather writes into a page size chunk in a small cache that is non-volatile or can persist itself on power failure. Then that chunk is programmed out to a free page. A mapping structure records the new NAND location of the logical block addresses you wrote. And a garbage collector comes along behind and compacts and frees data in block size units, erases them, and puts them on the free list.
It's a log structured filesystem with one file (the block device), if you've read any of those papers.
If you write+fsync to sector 0 of your disk 500 times, that data will get stored at 500 different places on the NAND (ignoring larger persistent caches in front of the NAND that some devices have). Selecting what pages to use and what to garbage collect etc is all part of wear leveling that is intended to prolong live of the drive. That's why endurance ratings tend to be in total writes to the drive, not writes to any particular sector.
To handwave the numbers, if you have a 100GB SSD that might be implemented with 105GB of NAND. Then if you had a program/erase endurance of 1,000 cycles you will be able to write that block 25.6 billion times. Your drive can do about 20-30 thousand QD1 writes per second, so about 12 days of that write+fsync block workload. Again assuming no front end cache on it.
You definitely can wear out NAND drives if you write a lot to them (large streaming writes would be easier to do it with), but regardless of what you do in the software layer, the drive will (should) last for its rated endurance no matter what kind of write patterns your application does. A block is a block as far as the NAND sees.
For consumer stuff you generally have to be pretty extreme or have malfunctioning software to wear them out as far as I've seen.
> A hard drive is fine with this access pattern, though you'd probably be doing a lot of seeking to update fs metadata and it would be so slow that you'd come up with some other way. So maybe this pathology only exists in my mind.