3 ms·
Short stroke in a SSD era is more about trying to keep your data within a page of a SSD. Depending on the flash drive you are using the controller may or may n
by sumtechguy 3y ago
Short stroke in a SSD era is more about trying to keep your data within a page of a SSD. Depending on the flash drive you are using the controller may or may not do some of it for you. One of the easiest perf gains I can sometimes get from a program that is write intensive is to put a small amount of write buffer into the mix. Depending on your OS that can help a lot too. Basically keep my code out of the kernel and off the bus and keep the write block to something that closely resembles what the drive considers a block. If you are doing a bunch of small 8 byte writes randomly in your file you probably are having a bad time as you may quickly exceed the amount of buffer the drive has for that sort of thing. It will start backing you off which bleeds up into the kernel space then very quickly into your program. Keeping them together can help, as you found out.
I see this sort of issue in a lot of programs. As it is a dead easy problem to make in your program. You need to write something out you just splat it out somewhere. With a dozens 1/4/8/16 byte writes instead of one big write. Basically not thinking about how that data is getting into your files. Most of the time that is just fine and not that big of a deal. But as your data set grows or you want better perf you have to worry about it. I usually use something like filemon and can see what is going on. You can see the pattern where there will be a large stack of I/O with hundreds of very small read and writes. While SSDs are an order of magnitude faster than the older drives. They still have their command structure and kernel context switching you have to deal with. You in some cases want to minimize that as it can become a large portion of writing and reading data. As with most optimizations (in this case the drive is faster) we just moved where the bottleneck is (to the kernel and bus typically).
- bcrl 3y agoShort stroking data doesn't writes within a page. It gives the FTL more time to shuffle data around in pages when garbage collecting and erasing pages. Erasing a page tends to take milliseconds of time vs microseconds of time for reads, tens to hundreds of microseconds for writes. NAND flash is typically written in smaller pages (512 bytes in older products vs 4KB-16KB ) but erased in larger pages (32KB+). It's only NOR flash that supports smaller byte granularity write sizes. NOR flash is not as dense as NAND, so it doesn't get used much beyond embedded applications.
- sumtechguy 3y agoexactly. Also My point was more about keeping yourself out of calling into the kernel and paying the price for the context switches. Which can add up quickly.
- bcrl 3y agoContext switches between processes are bad, but going into the kernel for an irq using things like aio and io_uring which can populate events into a result ring buffer doesn't add a ridiculous amount of overhead. Optane DIMMs avoided that, but the consequences were such that the hardware had a lot more complexity. Doing something like an FTL at DDR4 speeds and latency is very difficult. Plus you don't really want to expose hardware that requires wear leveling directly to applications, as a buggy app can wear out (damage) the hardware prematurely.