4 ms·
Hi! Thanks for taking interest in the paper. First of all I'm sorry the setting doesn't meet with your expectations. It's always the problem with benchmarking
by savrus 5y ago
Hi!
Thanks for taking interest in the paper. First of all I'm sorry the setting doesn't meet with your expectations. It's always the problem with benchmarking why one measure one thing and not the other. In my case I was interested in uniformy random distribution since the whole work has been motivated by a task of putting key-value storage backend to NVMe. Hot data is already present in main memory and NVMe is used for a long tail of infrequently-accessed but large amount of data. Alas said maybe some other distribution could match the real workload more closely, I just don't know any better estimation.
It's pretty interesting that you question the working set size. Given the complexity of FTL it won't be surprising if the latency depend on working set size. I dind't do such kind of experiments but I hope this discussion could motivate somebody to take such measurements. Anyway thanks for the provided timings, I should remember them and keep that in mind.
As for section 7.1 QD was something from 16 to 64 but I don't remember exactly. In section 8 more attention is given to to QD and I try to pick the best one.
Section 7.2 could be probably the most confusing due to the read pattern complexity. I mention there that I'm interested in requests where several disjoint blocks are asked at once. Obviously that is transformed into several AIO requests. You have every right to complain about asynchronousity here because the reader waits for all of them to complete before issuing the next request. It's just the AIO interface which makes such kind of compound requests more efficient than issuing them via pread which the section demonstrated.