3 ms·
I liked the post (well, as an HDF5 user, I found it depressing...). My main qualm with it was the claim about 100x worse performance than just using numpy.memm
by superbatfish 11y ago
I liked the post (well, as an HDF5 user, I found it depressing...).
My main qualm with it was the claim about 100x worse performance than just using numpy.memmap(). To the author's credit, he posted his benchmarking code so we could try it ourselves. (Much appreciated.) But as it turned out, there were problems with his benchmark. A fair comparison shows a mixed picture -- hdf5 is faster in some cases, and numpy.memmap is faster in other cases. (You can read my back-and-forth about the benchmarking code in the blog's comments.)
One minor complaint about presentation: Once the benchmarking claims were shown to be bogus, the author should have removed that section from the post, or added an inline "EDIT:" comment. Instead, he merely revised the text to remove any specific numbers, and he didn't add any inline text indicating that the post had been edited.
I think the rest of the post (without performance complaints) is strong enough to stand on its own. After all, performance isn't everything. In fact, I'd say it's a minor consideration compared to the other points.
When it comes to performance, I think the main issue is this: When you have to "roll your own" solution, you become intimately aware of the performance trade-offs you're making. HDF5 is so configurable and yet so opaque that it's tough to understand why you're not seeing the performance you expect.
- superbatfish 11y agoAnd one last point. In the blog comments, I wrote this, which I think sums up my view of the performance discussion: ...it's worth noting that many of the tricky things about tuning hdf5 performance are not unique to HDF5. For storing ND data, there will always be decisions to make about when to load data into RAM vs. accessing it on demand, whether or not to store the data in "chunks", what the size of those chunks should be (based on your anticipated access patterns), whether/how to compress the data, etc. These are generally hard problems; we can't blame HDF5 for all of them.
- ngoldbaum 11y agoYou can also see h5py maintainer Andrew Collette's response here: https://gist.github.com/rossant/7b4704e8caeb8f173084#gistcomment-1665072 https://gist.github.com/rossant/7b4704e8caeb8f173084#gistcom...