4 ms·
Now, if Hadoop & friends stored their data in such a way that you could do a random access read (and used a read syscall) then mmap would _still_ be about 1000
by cmccabe 13y ago
Now, if Hadoop & friends stored their data in such a way that you could do a random access read (and used a read syscall) then mmap would _still_ be about 1000 times faster than a read() call for access if the datum is already in the o/s buffers.
Yeah, it's too bad this method doesn't exist:
http://hadoop.apache.org/docs/current/api/org/apache/hadoop/fs/FSDataInputStream.html#read(long http://hadoop.apache.org/docs/current/api/org/apache/hadoop/..., byte[], int, int)
It's too bad that this doesn't exist: http://www.parquet.io http://www.parquet.io
It's too bad that this doesn't exist either: https://issues.apache.org/jira/browse/HDFS-4953 https://issues.apache.org/jira/browse/HDFS-4953
You are amazingly, astoundingly, misinformed.
It's no surprise that you are wrong about mmap as well. It's only a performance advantage when the PTE (page table entries) are already populated. The PTEs may not be populated, even if the file region in question is resident in memory. And even then it's only about 2x, not "1000 times faster". There was no "serialization and deserialization" involved in our benchmarks, either.
- beagle3 13y ago> Yeah, it's too bad this method doesn't exist: http://hadoop.apache.org/docs/current/api/org/apache/hadoop/... http://hadoop.apache.org/docs/current/api/org/apache/hadoop/..., byte[], int, int) And how, if you may help me, do you access object #n in the stream (for non-trivial n)? Not byte n in the stream, which this appears to give you - but object record #n ? > It's too bad that this doesn't exist either: https://issues.apache.org/jira/browse/HDFS-4953 https://issues.apache.org/jira/browse/HDFS-4953 Your patch is two weeks old, and you expect anyone to be aware of it? > You are amazingly, astoundingly, misinformed. No, I have inherited a hadoop project from people who did it "by the book". And then rewrote it in C with mmap to gain a 20-1000 fold improvement in speed depending on workflow. That had been in 2011, but I'd be surprised if things changed so significantly since then. > It's no surprise that you are wrong about mmap as well. It's only a performance advantage when the PTE (page table entries) are already populated. The PTEs may not be populated, even if the file region in question is resident in memory. And even then it's only about 2x, not "1000 times faster". There was no "serialization and deserialization" involved in our benchmarks, either. Are you using hadoop map reduce, or your own HDFS code? Because every single non-trivial map-reduce jobs that I've met serializes and deserializes objects, spends 90% of its time on I/O, serialization, network, compression and stuff - and that's actually the recommended way to do so. Which is why, incidentally, spark manages to be so much faster (100 times claimed - I've heard of >70) than hadoop - it keeps stuff in memory. It's no surprise that you are wrong about mmap as well - your description is correct if and only if you are going to read the entire file. Which is standard for hadoop map reduce, but actually not that common in a well designed program, and certainly not common in the K world. And you certainly did not understand the 2x vs 1000 times faster comment, which stems apparently from your (faulty) understanding that everything has to be read. In an mmaped file, if you need one byte from the middle, you pay for reading one block from the middle exactly once. You do not need to pre-read everything, and you don't need to manage caching yourself. That's where the 1000x times come from. If you do read everything, it's not 1000x times. But you've possibly wasted 1000x times memory -- and 1000x times I/O to get to that byte in the middle that you need.
- cmccabe 13y agoYou access object #n by using an index. Something that the Parquet file format includes. If you or your team wrote jobs that always read the whole file, you were doing it wrong. No, our HDFS mmap testing did not include MapReduce. I'm pretty sure I've been repeating that over and over as well. Why don't you write your own test program in C if you don't believe me? How about we listen to this Linus Torvalds guy? Have you heard of him? [http://yarchive.net/comp/linux/o_direct.html http://yarchive.net/comp/linux/o_direct.html] Right now, the fastest way to copy a file is apparently by doing lots of ~8kB read/write pairs (that data may be slightly stale, but it was true at some point). Never mind the system call overhead - just having the extra buffer stay in the L1 cache and avoiding page faults from mmap is a bigger win.
- justin66 13y agoI don't have a dog in this fight, but note that the message you quoted - which starts with "right now" - is from over a decade ago. An awful lot has changed.
- cmccabe 13y agoYep. But our benchmarks are from this year, and they show the same thing. The situation might improve when Linux gets support for hugepages on (regular) file-backed mmaps.
- justin66 13y agoWhat is it you're working on exactly? I'm having some trouble putting it together from context but I'm wondering if you're talking about manipulating files using mmap from... the JVM? That's kind of interesting.
- cmccabe 13y agoUsing mmap in Java is not really that interesting-- you just use FileChannel#map, which has been around for a long time. There is no need even for JNI. The only remotely "interesting" part is that Java doesn't provide munmap, so you have to work around that. This whole thread makes me very sad. I might have to write a "Mythbusters" post about Hadoop at some point.