4 ms·
You access object #n by using an index. Something that the Parquet file format includes. If you or your team wrote jobs that always read the whole file, you w
by cmccabe 13y ago
You access object #n by using an index. Something that the Parquet file format includes. If you or your team wrote jobs that always read the whole file, you were doing it wrong.
No, our HDFS mmap testing did not include MapReduce. I'm pretty sure I've been repeating that over and over as well. Why don't you write your own test program in C if you don't believe me?
How about we listen to this Linus Torvalds guy? Have you heard of him? [http://yarchive.net/comp/linux/o_direct.html http://yarchive.net/comp/linux/o_direct.html]
Right now, the fastest way to copy a file is apparently by doing lots of ~8kB read/write pairs (that data may be slightly stale, but it was true at some point). Never mind the system call overhead - just having the extra buffer stay in the L1 cache and avoiding page faults from mmap is a bigger win.
- justin66 13y agoI don't have a dog in this fight, but note that the message you quoted - which starts with "right now" - is from over a decade ago. An awful lot has changed.
- cmccabe 13y agoYep. But our benchmarks are from this year, and they show the same thing. The situation might improve when Linux gets support for hugepages on (regular) file-backed mmaps.
- justin66 13y agoWhat is it you're working on exactly? I'm having some trouble putting it together from context but I'm wondering if you're talking about manipulating files using mmap from... the JVM? That's kind of interesting.
- cmccabe 13y agoUsing mmap in Java is not really that interesting-- you just use FileChannel#map, which has been around for a long time. There is no need even for JNI. The only remotely "interesting" part is that Java doesn't provide munmap, so you have to work around that. This whole thread makes me very sad. I might have to write a "Mythbusters" post about Hadoop at some point.
- beagle3 13y ago> You access object #n by using an index And in a C (or K) mmapped file, you access the n's object by accessing a memory location. No index lookup, whether O(log n) or a hash table. I'm indeed not familiar with Parquet - Hadoop failed me so badly I do not feel like I need to keep up to date with every new thing. I might take another look when when it's been around for a while. > Right now, the fastest way to copy a file is apparently by doing lots of ~8kB read/write pairs That again discusses reading the whole file. When mmap delivers the huge gains is when you don't actually need the whole file - rather, when you need parts of it that you discover while reading. Which in many workloads is the case (or would be the case if you did not have to conform with some framework's constraints about the order and kind of I/O you can do). mmap gives you hardware virtual memory and assistance for caching; Every other solution is a software virtual memory and implementation for caching. While it has its own constraints (e.g., you have to store data in a directly usable layout, rather than "disk frozen" layout), if you do that, things do work exceptionally well when you do.