6 ms·
This paints a really misleading picture of what is available in Hadoop right now. Columnar data formats? Yep, we've got them-- see Parquet [http://parquet.io/
by cmccabe 13y ago
This paints a really misleading picture of what is available in Hadoop right now. Columnar data formats? Yep, we've got them-- see Parquet [http://parquet.io/ http://parquet.io/]. Yes, Hive spends a lot of its time reading and writing data. This is because it decomposes SQL queries into sets of MapReduce jobs, all of which must take input from the filesystem and write it to the filesystem. That's one of the reasons Cloudera Impala was developed [http://blog.cloudera.com/blog/2012/10/cloudera-impala-real-time-queries-in-apache-hadoop-for-real/ http://blog.cloudera.com/blog/2012/10/cloudera-impala-real-t...].
If you don't like for loops, yep, we've got that too-- use Scala or another functional programming language to write your MapReduce jobs. Or use SQL, which is a lot more powerful than APL, and a lot more accessible to businessey types than functional programming. SQL is a declarative programming language, by the way-- I don't see any for loops there.
The efficiency argument makes no sense either. Does the author understand that a just-in-time dynamically recompiling virtual machine is faster than an interpreter? If so, he doesn't mention it anywhere. You know, sometimes old technologies are just... old. You could at least compare Java and the JVM to something like Lisp machines.
People overestimate the gains that are to be had from mmap. I am currently in the process of adding mmap support to HDFS and I know what I am talking about. mmap gives gains, but only when the data is coming from memory, and only when you can reuse the mmap. Otherwise, you're better off reading even a 512 MB file via read() and write(). The reason is that syscall overhead is not that high on modern UNIXes (like Linux), and page faults involve a transition into kernel space anyway.
- gngeal 13y agoOr use SQL, which is a lot more powerful than APL How could that possibly be true? What does "powerful" mean? Certainly not that it can compute things that APL family languages can't, and the design of APL/J/K etc. as languages is so much better that it isn't even fun. SQL is the COBOL for the 21th century. Does the author understand that a just-in-time dynamically recompiling virtual machine is faster than an interpreter? As memory latencies relative to instruction execution times go up, this factor becomes less and less relevant. An even worse problem is that your "just-in-time dynamically recompiling virtual machine" can't even vectorize properly, whereas APL has high-level array manipulating operators that make it comparatively easy to generate vectorized code. Your "just-in-time dynamically recompiling virtual machine" would have to "intelligently" recover high-level patterns from seemingly nondescript for() and while() loops to do that. Guess what: it doesn't! When you consider the fact that memory latencies matter these days, the fact that an AVX vector instruction can do a whole lot of work in a single cycle (or a few of them at most), and the fact that the integer units that would otherwise be idle can handle dispatching the next primitive operation in the meantime, there's a potential for even an interpreter to kick your "just-in-time dynamically recompiling virtual machine"'s ass.
- cmccabe 13y agoYes, APL is Turing-complete, whereas SQL is not... until you consider vendor extensions. I guess I should have included the obligatory "I know that..." in the original post. The fact remains that SQL allows you to do complicated queries without having a computer science degree, something I have never seen with APL/J/K, Java, or any other programming language. SQL arrived on the scene in 1974, by the way-- it's not a "21st century" anything. I agree that the JVM does not use memory well. For example, it has a high per-object memory overhead. And there are things such as lock widening which are a real design problem. But even considering these things, the JVM still manages to consistently beat other choices such as Perl, Python, Ruby, and so forth. Yes, vector instructions are great, and Java can't really use them. There are also other builtin instructions such as the CRC ones which would be nice to have at the language level. We have been adding JNI segments to make use of these. It isn't the syntax of loops that is the problem, but rather, the presence of side effects in the language, that makes auto-vectorization difficult. Probably annotations should be added to make this easier, like Intel has done for C/C++. None of this means that APL is going to come back. The scientists and engineers who originally used it back when 4kb was a lot of memory are just going to use R, MATLAB, Mathematica, or another choice like that and gain a lot of the same micro-optimizations you love. People who want to do big queries over giant data sets are going to keep using Hadoop. The one thing that could maybe dethrone Hadoop is a system that made use of GPUs (graphics processors). Some of them have 512 cores now-- that's a lot of power, and Java can't really harness it at all. But GPU processors are really restricted in how they can communicate and how much memory they can access, so it would not be as much of a general purpose solution.
- beagle3 13y agoI have used Hadoop; it is miserably slow (as in 10 to 1000 times slower AND a lot harder to work with) then the comparable K program. If mmap doesn't bring you significant gains, you are doing stuff wrong. E.g. You serialize objects. Don't. Hadoop is horrible. I've heard Spark manages reasonable performance (subject to the shackles that Hadoop compatibility entails), but haven't used it. If you think SQL is "more powerful" than APL... Qualify that. Because e.g. APL is Turing complete, whereas SQL isn't.
- technolem 13y agoAs a sidenote, SQL is: http://stackoverflow.com/questions/900055/is-sql-or-even-tsql-turing-complete http://stackoverflow.com/questions/900055/is-sql-or-even-tsq.... Not that you'd even want to use it as such.
- beagle3 13y ago:) a corollary of Greenspun's tenth rule: every language spec will be revised at least until it is possible to implement a slow, incomplete version of Common Lisp inside it.
- justin66 13y agoWhy are people attributing k's magical powers to mmap? That's something which should be marginally faster when it's faster, if every other product that uses one or the other can be used as an example.
- beagle3 13y agoSome of k's magical powers come from mmap > That's something which should be marginally faster when it's faster, No. If used properly, it is crazy faster. Let's say what you need from the data is element no. 1821942 in the stream. Hadoop standard: read & deserialize & discard 1821941 items (just so you can get to the 1821942nd item). read & deserialize and use one item. K standard: (equivalent to): seek directly to element 1821942, read it and use it. except .. K does it slightly faster than that, using mmap: It just accesses the location for this element in memory. If it is not in memory, then the operating system will arrange for it to appear in memory in a way that's usually more efficient than read. If it is already in memory, then it is exactly one memory read. Now, if Hadoop & friends stored their data in such a way that you could do a random access read (and used a read syscall) then mmap would _still_ be about 1000 times faster than a read() call for access if the datum is already in the o/s buffers. The magic is not that K can read stuff faster -- e.g., it can't read a hadoop stream faster than hadoop can. The magic is that K makes it easiest to just store your data in a random-access mmapable way, so that mapping the data back in takes ~10ms, and from then on, you only pay for the data you actually use without having to explicitly read it -- and usually, you pay less than you would have paid if you explicitly read it.