7 ms·
Clojure on Hadoop: A New Hope
- alexatkeplar 15y agoThis is an entertaining comparison of 8 different map-reduce languages for Hadoop, including a flame-tastic take on Cascalog/Clojure: http://www.dataspora.com/2011/04/pigs-bees-and-elephants-a-comparison-of-eight-mapreduce-languages/ http://www.dataspora.com/2011/04/pigs-bees-and-elephants-a-c...
- swannodette 15y agoThat's not flame-tastic. It's just inaccurate.
- dirtyvagabond 15y agowell, i'd agree it's "entertaining"! had fun reading that, thanks.
- CodeMage 15y agoAs someone who would like to start experimenting with Hadoop in near future, I would appreciate it a lot if you could elaborate on that or point me in the direction of a better comparison.
- swannodette 15y agoPersonally, I dislike LISP odd syntax, the widespread use of side effects in a functional language and the poor abstraction that lists represent over RAM, from a performance point of view — indeed LISP variants often add additional data structures, somehow negating the “LIS” part of the language. In the specific case of Clojure, the fact that a compiled language is compiled into an interpreted one, JVM bytecode, combining a slow dev cycle with suboptimal performance, makes me think Clojure users must be glutton for punishment. Clojure has sophisticated state management. So much for widespread use of side effects. Clojure has high performance data structure implementations tuned for modern hardware. So much for performance. Clojure is compiled on the fly at the REPL and the JVM is one the fastest runtimes out there. So much for slow dev cycle. You may not like the syntax, but boy is that Hadoop query short and sweet.
- Raphael_Amiard 15y agoYeah it's almost hilarious how every critic he is aiming at LISP is fixed in Clojure, where immutable is the default and vector the data structure of choice. Also he clearly doesn't understand the subtleties/differences between compilation and interpretation, and how the two can interlace.
- benatkin 15y agoHe got Nathan Marz' name wrong, too.
- ekoontz 15y ago"LISP has been around some 50-odd years without taking off despite several attempts at its revival... I suspect something is wrong with it" I laughed out loud at that.
- alexatkeplar 15y agoThe guy is a genuinely funny writer, incredibly rare in tech. If you find anything else written by him, let me know.
- moomin 15y agoThe problem is, I'm not sure that line was actually meant to be funny.
- wicknicks 15y agoMy experience with Hadoop tells me that its great for all counting tasks. Makes a lot of sense that it was designed at Google to construct posting lists for their search index. Beyond this sweet spot, it gets really tricky to map your solution to a map-reduce task. The programmer has to rely more on the fact that map/reduce are java black boxes to express everything he needs. Hadoop's big victory is the scale it operates on. Are people here waiting for a "SQL-like-declarative-query-language + Hadoop" hybrid, which lets one write declarative processes to run on large amounts of data? Academia seems to be very motivated to produce something like this.
- scott_s 15y agoI work on Streams at IBM Research. Our solution to handling large amounts of data is called stream programming, and its more general than MapReduce. For a sample of what our language looks like, check out this example: http://publib.boulder.ibm.com/infocenter/streams/v2r0/index.jsp?topic=%2Fcom.ibm.swg.im.infosphere.streams.spl-language-specification.doc%2Fdoc%2Flangoverview.html http://publib.boulder.ibm.com/infocenter/streams/v2r0/index.... And a pdf version of the same: http://publib.boulder.ibm.com/infocenter/streams/v2r0/topic/com.ibm.swg.im.infosphere.streams.product.doc/doc/IBMInfoSphereStreams-SPLLanguageSpecification.pdf http://publib.boulder.ibm.com/infocenter/streams/v2r0/topic/...
- toisanji 15y agoisn't stream programming what storm is supposed to be?
- scott_s 15y agoYes, it is a similar programming model. Some differences (please put "to the best of my knowledge" in front of all of these): - Storm does not allow arbitrary state in operators (what Nathan Marz calls "bolts"). This makes implementing the runtime easier, such as being able to replay tuple sends for fault tolerance, but it limits what kinds of applications one can make. Yes, I'm on board with the idea that we should avoid mutable state as much as possible, but people who build real applications want it. Yet, fault tolerance in our system requires more work, so it's a trade-off. - Storm programs are implemented in Java. Streams applications are implemented in our programming language, which has the rather pedestrian name Streams Programming Language, but usually just SPL. This may seem minor, but it's a big deal. Marz is working on a higher level language in Clojure. Implementing programs in a higher level language enables developers to abstract away many issues related to high performance, distributed systems. I compare it to the difference between writing assembly code and writing C code. (Or the difference between writing Python code and writing C code.) The code that we generate is similar in principle to how one writes a Storm application. Which brings me to... - Storm runs on the JVM, we generate C++ code which gets compiled. Neither Storm or Streams are the first or only in this area. Stream programming is also popular for hardware, but that is usually synchronous and if there's state, it's shared-memory. Storm and Streams are distributed and asynchronous. There are academic distributed streaming systems such as Borealis. The research name for Streams is System S, and there are many academic papers about it, or that use it as a platform for other research: http://dl.acm.org/results.cfm?h=1&cfid=66087472&cftoken=23295126 http://dl.acm.org/results.cfm?h=1&cfid=66087472&cfto... And for the record, I am impressed with Storm.