3 ms·
As someone who has never used Map-Reduce before, something about this implementation makes the technology feel 100% more accessible to me.
by awhitty 12y ago
As someone who has never used Map-Reduce before, something about this implementation makes the technology feel 100% more accessible to me.
- andrewguenther 12y agoYou should check out mrjob[1]. It wraps the Hadoop streaming API and makes it super easy to write MapReduce jobs in Python. I find it much easier to understand than this implementation. [1] https://pythonhosted.org/mrjob/ https://pythonhosted.org/mrjob/
- JensRantil 12y agomrjob looks nice. If you have a Hadoop cluster. But for "medium sized big data problems" FileMap is a very viable alternative if you have people who know their way around a terminal. Especially if you'd like to have something set up pretty fast. For 500 GB of data setting up Hadoop (steep learning curve; name nodes, jobtrackers, thrift API:s, datanodes, zookeeper and whatnots) is a lot of heavy lifting. Not to mention administering it. Sure you have Cloudera et al., but I'm still trying to figure out if they really make things easier or worse when it two weeks later comes to figuring out why something is broken, or how to start/install additional Hadoop components.
- andrewguenther 12y agomrjob has really good integration to Amazon's Elastic Map Reduce, makes it totally painless. I had to analyze 1TB of logs for my thesis and in less than 8 hours I discovered mrjob, wrote my job, and successfully ran it on EMR. Granted I have prior experience with MapReduce, but even to a newcomer, I can't imagine that would add too much time.
- JensRantil 12y agoIf your security policy allow putting stuff into Amazon... ;)
- JensRantil 12y agoYes, FileMap was also extremelly easy to set up.
- jasode 12y agoIf you're talking about other "Map-Reduce" as in Hadoop MapReduce being "less accessible", it's because Hadoop has a ton of housekeeping infrastructure. Hadoop is aware of the health of other machines and therefore slave jobs can be self-healing, any job errors are properly routed and centralized, etc. Leaving out the housekeeping stuff, the core "map & reduce" part of Hadoop is very simple. This FileMap doesn't seem to have all that infrastructure intelligence, hence it's "simpler". It's like talking about bwr nuclear reactors. On the one hand, it's just a big glorified pot of boiling water. What's so hard about that?! The complexity isn't boiling the water, it's temperature regulation and redundant failsafes that prevent an explosion of radioactive materials[1] that makes it harder to implement than a kettle of hot water for brewing tea. [1]http://en.wikipedia.org/wiki/Chernobyl_disaster http://en.wikipedia.org/wiki/Chernobyl_disaster