5 ms·
This story reminds me of the (apocryphal?) claim that C++ running on a single machine can often beat a cluster of Hadoop machines. Scalability is another matte
by codex 15y ago
This story reminds me of the (apocryphal?) claim that C++ running on a single machine can often beat a cluster of Hadoop machines. Scalability is another matter.
- bpodgursky 15y agoThat's a totally problem/domain specific assertion. A single machine with 8 or so cores simply isn't going to make it through a hundred terabytes of data, and many (most?) hadoop jobs are IO-bound anyway. If you find that micro-optimization greatly increases your performance, you probably shouldn't be using hadoop anyway...
- rbranson 15y agoHave you any useful examples of non-trivial Hadoop MapReduce jobs that can process input data at hundreds of megabytes per second?
- jeffffff 15y agobuilding sharded inverted indexes
- beagle3 15y ago> If you find that micro-optimization greatly increases your performance, you probably shouldn't be using hadoop anyway... It's all a matter of what you optimize for. In a recent project, a C++ version, using memory mapped files with a fixed-length record, was about 20 times faster than the hadoop version, and that's just the CPU. The C++ version had: no deserializing needed, everything random access, multiple runs on same data had NO I/O requirements (mmap was cached between runs). It was about as hard to write. So, instead of running a 100-core strong Hadoop install (which is small, but far from trivial), I was able to do with one hefty 8-core machine. Scaling up is "easier" with Hadoop, in the sense that you can just throw money at it and get more EC@/rackspace nodes when needed. But it is a lot of money. Hadoop makes sense if you've got lots of money to waste, and actually need thousands of cores (cause you'll only need a few tens if you effectively use your hardware).
- jeffffff 15y agoIf your working set fits entirely in ram on a single machine, hadoop is the wrong tool for the job. Hadoop is more about getting obscene IO bandwidth from hundreds of hard disk spindles going in parallel, not raw number of cores/parallel computation.
- beagle3 15y agoWhile that is what Hadoop is about, and it is successful at it, it is very far from being efficient; In my experience, you pay 10x compared to good use of the same hardware. "obscene IO bandwidth" is right, but at 10% of the practical maximum. This is a tradeoff that people in computing have been happily taking for years -- e.g. using a higher level language is a similar kind of tradeoff (use Python instead of C - pay x10 in performance, get x10 shorter development time). But this tradeoff only makes sense if the x10 thing you get is cheaper than the x10 thing you pay for. With hadoop, you pay x10 for hardware/computing time and more for administration (unless you have some kind of EC2 elastic mapreduce, at which point you pay x20 for hardware/computing but no administration cost), and you save some development time. Two hadoop projects I was consulting on (or rather, consulting _off_ hadoop) were paying ~$10K/month each to Amazon for a while, when <$20K in programmer time for each reduced it to working on one beefy machine that was already located in the office, with ~$100/month cost (power, cooling).
- jeffffff 15y agoI've found the slowdown from hadoop to be very problem dependent. I've seen anywhere from no slowdown to 1000x slower. It sounds to me like those projects should never have been using hadoop in the first place. If your problem is not IO constrained naturally, moving it to hadoop will make it IO constrained. In my experience doing that will get you a 10x-100x slowdown versus a more appropriate solution on the same hardware. If your problem is IO constrained, for example processing 10 TB of log data, hadoop will get all your disks to 100% utilization just as well as anything else will. It will leave your processors horribly underutilized, but that's just the nature of IO constrained problems. If you could get 10 TB of ram dedicated to your dataset then you could easily get a massive speedup over the hadoop solution but that's just not realistic.
- varelse 15y agoDoesn't sound apocryphal at all. A well-written piece of custom code for a single multi-core CPU can easily run circles around an algorithm shoehorned into a useful but somewhat generic multiprocessor framework where the interconnect between the processors is the real culprit making the computation communication-limited. A related specific example would be molecular dynamics on GPUs versus CPUs. Not only are GPUs faster, they're faster than CPUs can be at this time because of communication issues. As in you can throw as many CPUs and servers as you want at the problem and they'll never beat a single $600 consumer GPU with its 120+ GB/s internal bus (note the absence of hyperbolic claims of 100x to 1000x faster, just faster). Or in other words, strong-scaling tasks don't work out so well on a weak-scaling architecture.