4 ms·
This is blown out of proportion... actually increase is probably a factor 10-20X, not 100s. The fact that EMR is used is a problem, provisioning, bootstrapping
by ec664 12y ago
This is blown out of proportion... actually increase is probably a factor 10-20X, not 100s. The fact that EMR is used is a problem, provisioning, bootstrapping the cluster alone accounts for probably half the time.
The fact that shell commands were run repeatedly means that the data ends up in the OS buffer cache and basically in memory.
I'm not discounting that CLI is faster than Hadoop by an order of magnitude on small datasets. Nor will I dive into Hadoop vs CLI. The answer to all that IMO is that it depends. And in this case, it's not well warranted.
What I do take exception to is the Fox News style headlines that are disproportional to the truth. EMR != Hadoop.