7 ms·
in other words, "I used a chainsaw to cut an apple and it SUCKED at it." If you're processing an amount of data that comfortably fits in memory on a single mac
by potatoyogurt 8y ago
in other words, "I used a chainsaw to cut an apple and it SUCKED at it."
If you're processing an amount of data that comfortably fits in memory on a single machine, then obviously Hadoop is going to perform poorly in comparison. The costs of scheduling a job onto N mappers/reducers, transferring code to each node, waiting for the slowest mapper/reducer to finish, transferring data from mappers -> reducers, replicating output on HDFS, etc. are well-understood. It's true that many people try to use Hadoop when they'd be better served with simpler solutions, but that does not justify the amount of shade that the author throws at it.
- itronitron 8y agoHadoop has a very narrow best fit use case but it has been oversold as the best solution for big data.
- potatoyogurt 8y agoBest fit is something you can argue a lot about. There are a lot of data processing tools out there now, many of which have come out after Hadoop. But if the comparison is against some process running on a single machine, then the use case is not narrow. It includes basically anything where you're processing more than 1TB of data in non-trivial ways (i.e. not just a map operation) and are okay with batch processing.
- PeterisP 8y agoThe issue with that is that for most organizations their key business data (and all its recorded history) fits in RAM of a sufficiently beefy workstation. They want to call it Big Data to stroke their egos, and properly acquiring, cleaning and integrating that data can take a LOT of effort so that data can be quite expensive and worthy of any glorious label they can think of; but my experience is that processing more than 1TB of meaningful data actually is a narrow use case, which matters in two specific categories: the (relatively few) very large multinational companies, and processing of raw video/audio/image data; and the majority of people working on data analysis end up with business needs that can be satisfied by relatively simple methods on relatively small datasets.
- potatoyogurt 8y agoI agree in general. But I think you underestimate how large the set of use cases are where people are processing > 1TB of data. This includes quite a bit of the adtech industry, for instance, even many startups. It also includes data warehouses for other industries, such as in health tech. Of course, these people generally are experts, since it is their core business, so they know well what tools they need. I agree that for some analytics department in a random company whose core business isn't processing data, Hadoop is more likely to be a resume item than something that's really needed.
- fiddlerwoaroof 8y agoIn my experience, people generally go the other way: they say “we need spark/Hadoop/Cassandra because we have Big Data” when they have a 30GB dataset that is best handled on a beefy EC2 instance with boring tools.
- itronitron 8y agoA lot of those > 1TB data sources are very standardized, they can be mapped to a schema, in which case indexing the data supports interactive queries and analytics. Hadoop seems well suited for data that needs to be processed in very different and changing ways.
- mmt 8y agoI think this is the other point that's so often missed/ignored in the "big data" discussion: there's a middle ground between everything-fits-in-memory and must-be-distributed. The Adam Drake article alludes to it only in the last sentence by mentioning traditional relational databases as an alternative. For workloads that are relatively I/O-heavy and CPU-light, it's very hard to beat local SSDs (or even HDDs in enough quantity) attached to a single [1] node, if the competition is distributed storage attached by ethernet. It only takes a couple 600MB/s SSDs to saturate a 10GE. A server with 48 lanes of PCIe 3 slots could take the I/O of 78 of them. 100Gb/s networks are getting close. For upwards of $1k per server (NIC and switch port) one can bring that ratio closer to 4:1 from 39:1. I'd expect this is the attractive route for anyone with CPU needs that can't be met by a 4-socket server. [1] Yes, there can be more than one node with copies for redundancy, as has been complained about elsewhere in the sub-thread, or even scalability
- ljw1001 8y agobut the chainsaw analogy was a winner on many levels.
- mmt 8y ago> are well-understood. It's true that many people try to use Hadoop when they'd be better served with simpler solutions I posit that these two assertions are contradictory. My own understanding of the term "well understood" is that it is synonymous with "widely understood". If many people are still making the mistake of using Hadoop when those costs outweight the benefits, it seems that understanding isn't quite wide enough. That said, although I have a grasp of when the tradeoff is so loopsided as to be obvious, I don't know where to go (or where to point other people to go) for a better understanding of where the boundary is. Where should we go to better learn that understanding of those costs?
- potatoyogurt 8y agoThey're well understood by anyone who has used these technologies professionally. I probably should have been ore precise with my language, though. It's really a grey area as to where the boundary is and it depends a lot on your specific application. But as a general rule of thumb, my feeling is that for tens to hundreds of GB, you should consider it. And for TBs or more, you almost certainly want to be doing something distributed. Hadoop isn't necessarily the best option then, but it's a powerful tool. I don't know if there's any resource out there that really goes deep into the tradeoffs involved though. There probably is, given how popular the subject is, but I'm not aware of one. The problem with the article is that if it's for a general audience that doesn't understand the tradeoffs of a system like Hadoop, it really paints a picture that it is just a bad, slow tool. It barely acknowledge just how rigged the comparison is at all, aside from mentioning that you might need something like Hadoop for really big data in the conclusion, while it is peppered with unnecessarily snide comments about Hadoop that will probably be more memorable. I think it is liable to leave readers more confused about the tradeoffs involved after reading than before.
- mmt 8y ago> They're well understood by anyone who has used these technologies professionally. I probably should have been ore precise with my language, though So.. It's well understood by True Scotsmen? :) I certainly understood that there are experts in the field who have a deep, even intuitive understanding of how and when to use which tools. To the extent that those experts don't communicate that knowledge to a broader audience while at the same time may be advocating for the use of the tools, they bear some responsibility for the misuse. My point wasn't so much that you used imprecise, but rather, that the statement about how well it's understood was inaccurate or irrelevant (depending on which definition you were going for). > But as a general rule of thumb, my feeling is that for tens to hundreds of GB, you should consider it. And for TBs or more, you almost certainly want to be doing something distributed While a 3-4TB cutoff makes some sense if ones workload has to remain in-memory for performance reasons, that can't be anywhere near the cutoff for any kind of workload that could stand to read from SSDs. > I don't know if there's any resource out there that really goes deep into the tradeoffs involved though. There probably is, given how popular the subject is, but I'm not aware of one. I would hope so, but I'm not so sure. It may not even need to be very deep, something akin to the "5 minute rule" for memory/disk caching. Mostly, I'm not convinced that the subject of tradeoffs is actually popular, so much as just using the tool without considering them is. > The problem with the article is that if it's for a general audience that doesn't understand the tradeoffs of a system like Hadoop, it really paints a picture that it is just a bad, slow tool. If we're talking about the adamdrak.com 233x article, I have to disagree, as my read of it was that it focused on evangelizing the "under-used approach for data processing" of "standard shell tools and commands". > it is peppered with unnecessarily snide comments about Hadoop that will probably be more memorable That's certainly not a charitable interpretation, and I would hazard that it's not even fair or factual (as to "peppered", at least). It's mentioned only a handful of times: > Command-line Tools can be 235x Faster than your Hadoop Cluster I agree that click-bait can be considered snide. > I was skeptical of using Hadoop for the task, but I can understand his goal of learning and having fun with mrjob and EMR To me, this comment, this very first mention of Hadoop in the intro, made it clear that this was "rigged" in that the "competition" was neither competing nor truly concerned about performance. > while the Hadoop processing took about 26 minutes (processing speed of about 1.14MB/sec). Merely a factual summary. Nothing snide that I could detect. > Although Tom was doing the project for fun, often people use Hadoop and other so-called Big Data ™ tools for real-world processing and analysis jobs that can be done faster with simpler tools and different techniques. This seems like just a restatement of the admission in the introduction plus the assertion (that I believe even you agree with) that many people mis-use Hadoop when it's not called for. > The resulting stream processing pipeline we will create will be over 235 times faster than the Hadoop implementation and use virtually no memory. Again, just another factual summary, with no snideness I could detect. > While we can certainly do better, assuming linear scaling this would have taken the Hadoop cluster approximately 52 minutes to process. This next mention is after at least half of the bulk of the article. It may be an assuming-spherical-cows estimation, but it doesn't strike me as grossly misleading on its face, and there's no editorializing. > This gets us up to approximately 77 times faster than the Hadoop implementation. > about 174 times faster than the Hadoop implementation. > gets us down to a runtime of about 12 seconds, or about 270MB/sec, which is around 235 times faster than the Hadoop implementation. The next three mentions, near the end, are just comparisons of the evolving demonstration implementation to the reference implementation. No detectable snideness. > Hopefully this has illustrated some points about using and abusing tools like Hadoop > but more often than not these days I see Hadoop used Here in the conclusion paragraph is where I agree that there is both snideness and where a reader may be confused about tradeoffs, if that's what they were expecting to be enlightened about. However, because that's not what the introduction promised, my criticism would be merely that the conclusion doesn't match the introduction (and maybe goes too far into inflammatory territory with "abuses"). Pretend that section isn't even in the article, and the article still reads OK.
- pmorici 8y agoThat's kind of the point of that post it's pointing out that you should consider weather you actually need to use something like Hadoop and that most people aren't actually working with data sets large enough for it to make sense.
- potatoyogurt 8y ago> Although Tom was doing the project for fun, often people use Hadoop and other so-called Big Data™ tools for real-world processing and analysis jobs that can be done faster with simpler tools and different techniques. "Big Data™" is unnecessary shade. > One especially under-used approach for data processing is using standard shell tools and commands. The benefits of this approach can be massive, since creating a data pipeline out of shell commands means that all the processing steps can be done in parallel. This is basically like having your own Storm cluster on your local machine. It is entirely unlike having a Storm cluster on your machine, and trying to do your data processing as chained shell commands will rapidly become cumbersome if you try to do actual complex processing. Yes, I get that the author is trying to point out that simpler tools can work for many cases, but the tone of the article makes it seem like that author is just generally saying that EMR/Hadoop is bad. He does not acknowledge just how weighted the test he did is against Hadoop or give any indication of what the tipping point is where you actually want to start considering something distributed. This paints a really misleading picture for anyone who does not already know a fair amount about these technologies.
- qaq 8y agoIt's actually an interesting question you can get 192 cores and 12 TB RAM in a single x86 box. At what point does it actually make sense to go for Hadoop.
- obelix_ 8y agoIt makes sense if you are handling millions of requests from all over the world per second and need failovers if machines go down. But...if you just want to run your own personal search engine say... Then for Wikipedia/Stackoverflow/Quora size datasets (50GB with with 10GB worth of every kind of index(regex/geo/full text etc) ) you can run real time indexing on live updates with all the advanced search options you see under their "advanced search" page one any random Dell or HP Desktop with about 6-8GB of RAM. Lots of people do this on Wall Street. People don't get what is possible on desktop cause so much of it has moved to the cloud. It will come back to desktop imho.
- nl 8y agoIt won't come back to desktops because as they get cheaper the costs of people to maintain them increases. There will always be people who need them and use them but that proportion is going to keep decreasing (I'm somewhat sad about this, but the math is hard to argue with).
- obelix_ 8y agoJust temporary. Nature did not need to invent cloud computing to perform massive computation. The speed and data involved in computation going on in a cell or an ants brain show us how far we can still go on desktop. The breakthroughs will come. In the meantime (unless you are dealing with video) most text and image datasets out there that avg Joe needs can easily be stored/processed entirely locally thanks to cheap terabyte drives/multicore chips these days. People just haven't realized there isn't that much useable textual data OR that local computing doesn't require all the overhead of handling millions of requests a second. This is Google problem not an avg Joe problem that is being solved with cloud compute.