4 ms·
You have a point, but it's also a fact that lots of people use "Hadoop" interchangeably with "MapReduce". And Spark can in fact replace the Hadoop infrastructur
by rs_atl 12y ago
You have a point, but it's also a fact that lots of people use "Hadoop" interchangeably with "MapReduce". And Spark can in fact replace the Hadoop infrastructure entirely, as it's not a component of that ecosystem. Just because the various Hadoop vendors also support Spark only validates the point that there's a need to "go beyond Hadoop".
- monstrado 12y agoJust because people use Hadoop and MapReduce interchangeably doesn't make it correct. I would love to hear how you think Spark can replace Hadoop, because that is an astonishingly inaccurate statement. Which part of Spark reliably distributes data? Which part of Spark handles enterprise level security? Which part of Spark can coordinate resources in multi-tenant environments? The answer is none of them, it relies on Hadoop for that.
- x0x0 12y agoum, you sound like a vendor. hadoop does mean map-reduce + hdfs as the common usage by the majority of devs + admins. Claiming hadoop is now some distribution of tools is fine, but that's simply not what the common usage is. It remains to be seen if yarn will carry the day or no; my suspicion is that many people are essentially going to be running spark on hdfs. I don't see much use for yarn unless you need to balance hadoop and yarn, and weren't their claims that yarn was going to support eg mpi style computation that didn't pan out?
- colin_mccabe 12y agoum, you sound like a vendor. hadoop does mean map-reduce + hdfs as the common usage by the majority of devs + admins. Claiming hadoop is now some distribution of tools is fine, but that's simply not what the common usage is. I am a Hadoop developer, and I can tell you that Hadoop does not mean "map-reduce + hdfs". That's also not what people are installing when they install Cloudera's distribution of Hadoop, Hortonworks' distribution of Hadoop, or even Intel's distribution of Hadoop (which is being discontinued in favor of adopting Cloudera's). This is more old information from 2008, being replayed as current. YARN even lives in the Hadoop source code repository, it's hard to get more "Hadoop" than that. Spark has its own repo, but it uses many classes from Hadoop like InputFormat, etc. It remains to be seen if yarn will carry the day or no; my suspicion is that many people are essentially going to be running spark on hdfs. I don't see much use for yarn unless you need to balance hadoop and yarn, and weren't their claims that yarn was going to support eg mpi style computation that didn't pan out? Much confusion. Much sadness. You run YARN (or its close competitor, Mesos) because you want to have multiple jobs going on in the same cluster at once. You need things like per-user queues, job control, reserving CPU and memory resources. The jobs going on at once may be multiple MapReduce jobs, or they may be multiple Spark jobs. Even Databricks, which employs many of the early Spark developers, doesn't ship a product that runs Spark in standalone mode. They run on Mesos.