4 ms·
Hadoop's whole appeal was it was cheap and scalable. Did it actually work without serious engineering teams maintaining each distribution? absolutely not. the h
by capkutay 7y ago
Hadoop's whole appeal was it was cheap and scalable. Did it actually work without serious engineering teams maintaining each distribution? absolutely not. the hype was purely VC funded.
fast forward 10 years and hadoop has basically been killed off by hosted storage services that are more expensive but 10000x easier to manage.
- dmix 7y agoI'm not really familiar with this industry. Are these Hadoop clusters hosted locally by these businesses? And MapR et all were providing the software/consulting to help manage them? Then I'm guessing better cloud hosted options came out offering similar capabilities. If that's the case was it really a big surprise that "cloud" hosting would eat any self-hosted platform's lunch?
- capkutay 7y agoyes mostly locally hosted back then. the idea was you'd have this massive distributed file system to hold all your data, and VC backed comopanies like Cloudera, MapR, Hortonworks promised everyone that they'd build the SQL layer, the data warehousing, BI, and all the other enterprise features you'd need to basically replace your expensive Teradata, Oracle Exadata, and other data warehousing systems. But of course it didn't play out that way.
- threeseed 7y agoActually it did play out that way. Most enterprises these days will have a massive distributed file system (S3 or equivalent) that has most of their data. And most will be running a decent percentage of data transformation jobs using Spark or some SQL layer e.g. Presto and then running BI tools like Tableau using Athena/Redshift Spectrum (or equivalent) as their SQL layer. It's just that this is all playing out on the cloud instead of on-premise with some vendor. But you have definitely been seeing the decline of the core enterprise data warehouse.
- threeseed 7y agoSo in the very recent past (i.e. a few years) businesses wanting to do Data Science ran a distro of Hadoop e.g. MapR that they bought from the vendor. They charged an exorbitant charge per node (e.g. $10k) because they figured they would get people to switch from Teradata or Oracle. Now what the cloud offered was so much more compelling. You had per hour pricing on the order of $20 for a minimal cluster. You had unlimited autoscaling so you didn't have to do capacity planning and go through procurement processes to pre-order hardware/licenses. And of course you had unlimited, ultra-cheap storage courtesy of S3. And it also allowed each team to have their own mini-cluster instead of everyone relying on some giant one. I don't think the recent explosion in Data Science would've happened without the cloud.
- disgruntledphd2 7y agoWhich is hilarious, given how expensive it is to run analytics on AWS. My current company looked into recreating our internal analytics environment on the cloud, and realised that it wouldn't be cost effective.
- notyourday 7y agoYes, MapR was local hosting.
- threeseed 7y agoDo you actually work in this industry ? Because Hadoop has not been killed off. In fact it is far bigger than it has ever been and growing each year. It's just that it is all transitioning to the cloud with parts like HDFS being replaced with more scalable solutions like S3/EMRFS. But you can't say it's been killed off when on AWS you have managed Hadoop (EMR), managed Hadoop pipelines (DataPipeline) managed Spark (Glue ETL), managed Hive Metastore (Glue Catalog) etc. And similar on Azure or GCP.
- capkutay 7y ago> Because Hadoop has not been killed off. > HDFS being replaced with more scalable solutions like S3/EMRFS. Thank you for re-affirming my point. I'm not going to go into the typical HN "I'm going to nitpick your point down to the bone and argue you over the pointless details", but 2012 Hadoop is not the same thing as the tools you're describing.
- parasubvert 7y agoExcept you’re not being just a bit inaccurate, you’re flat wrong. Simply put, Hadoop is not HDFS, it is a much larger Apache ecosystem that includes Spark, Hive, MR, NiFi, etc. All of which are doing fine.
- kristjansson 7y agoI’m not sure how common that usage is . ‘Hadoop Ecosystem’ surely encompasses Spark and the rest, but I’d argue ‘Hadoop’ most commonly refers to MR, HDFS, and maybe YARN. Which is only to say I’d be surprised to hear an application described as ‘using Hadoop’ and find it using Spark on data in S3
- capkutay 7y agoSpark is NOT hadoop. You don't even need hadoop for spark anymore. Spark is more or less completely independent from the hadoop ecosystem. Just because they're compatible doesn't mean they're synonymous. You can use spark with mesos or kubernetes instead of YARN.