5 ms·
We're seeing a lot of regret around sprawling Hadoop deployments, so this doesn't surprise me. Other Hadoop vendors (vendor?) pivoting to machine learning is a
by joehandzik 7y ago
We're seeing a lot of regret around sprawling Hadoop deployments, so this doesn't surprise me. Other Hadoop vendors (vendor?) pivoting to machine learning is a bandaid as the compute capabilities scale beyond HDFS's performance limitations. Look towards new-gen startups around NVME/NVMEOF (WekaIO, Excelero, E8, etc etc) to fill the void.
The question is going to be: will anyone provide an intelligent way to maintain compatibility with applications that expect to interface with HDFS rather than POSIX? It's a bit of a gap right now from what we see.
- joehandzik 7y agoNote that apparently the layoffs have already happened. Engineers are already interviewing elsewhere.
- capkutay 7y agoHadoop's whole appeal was it was cheap and scalable. Did it actually work without serious engineering teams maintaining each distribution? absolutely not. the hype was purely VC funded. fast forward 10 years and hadoop has basically been killed off by hosted storage services that are more expensive but 10000x easier to manage.
- dmix 7y agoI'm not really familiar with this industry. Are these Hadoop clusters hosted locally by these businesses? And MapR et all were providing the software/consulting to help manage them? Then I'm guessing better cloud hosted options came out offering similar capabilities. If that's the case was it really a big surprise that "cloud" hosting would eat any self-hosted platform's lunch?
- capkutay 7y agoyes mostly locally hosted back then. the idea was you'd have this massive distributed file system to hold all your data, and VC backed comopanies like Cloudera, MapR, Hortonworks promised everyone that they'd build the SQL layer, the data warehousing, BI, and all the other enterprise features you'd need to basically replace your expensive Teradata, Oracle Exadata, and other data warehousing systems. But of course it didn't play out that way.
- threeseed 7y agoActually it did play out that way. Most enterprises these days will have a massive distributed file system (S3 or equivalent) that has most of their data. And most will be running a decent percentage of data transformation jobs using Spark or some SQL layer e.g. Presto and then running BI tools like Tableau using Athena/Redshift Spectrum (or equivalent) as their SQL layer. It's just that this is all playing out on the cloud instead of on-premise with some vendor. But you have definitely been seeing the decline of the core enterprise data warehouse.
- threeseed 7y agoSo in the very recent past (i.e. a few years) businesses wanting to do Data Science ran a distro of Hadoop e.g. MapR that they bought from the vendor. They charged an exorbitant charge per node (e.g. $10k) because they figured they would get people to switch from Teradata or Oracle. Now what the cloud offered was so much more compelling. You had per hour pricing on the order of $20 for a minimal cluster. You had unlimited autoscaling so you didn't have to do capacity planning and go through procurement processes to pre-order hardware/licenses. And of course you had unlimited, ultra-cheap storage courtesy of S3. And it also allowed each team to have their own mini-cluster instead of everyone relying on some giant one. I don't think the recent explosion in Data Science would've happened without the cloud.
- disgruntledphd2 7y agoWhich is hilarious, given how expensive it is to run analytics on AWS. My current company looked into recreating our internal analytics environment on the cloud, and realised that it wouldn't be cost effective.
- notyourday 7y agoYes, MapR was local hosting.
- threeseed 7y agoDo you actually work in this industry ? Because Hadoop has not been killed off. In fact it is far bigger than it has ever been and growing each year. It's just that it is all transitioning to the cloud with parts like HDFS being replaced with more scalable solutions like S3/EMRFS. But you can't say it's been killed off when on AWS you have managed Hadoop (EMR), managed Hadoop pipelines (DataPipeline) managed Spark (Glue ETL), managed Hive Metastore (Glue Catalog) etc. And similar on Azure or GCP.
- capkutay 7y ago> Because Hadoop has not been killed off. > HDFS being replaced with more scalable solutions like S3/EMRFS. Thank you for re-affirming my point. I'm not going to go into the typical HN "I'm going to nitpick your point down to the bone and argue you over the pointless details", but 2012 Hadoop is not the same thing as the tools you're describing.
- parasubvert 7y agoExcept you’re not being just a bit inaccurate, you’re flat wrong. Simply put, Hadoop is not HDFS, it is a much larger Apache ecosystem that includes Spark, Hive, MR, NiFi, etc. All of which are doing fine.
- kristjansson 7y agoI’m not sure how common that usage is . ‘Hadoop Ecosystem’ surely encompasses Spark and the rest, but I’d argue ‘Hadoop’ most commonly refers to MR, HDFS, and maybe YARN. Which is only to say I’d be surprised to hear an application described as ‘using Hadoop’ and find it using Spark on data in S3
- capkutay 7y agoSpark is NOT hadoop. You don't even need hadoop for spark anymore. Spark is more or less completely independent from the hadoop ecosystem. Just because they're compatible doesn't mean they're synonymous. You can use spark with mesos or kubernetes instead of YARN.
- threeseed 7y agoSo I deal with hundreds of Data Scientists and none of this is true for us. Hadoop vendors are simply getting killed by the cloud. AWS EMR, Azure ML, Google Dataproc. And many people are simply swapping HDFS for S3 or similar object stores. Never seen anyone switching to very expensive, more HPC style NVME setups. That seems strange given the data sizes we are working with. And 95% of what Data Scientists do is ETL, Data Curation, Feature Engineering etc and which Spark is still the overwhelming dominant player. And Spark uses Hadoop filesystem API underneath which is trivial to replace hdfs:// with s3://
- joehandzik 7y agoIt really depends on what the workload is and if performance or scale are priorities. Our team and the products we build come from the HPC side of the fence and we've been engaged with all sorts of large businesses that know their performance requirements (think 100s of GB/s read bandwidth), and it leads them down the path of either HPC filesystems or newer NVMe-oriented filesystems. Does 'AI' or 'data science' always mean 'high performance'? No, but it definitely can. Everyone's data volume is different as well...working a proposal for several petabytes of NVMe storage right now. It's not everyone, but use cases are out there for large volumes of high performance storage. Fair point on the ability to swap the hdfs interface trivially, I hope it's that simple everywhere. Another issue we run into is that these companies have invested heavily in these Hadoop clusters, and would prefer to continue to get some use out of the mountains of HDDs that are captive in these environments. So tiering/HSM functionality is another facet of the issue that these environments will anchor for quite some time.
- jamesblonde 7y agoThe problem you have with Weko and those newer filesystems is that the tooling hasn't caught up. They are trying to get their APIs upstreamed into TensorFlow. But what about Spark, Pandas, Arrow, etc? It's not enough to have just the training part of the pipeline - you need to be the filesyste for the whole pipeline. That's what we provide with HopsFS.
- saganus 7y ago
- busterarm 7y agoI literally can't remember how many years it has been since the last time I heard something positive about Hadoop. It's at least 4.
- jamesblonde 7y agoWe build a next-gen version of HDFS (HopsFS) that has distributed, consistent, transactional metadata where small files (<1MB) can be stored in NVMe disks in the metadata layer. And it's open-source - cf WekaIO, etc. And it's performance is backed by peer-reviewed publications at tier-1 conferences (Usenix fast, ACM middleware, etc). https://www.logicalclocks.com/millions-and-millions-of-files-deep-learning-at-scale-with-hopsfs/ https://www.logicalclocks.com/millions-and-millions-of-files... Our business model is to build a data science platform, Hopsworks, around our distributed metadata layer. And yes we use YARN (training models) and also Kubernetes (serving models). The choice of resource mgr is really just an implementation detail, as the platform is backed by a REST API.
- joehandzik 7y agoVery interesting! I did not know that this existed. I know some people who will be interested.
- jamesblonde 7y agoSome background reading/viewing: https://www.logicalclocks.com/eventscustom/ https://www.logicalclocks.com/eventscustom/
- jamesblonde 7y agoAnd it works in the cloud - HA and high throughput. On Spotify's Hadoop workload, we got 1.6 million file system ops/secs, highly available over 3 availability zones. On GCE.