5 ms·
Hadoop didn't fail us. We failed Hadoop. The decline of Hadoop as a software category is Software Product Marketing 101: it did not identify pervasive killer u
by ktamura 10y ago
Hadoop didn't fail us. We failed Hadoop.
The decline of Hadoop as a software category is Software Product Marketing 101: it did not identify pervasive killer use cases critical to running . Yes, it's true that Hadoop was a revolutionary way to store and process massive datasets on commodity hardware, but what's the use case for that? If you are Visa/AMEX (fraud detection), Facebook/Google (various ML-based data products) and a few other types of companies with obvious applications of massive data processing, yes, Hadoop has been great.
But here's the thing: beyond a few such corner cases, it never found a use case that enterprise data warehouse couldn't handle.
Then came Redshift, then BigQuery, and now Snowflake (as a BigQuery on AWS, really). While there are some key technical differences between Redshift and BigQuery/Snowflake, they are all _much_ cheaper than the previous generation of data warehouses (Vertica, Netezza, Greenplum, etc.) The lower price meant greater access, and developers who previously couldn't imagine using data warehouses could finally spin one up with a credit card swipe.
Hadoop, too, took a lot of collateral damage because many developers realized that they didn't need much of Hadoop beyond SQL-on-Hadoop.
Redshift was a beautiful feat of product strategy and marketing: They just took what used to cost a lot and offered it for much less in an environment where developers already had a lot of data (AWS). This was much simpler to execute than what Hadoop had to do: introduce new technology, identify use cases, and finally compete with incumbent solutions.
We failed Hadoop (as you can see from www.cloudera.com, even Cloudera, the Hadoop company, hardly mentions Hadoop on its top page). Not the other way around.
- bsg75 10y ago> because many developers realized that they didn't need much of Hadoop beyond SQL-on-Hadoop I wonder how many of those are just SQL-on-HDFS (Drill, Spark, Presto) ?
- vgt 10y agoMinor correction - unless you consider Paraccel a part of Redshift (you probably should), BigQuery GA precedes Redshift release by around a year, and Dremel at least 6.