4 ms·
It seems that as the amount of data being produced continues to expand at an unprecedented rate, it will become essential to master Hadoop which seems to be the
by jackhoy 14y ago
It seems that as the amount of data being produced continues to expand at an unprecedented rate, it will become essential to master Hadoop which seems to be the gold standard for managing big data - would anyone disagree with this?
- chubot 14y agoWell, you will encounter data in many forms. Sometimes it will already have been "hadooped-down" by someone else, and you can analyze it on a single machine. Don't underestimate what a single machine can do these days, if you have say 16 cores and 32 GB of RAM. Or you can set up a system that will incrementally summarize the data, and then you could do smaller queries against those summaries. That is the goal of Storm AFAIK. I think that is better model for a lot of applications. The model of having your production systems save terabytes of raw data and then analyzing it in a big batch job leaves a lot to be desired. It works but it's not very flexible and has this latency problem. Hadoop is good in that it's the only open source solution I know of that can churn through hundreds of terabytes of data. But I wouldn't say it's a complete solution for "managing big data". It's part of one.
- paulgb 14y agoI wouldn't be surprised if Hadoop (or another Map/Reduce implementation) becomes a sort of "assembly language" for big data, with higher-level abstractions built on top. You can already start to see this happen with Hive and Pig.
- equark 14y agoWhile it is sometimes important to be able to manipulate large datasets, every single competition on Kaggle can fit in a laptop's memory.