3 ms·
I've had a similar question before. I have heard of as few as 60,000 observations was considered "big data" [1], yet at my company, we generate about 60 million
by christopheraden 13y ago
I've had a similar question before. I have heard of as few as 60,000 observations was considered "big data" [1], yet at my company, we generate about 60 million pharmacy claims every 3 months, and no one here calls it big data. In terms of storage, it's on the order of a couple hundred terabytes for all our data. This is considered small enough that we can query it with traditional SQL.
Big Data, the experts say, is more about the novel way you analyze data, relative to the difficulty of the problem. A speaker at PyCon who was talking about algorithms and data structures for handling genetic data had a term that I like quite a bit better: "Data of Unusual Size" (C. Titus Brown at MSU was the speaker).
"Big Data" is a really big buzzword right now, but the term is overused and often does not convey the meaning it's supposed to. "Big" is a relative term. The novelty of how much data is being used as opposed to how much used to be used (in the sumo case, they had never handled so much data before) is what makes it big.
As for Hadoop, you'd want to use it when it's no longer feasible to keep your data stored in an RDBMS, or when speed becomes an issue, or when you want your schema to be more flexible than an RDBMS. If you are not concerned with the reliability of your data (RDBMS make the safety of the data a paramount priority--read the wikipedia page on ACID to see these guarantees), there's plenty of reasons for choosing Hadoop.
[1]: http://www.wired.com/wiredenterprise/2013/03/big-data/ http://www.wired.com/wiredenterprise/2013/03/big-data/
[2]: http://en.wikipedia.org/wiki/ACID http://en.wikipedia.org/wiki/ACID
[3]: http://hortonworks.com/blog/4-reasons-to-use-hadoop-for-data-science/ http://hortonworks.com/blog/4-reasons-to-use-hadoop-for-data...
- Joyfield 13y agoI would say that "a couple hundred terabytes" IS pretty big.
- christopheraden 13y agoThe types of queries we run against it don't require real-time results, and we do a pretty heavy amount of subsetting. By the time it reaches the point where we do numerical summaries and statistics, the largest set I've worked with here was around 30GB. Most times it's around 5-10GB.
- xtacy 13y ago5-10GB seems small for analysis. Is this sampled across your historical records, or just the most recent? What's the turnaround time for stats today, and what's your pain point? (Slow IO? Lack of high level programming frameworks? Or something else?)
- christopheraden 13y agoMost of the work I do involves recent data--cycles are six months at most and 3 months on average. 3 months of data, sifting by a pretty strict filter, it's not unsurprising that hundreds of terabytes of claims over years and years gets filtered into a few GB. Start to finish on jobs is a few hours (though it can be a few days if the filter is less strict or the task is more complicated), including pulling the data from the warehouse. Without a doubt, the bottleneck of the process is the data warehouse query. I'm sure having a more distributed database (it's DB2--I've been more pleased with Teradata's speed) could make the queries faster, but things are slow to change. Second to the query, the latency from working with a remote server (I'm in CA--the server's in Minnesota) adds latency if there's something I need to pull to the local machine. The actual SAS code (Is it still "high level" if the syntax models Fortran? I kid, I kid.) takes a negligible amount of time compared to the query.
- xtacy 13y agoThanks; the way of analysing data seems like an interesting view. I do understand the consistency tradeoffs between frameworks like Hadoop and that of a database, and that strong consistency guarantees are not required for many applications, but glad to know the range of data sizes we normally deal with. but yes, a few 100 TB does seem like a lot. :-) What kind of analysis do you do on pharmacy claims (~1.7MB per claim seems a bit high!)