12 ms·
Out of curiosity, when does it effectively become "big data"? I ask not to be snarky, but it might be the case that it's "big data" to someone else, but not ne
by JonLim 13y ago
Out of curiosity, when does it effectively become "big data"?
I ask not to be snarky, but it might be the case that it's "big data" to someone else, but not necessarily to you. I figured it was a relative term for your industry/business, but the hacker crowd definitely seems to peg that amount in the millions of data points before calling it big data at all.
Seems fair, but I'd rather clarify.
- justincormack 13y agoBig Data used to mean petabytes, ie above the limits of performant scale up.
- ntoshev 13y agoI believe the accepted definition for big data is data you can't handle on a single machine and need a cluster to process. So Moore's law makes it a moving threshold.
- Malarkey73 13y agoI think Hadley Wickham has a decent description of big data in terms of the analytical process... to expand slightly on his description: On normal data you can iteratively explore and visualise it hitting return and seeing plots or model results instantaneously or at most a few seconds. When you have time to grab a coffee after hitting return then you have bigger data. If you carefully think through what you are about to ask the computer to do before pressing return then maybe you have big data. I actually think this is a better description than just size of files or data distributed across many computers as an algorithm that just streams over a massive dataset maybe in parallel can be less challenging than one that has to hold a much smaller e.g many Gb dataset fully in memory.
- RogerL 13y agoSo my complicated algorithm that processes 200,000 data points is big data because it takes 1/2 hour to run, but someone else's petabyte algorithm that takes 1 second on a cluster is small to "bigger" data? I don't think this makes sense. It's a measure of the size of the problem to be sure, but it is not a measure of the size of the data, or an indication of what techniques might be required to solve the problem.
- deleted 13y ago[deleted]
- Malarkey73 13y agoYes to a data analyst that is big data. If you are doing some MCMC and that is really what it takes on that size of data then you have a big data problem. The more sophisticated a statistic, the more high dimensional the data, the more sampling required, or the more of the dataset it requires to memorise at once - then the smaller your big data threshold will be. It depends a lot on your point of view too. If I google something now it may bounce across lots of crazy server farms but to me I don't feel like I'm doing big data.. the person who built it all probably feels differently.
- pedrosorio 13y agoI think defining "big data" as something that depends on the algorithm you are applying is not very useful. In that sense, almost anything is big data when you are trying to solve an NP-complete problem (nice course about algorithmic approaches to these problems: https://www.coursera.org/course/optimization https://www.coursera.org/course/optimization). The problem is, if you define "big data" as something that depends on the algorithm, then it makes no sense to include the word "data" in it. The expression "big data" as it is commonly used refers to flows of data so big that you need specialized approaches even when applying simple transformations to the data. Since the dawn of computing we've always wanted to solve problems, with small or large amounts of data, that required complex algorithms. The usage of a new expression is justified by the fact that huge flows of data are now available to many companies (mainly because of the web), not because these companies are attempting to perform extremely complicated transformations to the data. TL;DR: Big data means big volume/flow of data, and not (as you are defining it) using large Big-O complexity algorithms on some set of data. In fact, the size of big data precludes applying large Big-O complexity algorithms to said data.
- twic 13y agoI usually follow DevOps Borat's definition [1]: "Big Data is any thing which is crash Excel." Many a true word spoken in jest. [1] https://twitter.com/DEVOPS_BORAT/status/288698056470315008 https://twitter.com/DEVOPS_BORAT/status/288698056470315008
- oinksoft 13y ago"Small Data is when is fit in RAM. Big Data is when is crash because is not fit in RAM." https://twitter.com/DEVOPS_BORAT/status/299176203691098112 https://twitter.com/DEVOPS_BORAT/status/299176203691098112
- xroche 13y agoThis is very inaccurate/misleading IMHO. Big Data is something which does not fit in a regular machine for a given operation. You can sort billions of records on an iPhone, for example. You can grep a string within a terabyte-file data on a single personal computer, and I am not convinced you'd go faster with a distributed system (reading the file on cold storage will be the limiting factor). People claiming to do "big data" in these situations do not generally understand the underlying concepts.
- dannypgh 13y agoWith a distributed storage system you should be able to read said terabyte file using far more disk heads. It would also be easier to engineer it so the terabyte file was entirely in RAM by distributing it across multiple machines (although single machines with TB ram capacity are no doubt continuing to become more common) Sure, store it on a single tape or disk and distributing the computation won't help. You need distributed storage to properly leverage distributed computation for otherwise I/O bound processes.
- com2kid 13y agoYou underestimate Excel! You can point Excel to a DB table and use pivot tables on top. I know it can at least get up to several million, didn't have a chance to test it beyond that! :)
- 13y ago
- beejiu 13y agoWhen you are constrained to O(n) methods, you have big data.
- dj-wonk 13y agoSome people are constrained to O(WTF) methods and have no idea about O. So everything is Big Data.
- geocar 13y agoI prefer this definition to others involving the size of memories, or number of computers, because it underscores the data rate instead of just its' (instantaneous) volume.
- deleted 13y ago[deleted]
- entendre 13y agowhen you need statistical models and a tool for querying beyond the ken of an average sql dba
- ronaldx 13y agoI appreciate the definition of Big Data as requiring >1 machine. However... Small Data, that which traditional researchers handle, is normally much, much smaller than that: perhaps 10-10000 data points (and most often on the small end of that). An experienced researcher can essentially can know everything about this data set, including its outliers and quirky points, and get a good sense of it by drawing out simple graphs. There is clearly some disconnect between these two ideas: is that "Medium Data"? I would accept a concept of "Big Data" as data that cannot easily be eyeballed to get a sense of what's going on, so 10000+ points would count (under some circumstances). Maybe the concept of "six sigma" is useful - enough data that you would reasonably expect a six sigma outlier. Mathematically/statistically, the storage limit is not a particularly important milestone: the ideas and methods don't change once you reach this scale (except for potential parallelisation).
- gtrubetskoy 13y agoI was just going to post the ">1 machine" definition when I saw your comment. I think there are also at least two meanings of "Big Data". The more popular one is simply a trendy name for good old and boring "statistics", but with a twist that the data comes by way of the Internet, social media, all that. The second one (and a little closer to my heart) is what ">1 machine" means from a developer/sysadmin perspective. This is where the hadoops, hives, cassandras, etc. come into play, and it's A LOT to learn, even for seasoned developers. I think it's also a little intimidating for people who have become very comfortable with the typical rdbms stack. Parallel processing can be hard to understand, it's not something you can tinker with on your laptop over the weekend, and it's not surprising to hear all the "your big data thing is stupid" comments.
- frozenport 13y agoWhen you can't read all the data. ~(From a math professor I worked with)
- hnriot 13y agoBig Data is less about size and more of a characterization. Human generated data can never be big data - there's just not enough humans to make it all. Possibly with the exception of the biggest social networks. Big Data is machine generated by systems. Typically its logs, IoT etc.
- sam_sach 13y agoNot sure how you came to that conclusion. Humans generate voice data, which are then translated to a time series of frequencies for analysis. Comment data on this site alone would be a pretty big task to analyze.
- mark_integerdsv 13y agoFourteen years in IT, ten of them in BI... My definition is as follows: data pertaining to and generated by the source systems that govern various business processes are 'data' data (internal data owned by the business.) Data pertaining (in whatever abstract sense) to the business, generated by systems outside of the business are 'big'(external data.) Nothing to do with rowcounts directly IMHO.
- zmmmmm 13y agoI thought the definition in the article was actually really insightful: big data is when you start to behave as if you have N=all for a non-trivial sample.
- einhverfr 13y agoThe typical definition is where standard data management approaches do not work due to high volume, velocity, and/or variety of data sources. What are standard data management approaches? I don't know. Usually they mean single machine relational db's. But the thing is that once you get to a certain point on these three you need specialized solutions. High volumes of transactional data with real-time reporting might be handled well by something like Postgres-XC, but that won't handle data of sufficient variety. High velocity data may be best handled with something like VoltDB, but it can't handle volume. Etc....
- radmuzom 13y agoI think "big data" is a term characterized more by the analytical techniques you apply on them rather than the size of the data. Traditional inferential statistical techniques work on "small data", while newer Bayesian techniques work on "big data" - note that this does not imply that one cannot work on the other.
- sam_sach 13y agoNot necessarily. Try running basic descriptive statistics on terabyte scale data.