7 ms·
Do you have any tips on how to convince management they don't have big data? My company's largest database has been in production for 10 years and could fit int
by badthingfactory 9y ago
Do you have any tips on how to convince management they don't have big data? My company's largest database has been in production for 10 years and could fit into the RAM on my dev machine, yet I'm constantly pushing back against "big data" stacks. This post made me laugh because "big data" and Hadoop were mentioned in our standup meeting yesterday morning.
- pps43 9y agoA good definition of "big data" is "data that won't fit on one machine". Corresponding rule of thumb is that you don't need big data tools unless you have big data.
- cookiecaper 9y agoI don't think that's a good definition. "One machine" of data is highly variable; everyone has a different impression of the size of "one machine". Does "fit" mean fit in memory or on disk? Why is "Big Data" automatically superior to sharding with a traditional RDBMS, or a clustered document database?
- aurelianito 9y agoI usually change "one machine" with "my notebook" and "fit" with "can analyze". It is big data if I cannot analyze it using my notebook. So it depends both on the size of the data (a petabyte is big data), the performance requirements (10GB/s is big data even if I keep 1 minute of data in the system) and also depends on the kind of analysis (doing TSP on a 1000000-node graph is big data, even if it fits my notebook memory). I also define "small data" as anything that can be analyzed using Excel. It is usually dramatic enough to get buy-in :).
- marcosdumay 9y agoIt's the best definition. It makes "big data" the name of the problem you have when your data can not be worked in a coherent way, what must be solved by distributed tools. If you can buy a bigger machine, you can make "big data" bigger, and maybe evade this problem; if you must access it a lot of times, fitting on disk is useless and "big data" just got smaller; etc.
- vonmoltke 9y agoHow then does "big data" differ from traditional HPC and mainframe processing? Those fields have been dealing with distributed processing and data storage measured in racks for decades.
- marcosdumay 9y agoI dunno. Is it useful to separate them?
- acdha 9y agoI think the simplest answer is that it's often essentially the same thing but approached from a different direction by different people with different marketing terms. One area which might be a more interesting difference to talk about might be flexibility/stability. A lot of the classic big iron work involved doing the same thing on a large scale for long periods of time whereas it seems like the modern big data crowd might be doing more ad hoc analysis, but I'm not sure that's really enough different to warrant a new term.
- speedplane 9y agoIf you are sharding across machines with a traditional RDMS, I think that would qualify as big data. Once you start dealing with multiple computers, complexity goes way up, because you've added a very large point of failure: the network.
- dsacco 9y agoEhhhh....that sounds like you're defining big data as distributed data. Hadoop and Cassandra lend themselves to distributed nodes, but you can also use them without that. Or you can use solutions that work well with "big data" that aren't as opinionated about it, such as HDF5. I guess the point is this: if I have 20TB of timeseries data on a single machine, and I have 20GB incoming each day, do I get to say I'm working with "big data" yet? EDIT: My other complaint with this definition (perspective, really) is that it predisposes you to choose distributed solutions when you really do have "big data", which is not ideal for all workflows.
- xyzzyz 9y agoIf you have 20TB of data on a single machine, you're better off with just Postgres 90% of the time. If you predict you're going have more data to fit on a single machine by the end of the year, then it makes sense to invest in distributed systems.
- dsacco 9y agoOkay, why? If you can sell me on that I'd be eager to change my workflow. For reference - this is entirely timeseries financial data and PyTables. For basically everything else I use postgres.
- dintech 9y agoDefinitely stick with HDF5 and Python for what you're doing. Postgress doesn't lend itself well to timeseries joins and queries in the same way that a more time series specific database like KDB+ would. The end result is most likely that you'd be bringing the data from a database into python anyway, probably caching in HDF5 before using using whatever python libs you want to use. You could alternatively bring your code/logic to the data using Q in KDB+, but there will be a learning curve and you will have to code for yourself a lot of functionality that just isn't available in library form. The performance will be a lot better though.
- pps43 9y agoThere are many examples of small data that is distributed, so distributed data is not necessarily big. But big data has to be distributed simply because there is no single computer big enough to hold all of it.
- alexchamberlain 9y agoI like that as a rule of thumb; I think the following tend to bucket storage and processing solutions well enough to be a starting point Small data: fits in memory Medium data: fits on disk Big data: fits on multiple disks I've yet to come up with a rule of thumb for throughput though, and this can never replace the expertise of an experienced, domain knowledgeable engineering team. As always, there are lots of things to balance, including cost, time to implement and the now Vs the near future. Rules of thumb over simplify, but also give you a way to discuss different solutions without coming over as one size fits all.
- crypto5 9y agoIt depends. You can store 4TB on single HDD, but reading and processing it can take many hours, so you may want big data stack to have your task paralleled.
- nilkn 9y agoI've ended up using "big data" tools like Spark for only 32GB of (compressed) data before, because those 32GB represented 250 million records that I needed to use to train 50+ different machine learning models in parallel. For that particular task I used Spark in standalone mode on a single node with 40 cores, so I don't consider it Big Data. But I think it does illustrate that you don't have to have a massive dataset to benefit from some of these tools -- and you don't even need to have a cluster. I think Spark is a bit unique in the "big data" toolset, though, in that it's far more flexible than most big data tools, far more performant, solves a fairly wide variety of problems (including streaming and ML), and the overhead of setting it up on a single node is very low and yet it can still be useful due to the amount of parallelism it offers. It's also a beast at working with Apache Parquet format.
- plafl 9y agoSame here. Some problems in ML are embarrasingly parallel like cross validation and some ensemble methods. I would love to see better support for Spark in scikit-learn and better python deployment to cluster nodes also.
- jimbokun 9y agoForward this article?
- cookiecaper 9y agoIt's very hard because they see "big data" as something that important people and important companies do. When you say "We don't have big data!", it translates to "We aren't that important!" This, of course, makes everyone very angry. Be mindful of people looking to introduce big data without justification. They are playing a game of some sort (maybe just personal resume value, or maybe a larger vie for power), and you are positioning yourself as their opponent when you try to stop the proposal they're pushing. Do not go into this naively.
- hinkley 9y agoThis isn't that different than all of the 24 year old programmers with the BMW screensaver. Successful people have BMWs. If I buy a BMW that means I'm successful. No, son, it doesn't.
- kwillets 9y agoOne thing is to create a standard benchmark for your current solution, eg a dataset and some standard queries, and run it occasionally. When they propose a "better" solution, point them at the benchmark and wish them well. This will achieve the two goals of measuring raw performance and keeping them out of your hair.