6 ms·
In my experience, developers' top 3 mistakes in this area are: 1. Convincing themselves they have "big data" instead of just "data". If it fits in RAM on your
by tedchs 12y ago
In my experience, developers' top 3 mistakes in this area are:
1. Convincing themselves they have "big data" instead of just "data". If it fits in RAM on your laptop, it's definitely not "big".
2. Thinking an example of "sophisticated analytics" is "Average over time with no standard deviation".
3. Resume Driven Development over pragmatic solutions that get stuff done.
- collyw 12y agoThis is so true. The problem I have though is that looking around for jobs recently, people want "experience in MongoDB", rather than "ability to recognize when MongoDB is clearly not the best tool for the job".
- tedchs 12y agoYou could always take some Art of War inspiration and learn MongoDB really well before your interview, then once you get in there become the in-house Mongo go-to expert, then use your clout to recommend a replacement. :)
- F_J_H 12y agoI've often wondered why people don't use databases and SQL for more often data analysis. SQL is relatively easy to learn, and there is so much you can do, especially if you use analytical functions. I almost always load the data into a DB, and do as much as I can with straight SQL. And, if I can't do it in SQL, I write a query to the transform the data to the format I need, and then access it via python to do the rest.
- gaius 12y agoThey do. In real companies (those with you know, revenue) millions of ordinary people called business analysts, accountants, statisticians use SQL or tools like Business Objects to do real work on multi-terabyte databases called Data Warehouses every day. Have done for years before the first hipster called himself a "data scientist".
- ThrustVectoring 12y agoThere's more you can do with straight SQL than you'd think. Say, multiplication of sparse matrices.
- seccess 12y agoYep. For things that aren't "big data" (nearly everything), I've started to use in-memory SQLite databases for analysis. Just loading your data into a table makes doing all the selections/aggregations pretty simple if you're familiar with SQL and it fits in RAM.
- zo1 12y agoAny particular reason you pick SQLite over more featured DB's such as MySQL or Postgres? I am assuming that they would be faster, with the only downside being more difficulty in the initial install/setup.
- jlarocco 12y agoFor a lot of "one off" tasks (like analyzing a big log file), just configuring a full blown DBMS like MySQL or Postgres would take longer than it would to write a script using an in memory SQLite DB and running the necessary queries. I'm not sure MySQL and Postgres are faster enough to justify the extra setup work when everything fits in memory. My intuition is that they may even be slower, because they still have to write everything to disk, where SQLite can be memory only.
- mbreese 12y agoIt's very easy to make a throwaway SQLite in-memory database for processing... not so much for a traditional RDBMS.
- bunderbunder 12y agoI'm beginning to think that "big data" is the most unfortunate possible term. By any reasonable definition of 'big' even a gigabyte or two counts as big. And since 32bit isn't quite one for the history books just yet, many of us still have 2-4GB fresh in our heads as a threshold where you might have to start thinking about the data needing special treatment.
- mfisher87 12y agoThere have been conversations in my job where I refer to a 40GB file as "big" and someone else will go "that's nothing." Well, obviously there are bigger files somewhere, but that doesn't mean 40GB is nothing. "Big" is always a relative term. Anything is big in the right perspective, so to say that something "isn't big" is almost always wrong. We should really try to find a term with a more absolute meaning.
- andybak 12y agoI think a rough definition might be something like "datasets big enough to cause problems with standard tools on standard hardware" - i.e. where you need to change your tools to get things done. Still pretty vague but it captures what I think many people are getting at.
- blumkvist 12y agoBig data does not refer only to the size of the dataset. It's variance and velocity, in addition to volume.
- jghn 12y agoI was going to say something similar except for me it's a function of size and multidimensionality. We're sitting on several PB but it's generally in two dimensions - standard analysis techniques work, we just have problem moving the data around efficiently and processing it. I like your inclusion of velocity, I'm including that in my definition.
- 12y ago
- jghn 12y agore 1 - but it sounds so sexy! I recently had a recruiter talking up a local shop that was on the "cutting edge of big data techniques". He also talked up that they had a dataset which was "several terabytes". We churn out many times that a day and I don't consider us to be "big data".
- tjradcliffe 12y agoMistake #0: not looking at the data. Whether you're designing a processing pipeline or simply keeping track of what one is doing, decent visualization of inputs and outputs is critical to understanding what's going on and ensuring that nothing surprising is happening. I'm always surprised when people come to me with data analysis problems (big or otherwise) and have never bothered to do any kind of visualization. Even if you're just visualizing sparse samples it's remarkable how easy it can be to spot issues. Good visualization won't solve all your problems but there is a significant sub-set that go from hard to easy when you do it.
- jackmaney 12y agoYep. Even looking at quartiles, min, mean, and max for each of your (numeric) variables can yield a bit of insight for prettymuch no effort.
- tmarthal 12y agoRight. Using pandas may allow you to process more data than an excel spreadsheet, but if the file is on your laptop then it's inherently not very big. It's a great tool, but it is a data exploration tool, not a data processing workflow tool. IMHO, there are too many 'data scientists' nowadays that are taking averages and calling it 'analytics'. If the call you are making exists in a library, more than likely it's not "sophisticated".
- nightski 12y agoWhile the engineer in my agrees, the business person says it really doesn't matter what techniques are used if he/she is providing value. If all it takes is some averaging to save a company a good chunk of money, then that person may be earning their pay.
- tomrod 12y agoAveraging the right things in the right way is important too--such as averages within clusters.
- _dark_matter_ 12y ago> If the call you are making exists in a library, more than likely it's not "sophisticated". What a disparaging comment. The whole point of good software design is many of the algorithms you may need to use are packaged up for ease of use. The best ones are highly specified by the parameters you supply. "Sophisticated" should play no role in a data scientists workday - results should be verifiable and understandable, and any data analysis pipelines should be extensible and repeatable. Writing their own deep-learning implementation does not a good data scientist make.
- mathattack 12y agoI see an awful lot of #3. Hard not to laugh when I see it. Sometimes I think that describes 3/4 of the things I see Data Scientists working on. The challenge in the field is very few folks know programming, math/stats, and the domain they're working on. It's rare to even get 2 of 3.
- jackmaney 12y agoRegarding #1, it took a phone call and two emails (the last of which was CCed to my boss) to convince someone in another silo of the company that yes, I do need all 3.5 million rows of a particular table and that no, he wouldn't crash the network drive if he dumped it there. If all that you use is Excel, then there are a lot of huge datasets out there.