6 ms·
Good to see that data cleaning was #1 on that list. Whenever I do work on a side project, it takes way way more time to get and structure the data than it does
by jackschultz 10y ago
Good to see that data cleaning was #1 on that list. Whenever I do work on a side project, it takes way way more time to get and structure the data than it does running the algorithms. Granted, that's because I have to go out and get the data in the first place, and then make sure it's useable and in the correct format.
Like the recent project I'm doing trying to classify country music songs based on their topic on the data blog I write on (https://bigishdata.com https://bigishdata.com), the amount of time it's taking to scrape lyrics, remove duplicate / incorrect songs, and then do manual classification for training data is taking far longer than running the ml algorithms in the end one I've gone through that process.
I've been looking for jobs recently, and I've seen only one job posting that mentions data cleaning as a necessity, whereas the rest only talk about data science and algorithm knowledge, or overall ETL design on the data engineering side. Seems like data set knowledge should be emphasized more.
- makmanalp 10y agoYes! Came here to write this too. I've been spending the last few years thinking, inquiring and talking about this, and I think there definitely is a field emerging here. There are already some companies trying to think about this stuff, and classes give it lip service but I don't think we've even seen the tip of the iceberg. It's interesting to talk to different people about data quality and what they think it means, or how they choose to deal with it, and it's all over the place. Some people just mean open and consistent formats, some people have stylistic preferences for data shape, some people talk about accuracy of values, etc etc. In some ways it's an extension of the thought that the world is inherently noisy, and we've been thinking about that one already, it's just that it turns out you don't need sensor data a la robotics to get noisy data - it's already in the datasets we know and love, and you accumulate more of it, the more sources you pull into your analysis.
- sixdimensional 10y agoData governance and master data management spring to mind as existing ideas along these lines.
- tomrod 10y agoAye, but the locked-down nature of data governance actually fights against using it for insights. I've seen a situation where a small two-way frequency table was requested by Team 1, and it took six _months_ because Team 2 had access to data but no access to metadata, Team 3 had metadata and could help design a query but didn't understand the technical details or run the query, Team 4 had to approve the process, and Team 5 had to review Team 3's query before it could run. In the mean time Team 2 was reorged and Team 1, the original requestors, found upstream sources. Data in a regulatory regime can be excruciatingly difficult, and lend itself to "gut instinct" being used because fear of risk and regulation lock things down too tight to be useful.
- sixdimensional 10y agoI couldn't agree more. Bureaucracy creates silos and stymies sharing and agility. Balance is needed.
- dasboth 10y ago> classes give it lip service I agree with this. It's fine to teach machine learning using the iris dataset, but there is rarely, if ever, a section dedicated to "real" problems. It was a shock to me just how high a percentage of time is spent cleaning data. It is a fundamental skill that is not only underestimated but "undertaught".
- _delirium 10y agoI agree, and would add: data cleaning's importance to the quality of the result is also often underemphasized compared to the much bigger focus on the quality of the algorithms. A single bad decision on data cleaning can have a large effect on the end result (in many cases, more than choosing between algorithms, assuming you pick some vaguely reasonable algorithm). Especially any choice that ends up producing non-random effects, like deduplicating things in a way that ends up biased: it's common that missed duplicates in an automatic deduplication process aren't randomly distributed. Or a scraping process that ends up with biased samples. You can correct for these kinds of things in various ways (e.g. incorporating an estimate of the bias in a statistical model), but people who don't consider data cleaning a "real" part of the whole statistical modeling pipeline in the first place usually don't.
- eanzenberg 10y agoAbsolutely, having data that improves signal-noise will always trump a better algorithm.
- elevensies 10y agoAny specific resources you'd recommend on data cleaning, verification, etcetera? I've just started reading this: https://www.amazon.com/Accuracy-Economic-Observations-Oskar-Morgenstern/dp/0691003513 https://www.amazon.com/Accuracy-Economic-Observations-Oskar-... . I've seen a few other books on the subject which I'm planning to get into, but I'd be interested if anyone has specific recommendations.
- jackschultz 10y agoHonestly, as another commenter pointed out, seems like an emerging field. Best practices and processes are just being figured out, and I haven't seen any great resources online talking about what to do, especially since most of what you need to do depends on the data set and how you're storing the data, and that can vary widely. Like I recently dealt with finding duplicate song lyrics in my 5000 set of lyrics, and to do that, I just had to google around for StackOverflow answers or random blog posts before I found something that I could adopt and chance for what I had.
- chubot 10y agoDo you use R and the "hadleyverse"? (Or "tidyverse" I think as he prefers?) I'm a programmer by trade but I use R because the people who actually work with data use it, and they write good tools for it... I think there is some confusion in the programming world about this. Programmers work with data, but they don't do it nearly as much as "professionals". Tidy data is a good intro if you're not familiar with it: http://vita.had.co.nz/papers/tidy-data.html http://vita.had.co.nz/papers/tidy-data.html And I would recommend going through other publications by Wickam, all on his site -- they are quite readable.
- elevensies 10y agoNo, I don't use R, thanks for the reference. Most of what I've done lately has been as part of software development process, so along the lines of validating the effectiveness different techniques for solving a known problem with a smallish test dataset -- typical engineering style optimization. I'm looking to impose more structure on the process.
- eanzenberg 10y ago>> I've been looking for jobs recently, and I've seen only one job posting that mentions data cleaning as a necessity, whereas the rest only talk about data science and algorithm knowledge, or overall ETL design on the data engineering side. Seems like data set knowledge should be emphasized more. Actual data cleaning, usually in an automated sense, is more 'data engineering' than 'data science' or applied statistics. Feature engineering and 'massaging' training data is more related to DS but it's understood that this data being consumed by the DS is already in decent shape.
- IanCal 10y agoI'd hesitate to call a lot of the work I do cleaning data "engineering". I think perhaps the problem here is the term science covers a lot of disciplines. I propose harder stats be data theoretical physics, with data biology and similar referring to cases with harder messy real world complications. I'm sure we can come up with a full spectrum.
- dasboth 10y agoThat also depends on the team. If your company has dedicated data engineers, great! Otherwise you're probably stuck with it.
- garysieling 10y agoI agree - I have a project that has similar problems (cataloging standalone lectures - https://findlectures.com https://findlectures.com). The biggest advantage of it being a side project is there's no pressure to get the data cleaning done, but in a work environment with time pressure this type of project is a huge pain.
- maverick_iceman 10y agoI don't get why companies would hire data scientists to do ETL jobs. These are properly left to engineers with expertise in data warehousing. From what I see though, this is a pretty common occurrence in Silicon Valley.
- mrhektor 10y agoAgreed! However, data cleaning is a pretty hard problem. My previous company stored merchant credit card transactions, and these transactions were large unwieldy beasts whose data model had changed many times over the course of the company. Old data was completely invalid, and yet we couldn't remove it because the dollars and cents had to add up. The cleaning significantly hindered new development. Validations when storing new data definitely help, but changes to the data model are tough to reconcile with old data.
- gerhardi 10y agoHundred times this. You see Qlik/Cognos Analytics/PowerBI/Alteryx/whatever sales guys making demos that make executives drool over the seeming easiness and wow-factor these tools are capable of producing. When the time comes to plug those over your production operative systems, CRM, whatelse, there comes "the now wait a minute" moment especially if your systems and their data models happen to be even slightly on the more complex side. Edit: The basis of succesful implementation of these tools is to have the data in digestible format and I feel that transforming the data to that business usable format is where the big job is. In my opinion well done ETL and DW are not going anywhere, even though in some circles they are said to be things of yesterday. Then there's a huge difference between an OK ETL/DW and a Brilliant ETL/DW. Designing a good ETL process is as large parts business and context knowledge as it is a application of data engineering skills. For example, it requires business knowledge AND data engineering knowledge to determine what kind of granular level advanced metrics could or should be calculated during ETL. Service level metrics and service level categorization for different kind of customers/claims/orders/... would be a perfect simple to understand example problem - there could be attributes and value ranges behind multiple relations that probably need to be taken into account and understood. Edit 2: I've been involved in both sales and execution of so called data discovery sprints, which are a 4-6 week periods where we bring a data engineer, a subject matter expert and client key personnel working together and let them go "fishing". The key thing is that this provides an low cost way for the clients to possibly gain insight on the potential their data could provide. On the other hand, many prospective clients just have so messy data that this data discovery job can't be recommended, which leads to other possible opportunities (MDM, ETL, DW).
- vram22 10y agoMDM is Master Data Management? (from a quick google). Hadn't heard of that specific term before. I'm interested in data projects.
- gerhardi 10y agoYes correct. In a way it's a combination of philosophy, agreed practices and techical solutions. I think this one is a good introduction: The What, Why, and How of Master Data Management https://msdn.microsoft.com/en-us/library/bb190163.aspx https://msdn.microsoft.com/en-us/library/bb190163.aspx
- mrweasel 10y agoI'm currently preparing a lecture on the topic of logging for my students. Part of the lecture is of cause how to use various logging frameworks, but the main part is what to log and how to structure logs. Basically we're trying to get them to pre-emptively do data cleaning, so their logs will actually be useful for potential future data projects.