3 ms·
In my company, the biggest issue is finding the right data sets in company's vast data landscape, figuring out the exact definition/meaning of each column etc.
by cyberdrunk 6y ago
In my company, the biggest issue is finding the right data sets in company's vast data landscape, figuring out the exact definition/meaning of each column etc. Then, it's dealing with the data quality issues. Then, it's getting access to it and setting up an ingestion job for the data to be copied to some common storage (e.g. Hadoop). At the very end, it's the actual data science. I suspect a lot of the PhDs we hire start drinking before they reach the data science stage :)
- karishmakunder 6y agoYeah, that! Do you use any tool to centralise the data that you use across your Data Science teams? Like a central repository of sorts, so that each one can be given access to it and from there you can start off building your models for training etc?
- cyberdrunk 6y agoYes. We ingest all that data onto central Hadoop, where data science team can access all of it in an uniform way. This solves the physical access problem. Unfortunately, the DQ and meaning of data are harder to solve. They require essentially caretaking of the datasets done by the data owners (cannot be done by a centralized unit). My organization is currently undergoing a transition, where it will be a responsibility of the data owner to maintain the metadata of his/her dataset and also to measure the data quality, but implementing it across the whole org is a journey that will take a long time.