5 ms·
Four or five years ago, this was a tool I was using almost every day for work. Doing data consolidation and migrations for small nonprofits, we were faced with
by cdcarter 3y ago
Four or five years ago, this was a tool I was using almost every day for work. Doing data consolidation and migrations for small nonprofits, we were faced with so many loosely structured excel sheets and CSV exports from various mailing programs. OpenRefine was absolutely instrumental in cleaning up lots of disparate data when the data sources were too many and too variable to make a scripted solution valuable. Glad to see it lives on.
- yawnxyz 3y agoWhat tool did you move on to using instead? This tool seems super powerful!
- a5seo 3y agoCan’t speak for OP but I moved to Exploratory.io. And the beauty of it is, it’s a GUI for R so you can export your transformation steps to R if needed.
- layman51 3y agoI have been working at a nonprofit and have only recently started using this for cleaning up Excel or CSV files that we want to import. I am not as familiar with doing this with code, but I love that this tool gives me the steps I have taken in case I ever want to audit the changes I made to the data. The one disadvantage I see is that it seems like it’s only for a single user and it might be burdensome to collaborate since you have to share the project file. I’m still excited to learn more about OpenRefine, but I guess maybe something like Google Colab might be better in terms of sharing and having direct access to our G Drives.
- arrowleaf 3y agoI need to take another look at this. I've recently started working with RDF data on a scale of ~50 million URIs not including properties and statements, there's a ton of suplicates in there. I loaded a 10k entity subset of this and OpenRefine found all of the duplicates I had manually found plus others that I guess were 'similar'. Really cool, but it crashed when attempting to merge entities together. I've got a pipeline transforming the original JSON dataset to RDF, maybe it would work better working with the looser structure. What scale of CSV data did you have?