3 ms·
I need to take another look at this. I've recently started working with RDF data on a scale of ~50 million URIs not including properties and statements, there's
by arrowleaf 3y ago
I need to take another look at this. I've recently started working with RDF data on a scale of ~50 million URIs not including properties and statements, there's a ton of suplicates in there. I loaded a 10k entity subset of this and OpenRefine found all of the duplicates I had manually found plus others that I guess were 'similar'. Really cool, but it crashed when attempting to merge entities together. I've got a pipeline transforming the original JSON dataset to RDF, maybe it would work better working with the looser structure. What scale of CSV data did you have?