4 ms·
Interesting idea to do two passes and build a merge list ahead of time. I am not sure if there's enough memory to build up such a list but it's worth a try. Ed
by nullbytesmatter 4y ago
Interesting idea to do two passes and build a merge list ahead of time. I am not sure if there's enough memory to build up such a list but it's worth a try.
Edit: I tried, and OOM'd by the OS. Apparently 64GB of RAM isn't enough to key off the unique column.
- Minor49er 4y agoIf you are finding that many matches that early, it might help to do multiple passes. On the first pass, identify up to X number of duplicates. Once you have X number of duplicates or reach the end of the file, iterate over the CSV and merge those rows. Then start again identifying a new set of duplicates, then apply those updates, over and over, until the file has no more duplicates