4 ms·
Naive Bayes is also surprisingly powerful for deduplication of large datasets of messy data (and the related problem of record linkage). It's still competitive
by RobinL 3y ago
Naive Bayes is also surprisingly powerful for deduplication of large datasets of messy data (and the related problem of record linkage).
It's still competitive with cutting edge approaches if carefully applied, and usually much faster so it can be applied to big data.
It's the engine behind Splink, the free Python library that I develop:
https://github.com/moj-analytical-services/splink https://github.com/moj-analytical-services/splink
- greazy 3y agoThis is a very cool tool. I was testing fastlink recently but found some issues with the docs and implementation. I will check out your tool. Congrats on winning all those awards.
- RobinL 3y agoThanks, appreciate it. Fastlink was the inspiration for Splink - the fundamental statistical model is very similar. The first version of Splink was essentially a port to make it work faster and at greater scale, but we've subsequently added quite a bit of additional functionality Feel free to ask questions if you run into any issues - we're usually fairly good at responding: https://github.com/moj-analytical-services/splink/discussions https://github.com/moj-analytical-services/splink/discussion...
- reuben364 3y agoI had a friend that had a record linkage problem which interested me and so I dove down a rabbit-hole of Markov logic, but didn't see any mention of Naive Bayes approach when doing so. Although I don't really have any expertise in the area anyway, so it would be easier for me to miss.
- RobinL 3y agoThe model is usually called the Fellegi Sunter model, but once you get into the maths, it's actually the same as Naïve Bayes. If you're interested I've got a blog post that explains this here: https://www.robinlinacre.com/maths_of_fellegi_sunter/ https://www.robinlinacre.com/maths_of_fellegi_sunter/
- rkwz 3y agoThanks for this, been meaning to read up on the Fellegi Sunter paper for a long time!
- rkwz 3y ago> Naive Bayes is also surprisingly powerful for deduplication of large datasets Thanks for this! I implemented Entity Resolution from scratch 2 years back and was mostly able to do deterministic linkages using identifiers or using TF-IDF on names as a suggestion for human intervention. Will be writing about my learnings in upcoming months [1]. Was mostly looking at research papers before but will take a look at your project - your documentation looks good! [1] https://www.sheshbabu.com/posts/entity-resolution-challenges/ https://www.sheshbabu.com/posts/entity-resolution-challenges...