4 ms·
Simhashing (Hopefully) Made Simple (2012)
- deleted 5y ago[deleted]
- Nzen 5y agoSimhashing is a style of characterizing the similarity of data. The author begins with the idea that we can discard the first characters of { aaarock, aabjeep, aaareep } to prefer the latter two as most similar and concludes with computing the hamming distance of data.
- jonathankoren 5y agoI wrote something similar several years ago on minhashing for near duplicate detection. https://medium.com/@jonathankoren/near-duplicate-detection-b6694e807f7a https://medium.com/@jonathankoren/near-duplicate-detection-b...