4 ms·
Ask HN: How to prevent duplicate content in a distributed webcrawler?
So, we've a distributed webcrawler and we are gathering images + 180 character context. Then display it on a page for users to search, filter, group. We are suffering from a large number of duplicate or slight variations of those images/text.