4 ms·
With modern NLP models being extremely good, can't you bundle together all sites whose content embeddings are very close? Presumably there's an acceptable thres
by _hl_ 4y ago
With modern NLP models being extremely good, can't you bundle together all sites whose content embeddings are very close? Presumably there's an acceptable threshold somewhere that achieves a good tradeoff of false positives vs negatives?
- dang 4y agoIt's possible. I haven't had time to look into it. If someone wanted to do a proof of concept it would certainly be interesting.