4 ms·
Are you running into those autogen crawler labyrinths that have been posted about recently? Certainly the .edu and .gov ones are probably just huge sites. Am I
by outer_web 2y ago
Are you running into those autogen crawler labyrinths that have been posted about recently? Certainly the .edu and .gov ones are probably just huge sites.
Am I correct in assuming you are recrawling your entire corpus to get fresh results? Would there be a downside to replacing crawl epochs with a continuous crawl that is random but with age priority?
- marginalia_nu 2y agoCrawler labyrinths generally tend to be on paths disallowed by robots.txt, or behind nofollow links. I haven't seen any indication that they're playing much part in this. Seems the bigger problem is very large domains with slow response times and long crawl delays. > Am I correct in assuming you are recrawling your entire corpus to get fresh results? For known links, I'm sampling them and based on whether I find changes (first via if-none-match and if-modified-since or alternatively via locality sensitive hashing), I'm recrawling only a part or all of the links. > Would there be a downside to replacing crawl epochs with a continuous crawl that is random but with age priority? The drawbacks to this is a much more mutable crawl data (being able to read or write crawl data top down in an append-only format is a huge performance improvement), as well as problems with the indexing software which takes ~1 day to complete, and can't currently be partially rebuilt but is rebuilt from scratch every time at significant computational expense.