4 ms·
I’ve already caught their crawler ignoring robots.txt directives on one of my sites, aggressively indexing explicitly excluded information.
by hrbf 4y ago
I’ve already caught their crawler ignoring robots.txt directives on one of my sites, aggressively indexing explicitly excluded information.
- lizardactivist 4y agoOut of curiosity, what's the url for your website, and from what IP or host do their crawlers connect?
- hrbf 4y agoThe main connecting IP was 195.113.175.41.
- arjenpdevries 4y agoThat cannot be true, as the project has yet to start. But anyone can start a crawler, so you may have encountered other people's software. We wouldn't be so unknowledgeable to ignore robots.txt ;-)
- hrbf 4y agoIt was a crawler with the user agent "hgf AlphaXCrawl/0.1 (+https://www.fim.uni-passau.de/data-science/forschung/open-search https://www.fim.uni-passau.de/data-science/forschung/open-se...)", operated by the University Passau and Open Search Foundation named on your landing page. It would be a mighty big coincidence if this wasn't a project connected to this endeavor, especially when it confirms being an experimental crawler of said project at the UA URL.
- jaimex2 4y agoWouldn't it be impossible to know if it ignored robots.txt? Just because it crawled it doesn't mean it stored it.
- hrbf 4y agoStorage or not is entirely irrelevant to robots.txt directives. It guides automated access. It must be parsed first and excluded URLs must not be accessed at all.