2 ms·
It seems to me that Twitter's robots.txt only allows Googlebot: https://twitter.com/robots.txt https://twitter.com/robots.txt Therefore, this disallows other
by nmlracx 3y ago
It seems to me that Twitter's robots.txt only allows Googlebot:
https://twitter.com/robots.txt https://twitter.com/robots.txt
Therefore, this disallows other bots in "maschinenlesbarer Form" and the scraping is illegal.
- mschuster91 3y agoThis is fascinating. Seems like at least Bing openly ignores robots.txt?
- gnfargbl 3y agoGoogle is, to a first approximation, the web's only search engine. As such I think there's an argument to be made that declaring your bot's User-Agent as Googlebot is a technical necessity, similar to how every browser declares itself as Mozilla.
- deleted 3y ago[deleted]