3 ms·
you mean AI crawlers from Microsoft, owners of Github?
by knowitnone 1y ago
you mean AI crawlers from Microsoft, owners of Github?
- haiku2077 1y agoThe big companies tend to respect robots.txt. The problem is other, unscrupulous actors use fake user agents and residential IPs and don't respect robots.txt or act reasonably.
- internetter 1y agoBig companies have thrown robots.txt to the wind when it comes to their precious AI models.
- sph 1y agoYeah, they have openly disregarded copyright law, it's not a puny robots.txt file that's gonna stop them.
- haiku2077 1y agorobots.txt isn't just an on/off switch. You can set crawler rate limits in there that crawlers may choose to respect, and the big companies respect them- because it's in their interest to reduce their crawling cost and not send more requests than they need to. However, these smaller companies are doing ridiculous things like scraping the same site many thousands of times a day, far more often than the content of the sites change.
- PaulDavisThe1st 1y agoI have no idea where they are from. I'd surprised if MS is using a network of 1M+ residential IP addresses, but they've surprised me before ...