7 ms·
IA stopped respecting robots.txt in 2017. I had to issue a DCMA takedown to get my sites delisted. They are arrogant and think themselves above normal netiquett
by kwo 4y ago
IA stopped respecting robots.txt in 2017. I had to issue a DCMA takedown to get my sites delisted. They are arrogant and think themselves above normal netiquette. They deserve to lose.
https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
- tech234a 4y agoA relevant (yet incomplete) list of exclusions: https://wiki.archiveteam.org/index.php/List_of_websites_excluded_from_the_Wayback_Machine https://wiki.archiveteam.org/index.php/List_of_websites_excl...
- account42 3y agoMore as a hall of shame list than anything. And of course it contains websites that themselves are all about users sharing content without regards to copyright or even attribution.
- account42 3y agoIf anyone is arrogant it's those that think they can make information publicly available but then still control where it is shared. You don't get to define "netiquette" to whatever you want it to be. robots.txt in particular has been a giant mistake that only seves to entrench Google. Anyone that wants to archive or index the web as it appears to humans SHOULD ignore it.