3 ms·
That file includes at least two non-standard syntax extensions[0]. Robots is just a de facto standard and respect of some directives varies[1]. So much for it b
by FRex 9y ago
That file includes at least two non-standard syntax extensions[0]. Robots is just a de facto standard and respect of some directives varies[1]. So much for it being 'not difficult' while the task is not even clear because there isn't even a clear standard.
Archive.org also dislikes how robots.txt is being used mainly for search engines and goes against their mission in particular[2]. Are they now hackers for not throwing away information just because someone was overzealous with robots.txt or retired a certain website and uses robots.txt as SEO to let another one take its place in Google search results?
If some big corp wants to cry and bring legal matters into software they should first be accountable themselves for not securing themselves and the data of their clients (see the LinkedIn hack people mentioned elsewhere here and in general the high profile hacks like Equifax, Sony, etc.). Or should software shape up to be like many other areas today are - multi-million corporations are free to play fast and loose and endanger people while small guys get fried over meaningless bullshit and vaguely defined "crimes".
[0] - https://en.wikipedia.org/wiki/Robots_exclusion_standard#Nonstandard_extensions https://en.wikipedia.org/wiki/Robots_exclusion_standard#Nons...
[1] - https://intoli.com/blog/analyzing-one-million-robots-txt-files/ https://intoli.com/blog/analyzing-one-million-robots-txt-fil...
[2] - https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
- ikeboy 9y agoIt contains User-agent: * Disallow: / I am pretty sure none of the standard libraries/ tools that respect robots.txt would continue after being fed that file. >throwing away information This is entirely irrelevant. If they receive data from someone they have no obligation to discard it because of the current status of robots.txt. The question would be if they should continue to actively scrape that website. It seems like they've done that for gov sites, but nobody particularly cares about enforcing gov robots.txt. It would've been interesting if the government sued them, although if they cared they probably would've just told them to stop.
- FRex 9y agoSo we have an unclear "standard" that is only a de facto standard (and still varies in more advances directives between few big bots) that you're "pretty sure" about but that's seemingly not written down in its entirety anywhere and it'd also be enforced selectively depending on whether or not "someone particularly cares". Truly perfect and foolproof law that would be. And all this to protect some corp's business model of not letting others collect automatically the public information they provide, while they are free to use outdated or buggy software, store passwords in plaintext, etc. and get away with leaking data of millions of customers that should never be public. And it'd fail to stop anyone except benign, private and low fund actors because instantly Indian (or other low wage country) services for "scraping by human thus not a bot ignoring robots.txt" would pop up, just like there are captcha solving services that employ humans already, and malicious bots wouldn't care anyway just like they make 0 effort to respect it now and run from servers in some country that isn't friendly towards USA so there is 0 potential for catching the perpetrators.
- ikeboy 9y agoI disagree that laws that can only be enforced against US companies / people are worthless. Requiring a human would increase costs and it doesn't seem like a good argument against anything.
- FRex 9y agoBut they are in this case. They would not stop any scraped data from popping up for sale in shady places. That can be done by LinkedIn or whoever themselves using some smart way to detect bots and stop them from scraping their website. The only people a robots.txt law would affect are private users who set up a Python script to scrape a single page for themselves to check for something, things like archive.org, researchers, automated website testers, etc. while anyone nefarious can just rent a shady VPN or use a server in Russia, China, Middle East, etc. Requiring a human barely increases the cost if that data is so valuable in the first place and would be last resort anyway, far after just running the bots from a shady country, for captcha it's done because it's technically easier/cheaper (although supposedly automated solvers exist too). But laws that punish outright gross negligence would help protect everyone who uses these American websites (and most of the world does) from data leaks of data that is arguably way more sensitive (emails, unhashed passwords, SS and CC numbers, real names even like in Ashley Madison case, etc.). LinkedIn used sha1 with no salt as recently as 2012 (when they were hacked) for passwords and over 100 million such username + password combinations got stolen. Not only is sha1 not good enough for passwords but for many common and simple words (yes, yes, they are bad passwords, but people do use them) just googling can "crack" them due to lack of salt. The law should either go both ways or neither. To suggest such heavy handed laws like considering robots.txt ignorance hacking while multi million corporations with millions of users get away with stuff like that (and I mean true negligence of most basic practices, not some obscure bug in the underlying software or something else that isn't absolutely obvious) over and over and over again that every random free my-first-login-page and my-first-SQL-injection-prevention tutorials advise against is absolutely ridiculous and anti-consumer.