4 ms·
Surely the ethics are more complicated then just following robots.txt or not. The intended usage counts, and that isn't captured in robots.txt.
by lambdaba 2y ago
Surely the ethics are more complicated then just following robots.txt or not. The intended usage counts, and that isn't captured in robots.txt.
- marginalia_nu 2y agoIf you have a noble intent, you ask the webmaster for permission to use the data. Surely if they agree with your assessment that your intent is indeed noble, then you'll be given consent. I run a search engine and an internet crawler. I do this all the time. To this date I've never had a webmaster that didn't permit my crawler access when I've asked nicely.
- bryanrasmussen 2y agoIf you have a noble intent - identify members of fascist organizations - then obviously when you ask the top online fascist sites if you may scrape them to build up your list of online fascists - they will say no. OK less provocative, you have new algorithm to identify inaccessible websites, your automation is scary good, crawling a site you can identify many issues that most sites would have to pay for a full audit to get, but now these sites have problems - if you can identify their sites as being inaccessible then they have to fix these problems due to various accessibility standards that apply in the regions they operate in. But if they don't allow you access then they can maybe make an argument they are accessible due to audit they did last year, at any rate they don't want to be forced to spend money on accessibility issues right now which it sounds like they might have to if they let you crawl their site. Version 2 of above, some years ago I spoke about a job with a big time magazine publisher in Denmark and said one of the things that would make me a good employee is my knowledge of accessibility and their chief of development said they didn't have anyone with disabilities that used their site - so if I ask that guy to crawl their site why say yes? They have no users that would benefit!! Stop abusing our bandwidth bleeding heart guy.
- marginalia_nu 2y agoAll of these seem like variations of the-ends-justify-the-means, which generally tends to cut both ways in unanticipated ways. Bullying websites into accessibility compliance will most likely lead to them following the letter of the standard without giving a second of thought as to whether the content is in fact actually accessible. It's very difficult to get someone on board with your cause if your initial contact is an antagonistic one.
- greenbandit 2y agoThis might work in cases where those with the data are engaged in noble acts, but not ever actor is. I scrape and process websites of actors engaged in fraud. I do this to make the data more presentable to the proper authorities and to help uncover further evidence of their activities. I suspect that asking for consent would be quickly denied and the data/evidence would quickly become inaccessible.
- CaptainFever 2y ago> If you have a noble intent, you ask the webmaster for permission to use the data. Is Marginalia opt in, then? Surely "not having a robots.txt" ("you didn't say no!") does not equal consent. And surely you could just ask all the webmasters you are scraping from for permission, since you have noble intent. My point is that this is just hypocritical; you are placing the moral boundary right below what you are doing, while claiming moral superiority. If you ask others (e.g. anti-search Fediverse), they would think you are immoral too.
- marginalia_nu 2y agoYou really see no difference between following the robots exclusion standard, doing nothing to conceal your origins and intents, and respecting blocks when they appear; vs concealing your origins and intents, willfully ignoring the robots exclusion standard, and going to great lengths to circumvent IP blocks and other bot mitigation measures? Both of these are the same?