4 ms·
> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to thi
by threeseed 2y ago
> Google itself got big by indexing other people's data without compensation
Wrong.
a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest.
b) The difference here is that OpenAI, Meta etc have not even tried to honour the wishes of copyright holders. They just considered everything as theirs.
c) Google grew big because it had no ads, fast interface and PageRank was significantly better. It wasn't because it had the most comprehensive index.
- RALaBarge 2y agoTo your first point, the op said without compensation, not without permission.
- fredgrott 2y agopoint c is wrong...they had ads since the original yahoo contract....
- threeseed 2y agoYahoo contract was 2 years after it launched. I remember using Google the day it went public and it had no ads which made it unique compared to Altavista.
- karamanolev 2y ago> Web site owners chose to make it available to Google. Strong disagree. Since robots.txt is optional and the default is "crawl me as you please", website owners don't "choose to make it available", they just don't choose to make it non-available.
- XorNot 2y agoThat's a functionally meaningless distinction. If you setup a web server that responds to requests, then you're choosing to make content available because your server can choose to not respond to requests. The entire protocol includes mechanisms to negotiate access.
- jokethrowaway 2y agoGranting access and granting right to redistribute (even just title + snippet) and use your content commercially are two completely different things.
- XorNot 2y agoAnd yet it is legal to produce and redistribute summaries as sufficiently transformative derivative works, and this has been court tested[1]. Of course in Australia we passed rather specific laws to the contrary, because lo and behold Rupert Murdoch wanted money and gosh darn it our government was going to give it to him[2]. [1] https://www.practicalecommerce.com/Search-Engines-Indexing-and-Copyright-Law https://www.practicalecommerce.com/Search-Engines-Indexing-a... [2] https://www.alrc.gov.au/publication/copyright-and-the-digital-economy-ip-42/caching-indexing-and-other-internet-functions/ https://www.alrc.gov.au/publication/copyright-and-the-digita...
- eviks 2y agoThis is a meaningless simplification. In this framework "robots.txt" has no role, because your server "can choose" not to respond. Heck, even DDOS is fine, because "protocol"
- boesboes 2y agoWrong. Google ignores robots.txt entirely
- threeseed 2y agoI wasn't aware. Can you please update Wikipedia then: https://en.wikipedia.org/wiki/Robots.txt https://en.wikipedia.org/wiki/Robots.txt Maybe also get Google to update their docs: https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt?hl=en https://developers.google.com/search/docs/crawling-indexing/...
- nottorp 2y agoIt must be nice to believe everything people say by default... ;)
- phit_ 2y agotheir own docs also specify that the robots.txt does not stop indexing or showing up in search, they even bolded it "it is not a mechanism for keeping a web page out of Google" https://developers.google.com/search/docs/crawling-indexing/robots/intro https://developers.google.com/search/docs/crawling-indexing/...
- alphan0n 2y agoThe only way for links to appear in a Google search would be to host a public resource, that is linked from another public resource. If you have specified in your robots.txt that you do not want the page(s) or directories ingested then only the url is indexed (if it is linked from another page). It does prevent the public display of the content of a page and creation description/summary. https://support.google.com/webmasters/answer/7489871?hl=en https://support.google.com/webmasters/answer/7489871?hl=en
- threeseed 2y agoFrom the docs: “While Google won't crawl or index the content blocked by a robots.txt file” They will show the URL if someone else has linked to it. But the content itself is not indexed.
- tobyhinloopen 2y agoa) If you don't have a robots.txt, you're indexed by default. It's opt-out, not opt-in. If you do nothing, you're being indexed.
- antiframe 2y agoIt's an opt-out of an opt-in. If you run a webserver hosting your files, you already opted-in to people accessing that data. If you then don't go ahead an configure it properly, that's not exactly "opt-out" anymore. By default your files are not accessible to the network, you have to first opt-in to serving them.
- tobyhinloopen 2y agoGoogle makes a copy of your data and serves that data to users before they visit your site. Also google cache allows users to get a copy of your site without visiting your site. Why can they republish your data while we cannot? Why do we have to opt-out?
- veggieroll 2y agoRobots.txt is irrelevant after hiQ Labs v. LinkedIn (2019)