4 ms·
Google has a history of scraping content that they want, their business is built on the back of scraping other peoples content. The story I read just recently o
by DigitalSea 6y ago
Google has a history of scraping content that they want, their business is built on the back of scraping other peoples content. The story I read just recently of what happened to Celebrity Net Worth was an interesting read where Google asked for an API, they refused and Google just scraped the content anyway. There was no lawsuit, but CNW put up fake content and sure enough, it made its way to Google.
It is all ironic given how aggressive Google are in blocking any attempts to scrape its content.
- smabie 6y agoThat's like saying it's ironic that a soldier fights for his life when he tries to kill other people. It's just the war that is being fought, not some sort of hypocrisy or irony.
- harry8 6y agoGarbage. We live in a society of laws. Even soldiers. Google have shown they have no respect for the law not equality before it and will cheat while using the law as a cudgel. Recall law exists that the strongest might not always get their way. "Ironic" is the pole way of pointing this out. Without law, Google cease to exist immediately. They are incapable of enforcing property rights without it. Pardons aside, soldiers go to jail for taking an attitude like Google's.
- TAForObvReasons 6y agoJust like Genius, Google licensed the lyrics. If they didn't, the publishers definitely would have sued. Ironically, it is Genius that seems to have no respect for copyright law. Genius ended up having to settle a case years ago because they were using lyrics without the appropriate licensing [1]. https://www.nytimes.com/2014/05/07/business/media/rap-genius-website-agrees-to-license-with-music-publishers.html https://www.nytimes.com/2014/05/07/business/media/rap-genius...
- bawolff 6y agoWhich law did google break? Scraping in and of itself isn't illegal last time i checked, and usa doesn't have database copyrights unlike some juridsictions.
- a1369209993 6y agoIt's blocking scrapers that is (somewhat, per things like the Americans with Disabilities Act) and/or should be (in general) illegal. (And to head off the obvious: rate-limiting is orthogonal to whether the high-request-rate querient is scraping.)
- nl 6y ago> should be (in general) illegal But.. it's not illegal? > somewhat, per things like the Americans with Disabilities Act This is just not right at all. There is nothing in the Americans with Disabilities Act that make blocking scrapers illegal. I think you mean you don't like the power imbalance of the large company taking away from smaller companies while using technological means to stop the same thing happening to them. I don't like it either, but that doesn't magically make it is illegal. I'm not even sure it should be.
- a1369209993 6y ago> There is nothing in the Americans with Disabilities Act that make blocking scrapers illegal. Retrieving, processing, and displaying information in a manner contrary to the wishes of the provider of that information is necessary for accessibility to disabled users. As a specific example, any attempt to block use of wget for scraping also blocks use of wget as part of a `wget | filter | text-to-speech` pipeline[0], and is thus a discrimination against blind or otherwise visually impaired users. The ADA is, as mentioned, only somewhat effective in prohibiting such things, though. > it's not illegal > that doesn't magically make it is illegal. I don't think anyone is claiming that scraping itself actually is legally protected - I interpreted DigitalSea and harry8 as implying that it should be. 0: in either the shell sense or the workflow sense
- deleted 6y ago[deleted]
- anonytrary 6y agoProbably a silly question, but why not just use robots.txt? That was designed for preventing exactly this.
- asutekku 6y agoI’d say most of Genius’ visitors comes from the “song x lyrics” so hiding those with robots would ultimately make them lose almost all of their traffic.
- Polylactic_acid 6y agorobots.txt is designed to keep garbage off search results. It has absolutely no power to prevent a bot to do anything. Also if the site added robots.txt they might as well shut down because their entire userbase comes from people searching lyrics on google.
- mikemotherwell 6y agoOther way around. It was invented to stop crawling. Indexing is still technically allowed even when blocked by robots.txt from crawling.
- jtxx 6y agorobots.txt isn’t enforced by anything
- dewey 6y agoNot due to robots.txt but you can see what happens to genius formerly rapgenius when they get removed from the index: https://techcrunch.com/2013/12/25/google-rap-genius/ https://techcrunch.com/2013/12/25/google-rap-genius/
- encom 6y agoWear a condom before clicking techcrunch.com links: https://archive.ph/9eUkv https://archive.ph/9eUkv
- Avamander 6y agoThey also scrape MusicBrainz, but even if they don't index MusicBrainz at least they donate to it
- niknetniko 6y agoThey have an contract with MusicBrainz. They are listed on https://metabrainz.org/supporters/tiers/4 https://metabrainz.org/supporters/tiers/4. > The Unicorn tier is for large companies or companies that would like to have a reciprocal relationship with our foundation. If you need special guarantees, indemnities or require us to sign your contract for a data license, please select this tier. If you have another creative idea you would like to propose, please also select the unicorn tier. > For any of these cases, please detail your request in the company information field and we will work with you to fit your company's mythical situation. We will also find an appropriate monthly support amount to our non-profit foundation of $1500 or more per month. Please always consider enabling the growth of our non-profit foundation and the continuous growth of our metadata!