13 ms·
Does this break part 4 of the Goodreads TOS? "[...] you agree not to sell, license, rent, modify, distribute, copy, reproduce, transmit, publicly display, publ
by voidUpdate 11mo ago
Does this break part 4 of the Goodreads TOS?
"[...] you agree not to sell, license, rent, modify, distribute, copy, reproduce, transmit, publicly display, publicly perform, publish, adapt, edit or create derivative works from any materials or content accessible on the Service. Use of the Goodreads Content or materials on the Service for any purpose not expressly permitted by this Agreement is strictly prohibited."
Also did the reviewers give you permission to fed their content into an LLM?
- petralithic 11mo agoWhy ask questions you already know the answers to?
- hananova 11mo agoBecause some tech adjacent people still have morals?
- onetokeoverthe 11mo ago[dead]
- kosolam 11mo agoI’m not taking sides in this debate, however since feeding whole books into LLMs is considered legal fair use now, I guess these reviews don’t require a permission as well. Would be great to hear a professional lawyer take on this.
- saaaaaam 11mo agoThe hidden gotcha in the Anthropic judgement (which I think is what you’re referencing?) is that feeding whole books into LLMs is considered legal fair use if you obtain them legitimately. I suspect we need to wait for the NYT (and others) case to be decided before we know whether scraping sites in contravention of their terms is also fair use for LLM training. My own opinion (as someone who creates written content on an occasional professional basis) is that if you can’t monetise your content in some other way than blocking people from accessing it then your content probably isn’t as valuable as you think. But at the same time that’s tricky when it’s genuine journalism, as in NYT’s case. Obviously user generated content reviewing books online is rather different because the motivation of the reviewers was (presumably) not to generate money. And, indeed, with goodreads there’s a strong argument that people have already been screwed over after their good faith review submissions were packaged up as an asset and flogged to Amazon. A lot of people were quite upset by that when it happened a decade or so back. So from a ‘moral arguments’ perspective I don’t think scraping goodreads is as problematic as other scraping examples. (Sorry, none of this was aimed at you - your comment just got me thinking and it seemed as good a place as any to put it!)
- IncreasePosts 11mo agoGoodreads offers those reviews up publicly by serving them from their webservers to anyone who asks for it.
- saaaaaam 11mo agoSorry, I don’t understand the point you’re making. I know that these are publicly available - the point I was making, drawing off the parent comment, is that where it has been deemed fair use in copyright to use books to train LLMs when the content has been legitimately obtained then a similar assessment might apply for this sort of ingestion. If content is publicly available that does not necessarily mean it’s free of copyright control: the justification for using the reviews to train an LLM would be based on the fact that fair use means it is not an infringement of copyright. But if the publisher has terms that forbid scraping then that may mean the fair use argument is undermined if it is precedent in the content being legitimately obtained. I’m not a lawyer but it’s quite easy to see how “books can be used for LLM training under fair use but not if you pirate them” extends to “content on the web can be used for LLM training under fair use but not if you’ve breached the terms set out by the publisher”.
- esskay 11mo agoFairly meaningless in this day and age. Also IIRC scraping legality depends heavily on jurisdiction. Some places take a more permissive view of accessing publicly available information, even if a site's TOS forbids bots. In the US there’s a major precedent [0] which held that scraping public-facing pages isn’t a CFAA "unauthorized access" issue. That’s a big part of why we’ve seen entire venture-backed scraping companies pop up - it’s not considered hacking if the data is already public. [0] https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
- deleted 11mo ago[deleted]
- voidUpdate 11mo agoSo if you are legally allowed to "adapt, edit or create derivative works from any materials", what's the point of the TOS?
- hopelite 11mo agoThat’s a good question. It also would not be the first time that companies use trickery and manipulation or even deliberately illegal practices for various business/financial reasons. At the very least it could be used as a tool to underpin intimidating lawsuits and another step up, regardless of the legality in the relevant jurisdiction, it could be used to influence official government foreign policy to exert pressure on a jurisdiction that permits scraping.
- zigzag312 11mo agoI believe TOS is binding as long as it doesn't conflict with the law. If something is deemed fair use under the law, TOS cannot override those legal rights.
- hrimfaxi 11mo agoLegal rights are signed away all the time in contracts though.
- pantropy 11mo agoTechnically speaking none of Goodreads material or content is being used publically, the only information displayed on the site is freely available (Title, Author) and not Goodread's property. You could try to argue that this falls under "create derivative works from any materials or content accessible on the Service" but even then it seems really flimsy to say that recommending books based on Goodread reviews is an infringemnt. It's just not that different to a youtuber saying "I read reviews for 50 books, here's the ones to read"
- voidUpdate 11mo agoI'd be impressed if a youtuber could read 3 billion reviews and recommend books to you based on that
- tonyhart7 11mo agowhat about youtuber that build a machine that scrape 3 billions books and make recommendation based on the data????
- bravoetch 11mo agoSkip that step. This project enables a Youtuber that automates pulling related booklists from this site, and uses AI to make the recommendation videos. Thousands of videos.
- croes 11mo agoI visit your garden and take 1 apple from your tree I visit your garden and take 1000 apples from your tree. Not that different.
- kemotep 11mo agoNot only am I taking 1,000 apples, but I use those 1,000 apples to start my own orchard and encourage people to come to it instead of yours.
- MichaelBosworth 11mo agoWhat expectation of confidentiality are you ascribing to people having posted publicly accessible opinions on the internet? Out of curiosity, is your point about TOS out of concern for the poster or for Goodreads?
- voidUpdate 11mo agoMy expectation isn't of confidentiality, but of attribution. Sure, my website is perfectly accessible on the internet, and I'm fine with being able to find it on google, but if you pipe it into an algorithm that will start throwing out stuff based on what I wrote, with zero reference to me at all, I'd get a bit annoyed. This website has taken the combined output of probably thousands of people, shoved it into an algorithm and is then using their work to give "original" ideas. If one person wanted their content removed from the system, how would you do that?
- caconym_ 11mo agoWhat does that comment have to do with confidentiality?
- MichaelBosworth 11mo agoThat he viewed a review on Goodreads as the reviewer’s intellectual property hadn’t occurred to me. I see why, in aggregate, many such opinions become valuable, but the whole is more than the sum of its parts. So does it feel to you guys like your comments, say, here in this Hacker News thread should be considered effectively copyrighted as your personal IP? If so, do you feel the same way about opinions you share out in a supermarket or on the street?
- caconym_ 11mo agoThere are well established legal standards for what is copyrightable and I believe written literary criticism trivially qualifies (as it should). Stuff you yell at the supermarket doesn't, IIUC, as it isn't fixed in a tangible form. Social media comments are, IIUC, generally protected. The exception would be comments that don't meet the bar to be considered "original", "creative", etc. (not a lawyer)
- lunias 11mo agoIf it's on the internet, and people can access it, then it's public. I would have no expectations for what people do with public data; that just seems like setting yourself up for disappointment.
- contravariant 11mo agoAt what point are they feeding reviews into an LLM? From what I got the only personal data they're using is which user read which books.
- irl_zebra 11mo agoThis is, essentially, why I've withdrawn from posting content from my human brain almost anywhere on the open internet (except here, sometimes) and have retired blog posts, opinions, and so on to our friends WAN.