33 ms·
It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet
by munchinator 3y ago
It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim.
And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link.
I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP.
And the reason I bring this up, is that it seems like Open AI has the same attitude: scraping news articles is OK, or at worst a gray area, but what if they were also scraping, for example, Netflix content to use as part of their training set?
- Erratic6576 3y agoI find “4nn4’$ 4rch1v3 dot ORG” actually way better than pirate bay for pirating knowledge. It’s amazing the amount of books that copyright laws prevent us from finding https://www.theatlantic.com/technology/archive/2012/03/the-missing-20th-century-how-copyright-protection-makes-books-vanish/255282/ https://www.theatlantic.com/technology/archive/2012/03/the-m...
- munchinator 3y agoSure. It's just curious to me that news article have a pirated knowledge link as the de facto top comment, but link submissions to, for example, books for sale on Amazon don't have a link to Anna's Archive or equivalent.
- Txmm 3y agoI think the archive of an article is more preservation of history and maintaining records of events which often disappear if not archived. The number of threads referencing articles which are defunct is always increasing. A book or movie or original content on the other hand will continue to hold its own commercial value so reproducing it is more akin to an actual loss for the license holder. Definitely a grey area when that content is then used to train models though.
- Baldbvrhunter 3y agoI would say 9 times out of 10 it's to get around the paywall and absolutely not some higher moralistic preservation of history. And everything is a grey area, determining the line is the existential purpose of these court cases. We've been here before with hyperlinking, then indexing and then linking with previews and the Canadian Facebook stuff but I think this has more standing.
- nulbyte 3y agoIf I buy a book, I get a work of literature. But if I buy a news subscription I get a series of facts riddled with advertisements. I accept the former, but I oppose the latter. I suspect I'm not the only one.
- deleted 3y ago[deleted]
- Baldbvrhunter 3y agoI don't fully understand what you're opposing. is it? 1) that you paid for news 2) that it included ads both are just the price you want to pay. There are various state news outlets that you're probably already paying for - npr, pbs, bbc, cncb depending on your region
- cantSpellSober 3y agoThat's why you don't pay for news? There are browser extensions that block ads. They are called ad blockers.
- ralfd 3y agoThat is an apples to oranges comparison. An article about a video/book would have the relevant information in text form without needing to show the video "here is the new stuff shown in Apples 2 hour long WWDC keynote". If not is common that a comment in the discussion gives a summary as a tl;dr With text articles behind paywalls the relevant information is hidden and only hinted at as a teaser.
- munchinator 3y agoTo make it an apples to apples comparison, look at submissions where the link submitted is the retail link to the IP. For example, look at all the book link submissions on AMZN... https://news.ycombinator.com/from?site=amazon.com https://news.ycombinator.com/from?site=amazon.com None of these have the Pirate Bay or Library Genesis or Anna's Archive or the equivalent as the top comment. Compare that to... https://news.ycombinator.com/from?site=nytimes.com https://news.ycombinator.com/from?site=nytimes.com And almost all of these have an archived version as the top comment.
- tmhrtly 3y agoI wonder if this is because the purpose of linking to a book is to share awareness of that book’s existence - nobody is about to go and read it then and there to comment on its contents. Whereas the purpose of an article is to discuss it now, in the comments - the consumption horizon and bulk of the content is different.
- cactusplant7374 3y agoIf NYT was a HN startup the link to the archived version would be banned and dang would be slamming the ban hammer.
- 1f60c 3y agoPlease don't post baseless accusations. I think dang has said that he tries to moderate less, not more, when YC companies are involved. (Although it's impossible to say what he would do in this situation.)
- cactusplant7374 3y agoHN is currently facilitating piracy. Something your comment failed to address.
- quickthrower2 3y agoLike I said in another comment it is simpler than that. They just serve the login page/payment page to all HTTP requests. If they do that then the submission itself likely get's flagged as there is no workaround (just like if I submit my blog with a banner saying "hey you pay me $1 to read my cool post")
- Popeyes 3y agoPossibly because once an article is published the author receives no further payment. In all other mediums, there are residuals and royalties to be paid to the creators of the work.
- manojlds 3y agoAnd add to that fact that NYT subscription is hard to unsubscribe from. People have aversion to NYT, even setting aside the bias.
- hef19898 3y agoIt took me all of 5 minutes to cancel my digital NYT subscription from the following month onward. No idea what you are talking about.
- lupusreal 3y agoWhy did it take you five minutes instead of twenty seconds? It should be as simple as clicking on the link to your profile then clicking unsubscribe, mere seconds not minutes. Assuming you just said five minutes figuratively... Do you live in California or some other legal jurisdiction that forces them to play nice? Did you subscribe through some other company, like Apple? Horror stories about unsubscribing from the NYTimes are easy to find in the archive if you search for it. They make you call and chat to a retention specialist on the phone. This should help you have an idea of what he's talking about: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=NYTimes%20unsubscribe%20call&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
- hef19898 3y agoInternational one, as szraight forward as it could be: go to profile, go to manage subscription, cancel subscription, answer question why if you want, confirm cancellation, done for date depending on subscription.
- deleted 3y ago
- perihelions 3y agoIf it takes 120 seconds to read a newspaper article, the archive.is workflow is a significant overhead over that, a significant friction. Those links are a courtesy to other HN readers. This is very different from the economics of buying and reading a book. "Piracy is almost always a service problem and not a pricing problem." edit: It didn't even occur to me to compare the time-cost of "just pay for the article", but: last I read, it's half an hour of work to cancel a New York Times subscription [0]. So, that option's not even on the table. [0] https://news.ycombinator.com/item?id=26174269 https://news.ycombinator.com/item?id=26174269 ("Before buying a NYT subscription, here's what it'll take to cancel it", 812 comments)
- eropple 3y ago> edit: It didn't even occur to me to compare the time-cost of "just pay for the article", but: last I read, it's half an hour of work to cancel a New York Times subscription [0]. So, that option's not even on the table. I canceled mine two weeks ago. It was four clicks. One annoyed me because they tried to get me to stay with an offer, but I didn't drop them because of the price.
- dillydogg 3y agoSame experience here, it was effortless. But it is enough to justify stealing from those journalists, it seems.
- Germont 3y agoTo me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.
- afavour 3y agoIt’s a difficult problem with no great answers. If you want news to be free at the point of delivery you want public service news agencies. But that means they’re owned by the government… who are frequently the target of critical reporting.
- bongripper 3y agoThat's not true. You can have Independent public broadcasting that is not owned by the government and is reporting critically on it.
- afavour 3y agoIt’s still a difficult tension. The government will always control the purse strings so independence is always going to come with conditions.
- vidarh 3y agoThe Guardian in the UK is an example of an alternative: It is owned by a trust, which funds it. Norway has substantial public media funding across the political spectrum, but as you point out it always comes with conditions, even is less so than the funding for the state owned broadcaster. Combining the two models and putting public funds into several perpetual trusts intended to provide funding from their profits at arms length from any sitting government similar to the (private) trust funding The Guardian might be an interesting alternative. (EDIT: Norway also has its own variation over The Guardian model - the second largest media group was founded by unions but is now majority owned by the combination of two public benefit trusts)
- quickthrower2 3y agoIt is am ethical grey area, but if the paywall applied to all user agents, which would make it similar to say buying a Kindle book, then you might see that as pirating, whereas if you use an archive service that was served the HTTP response and cached it, then you are using a proxy UA. If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they want that SEO, and they want that marketing.
- EvgeniyZh 3y agoIndeed there are media that are hard paywalled, e.g., the information. However these are prohibited on HN, which possibly create additional bias towards non-hard-paywalled publications
- Yizahi 3y agoWe can extend this analogy. What if someone put up a proxy, that has a legal Netflix subscription and which "watches" streams of Netflix shows, captures actual RGB values of pixels and re-streams the resulting video to anyone else? Isn't it the same "proxy" excuse?
- quickthrower2 3y agoI would say no because the site was happy to serve the content publicly, whereas your proxy is breaking a contractual agreement. Now we get into terms of service of a website, and even if you visit for free you agree to them. Which is a possible point. It is quite grey IMO. In terms of HN I reckon a mag would love the free brand rec. vs. the archive not being shared. Where it hurts them is if someone is avoiding paying for a subscription by continually using archive sites.
- kriro9jdjfif 3y ago[flagged]
- Yizahi 3y agoGood comment, it was very funny to see how people desperately try to find moral justification for pirating media A but not B. "It's apples to oranges, you see, there are less letters in the NYT article than in the book and they are rendered differently, so it is ok to pirate their work. I did nothing wrong!" :)
- sjfjsjdjwvwvc 3y agoOf course pirating any media is totally fine from a moral standpoint.
- fodkodrasz 3y agoIt is actually pirating content by companies for humongous profit, or pirating by individual human beings for free access to culture and entertainment, oftentimes for content one has already paid for, but rendered inaccessible by megacorporations.
- lotsofpulp 3y agoWhich content making businesses earn humorous profit margins? Are all the journalist layoffs a fever dream? This is one of the more profitable ones, and only because they employ unscrupulous tactics: https://www.macrotrends.net/stocks/charts/NWS/news/profit-margins https://www.macrotrends.net/stocks/charts/NWS/news/profit-ma... This is NYT, the most successful news business: https://www.macrotrends.net/stocks/charts/NYT/new-york-times/profit-margins https://www.macrotrends.net/stocks/charts/NYT/new-york-times... As for movies/tv show/music makers, let’s just say most people in the software engineering business would look at their numbers and count their lucky stars that they are not in the movie/tv show/music business. (It is also true that excessive copyright lengths have removed access to content that the public should have a right to).
- defrost 3y ago> Which content making businesses earn humorous profit margins? https://en.wikipedia.org/wiki/Mad_(magazine) https://en.wikipedia.org/wiki/Mad_(magazine) https://www.theonion.com/ https://www.theonion.com/
- throwaway22032 3y agoBlocking ads and avoiding payment are two different things.
- iinnPP 3y agoThe archive link doesn't threaten their jobs and helps them avoid paying for NYT. It's NIMBY, or rather it's true form of NIIIM (Not if it impacts me). Hypocrites are EVERYWHERE and are the majority.
- bnralt 3y agoIt is pretty funny. If you go back and read the comments made yesterday about ChatGPT doing something much milder (using old articles to train data, some prompts fused to allow you to reproduce some of the articles though now don't work), you have a lot of comments talking about how The New York Times needs money and Open AI is using their work without paying for it. Now a comment points out that HN News (and most of the internet) routinely does something much worse - allows people to bypass completely new articles in their entirety without paying - and almost all the comments are about how it's the New York Times fault for making it difficult to cancel subscription, the importance of news being available to everyone, the problems with copyright laws, etc.
- lexicality 3y agoFunny, I don't see it as a moral thing but more a "what can you get away with" thing. I fully assume that if I was to post a magnet link to a torrent for whatever the link was about, I would be banned. Morally speaking, I think it's perfectly reasonable to download a copy of something and either read the relevant info for my current task or to sample it to decide if I want to buy it. I see it no different to using the library or browsing at a book store. Perhaps once news organisations can work out how to effectively wield the DMCA hammer against archive links we'll see the practice of posting them stop.
- midasuni 3y agoSo downloading a movie from piratebay is no different to using the library?
- kolinko 3y agoIn some jurisdictions (Poland, possibly whole of EU), downloading any kind of materials - be it movies, books or music - is legal. Uploading/sharing - if not between friends&family members - not so.
- mistercow 3y agoI’d argue that morality always has a “what can you get away with” component. Things that are normalized tend to be seen as morally permissible, and things that are seen as abnormal are more likely to be seen as immoral. The problem with the thinking in the root comment is that it implicitly assumes that people’s behavior is morally consistent, or that they even try particularly hard to behave in a morally consistent way. That’s not really how people work. If you ask them to discuss morality in the abstract, they’ll try to come up with a consistent system. But their actual behavior is mostly dictated by social norms. And if you try to pin them down on the morality of their concrete actions, they’re more likely to stretch their moral system to accommodate their actions than the other way around. None of this is to say anything about my own opinions on news sharing or OpenAI’s situation. It’s just that someone decrying piracy but also posting/sharing/upvoting links to copies of news articles is neither surprising, nor indicative of some deeper nuance to how people view morality around IP.
- lupusreal 3y agoThis is only an interesting juxtaposition if you have fully internalized and accepted the myth of people and corporations being interchangeable.
- seydor 3y agoit s also audacious how these news companies reproduce stories from social media and other electronic media of facts that are, like, freely available in nature. Or how they get embargos and exclusivity to government information as if they are some kind of information-bouncer
- puttycat 3y agoThey pirate movies as well: https://garymarcus.substack.com/p/an-artist-fights-back-and-midjourney https://garymarcus.substack.com/p/an-artist-fights-back-and-...
- unyttigfjelltol 3y agoHistorically newspapers leaned more on competition law than copyright, because their pages are supposed to be filled with non-copyrightable facts.[1] Copying part, but not all, of a factual article, significantly after the relevant event, was considered to be a promotion (not unfair competition) and a nice thing to do for the journalists. Things change, people lose sight of the original principles. [1] https://en.m.wikipedia.org/wiki/International_News_Service_v._Associated_Press https://en.m.wikipedia.org/wiki/International_News_Service_v...
- nsagent 3y agoThese days most news is mixed with analysis [1] (which is often biased). I wonder if part of the reason for this shift is that analysis is copyrightable. It also seems like the number of opinion articles is ever expanding [2], though I don't have any hard numbers on that. [1]: https://guides.library.cornell.edu/evaluate_news/source_bias https://guides.library.cornell.edu/evaluate_news/source_bias [2]: https://www.newsmediaalliance.org/rise-of-opinion-section/ https://www.newsmediaalliance.org/rise-of-opinion-section/ Interestingly there's a banner at the top of that link touting an agreement between Axel Springer and OpenAI. EDIT: formatting
- guipsp 3y agoEven the facts are not copyrightable, the prose is.
- vel0city 3y ago> their pages are supposed to be filled with non-copyrightable facts This is rather inaccurate. A fact is Hitler invades Poland. You're right, nobody can copyright this idea, as it is just a fact. However, if I then write a 500-word article describing the scene of Hitler invading Poland, have short quotes from some civilians there, etc. that particular arrangement of ideas and words is copyright. AP can't go and sue INS for just reporting the fact Hitler invades Poland, but if INS takes a whole article word for word and reproduces it that's still violation of copyright. The actual printed words of the news always had copyright. The WSJ can't claim copyright on the markets going up yesterday. They can claim copyright on something like "After the bell rang in the NYSE, the tech industry ticked up 1.2% over last week. Meanwhile the whatever market took a hit of -0.5% ending the quarter slightly lower than our analysis expected. Blah blah blah..." If Investor's Business Daily wrote a different article that also talked about the markets ending up at the end of the day, that's not a violation of copyright. If they literally write "After the bell rang in the NYSE, the tech industry ticked up..." then they're violating WSJ's copyright. This was true before and after International News Service v Associated Press.
- phpisthebest 3y agoLargely because "news" aka facts is not and should not be copyrightable, so while the style, and exact format of the article may be copyrightable, the facts contained within are not. This makes a news story copyright murky in the eyes of wider society unlike a clearly 100% creative work like a TV Show or Movie. Further the news themselves self cannibalize, how many stories are just rewrites of stories from other outlets? why it is OK for the Washington Post to copy the NY times, but not ok for OpenAI or Archive.org?
- gnz11 3y agoCreative works like books, TV shows and movies contain facts too.
- phpisthebest 3y agoNone of which are copyrightable and infact has been the subject of DMCA abuse like when a Movie uses NASA footage and claims copyright on YouTube videos with the same footage. Copyright is a complex subject, and not as vast as many believe, at the same time ironically it is more vast than i believe it should be. copyright should be much more limiting than it is. Which is at odds with people that believe copyright should be maximized. Keeping in mind commercial success of a work, author or company is not why copyright exists. For the US, the only reason copyright can exist in our framework of law (i.e the constitution) is for the promotion of the useful sciences. No other purpose for copyright would be constitutional under the US Constitution
- gnz11 3y agoCopyright doesn’t exist solely for the “promotion of useful sciences”. https://en.m.wikipedia.org/wiki/Copyright https://en.m.wikipedia.org/wiki/Copyright
- phpisthebest 3y agoCiting Wikipedia you already failed.. That is a General Article about Copyright world wide, I Specifically stated US Copyright, which is Authorized by Article I, Section 8, Clause 8 of the United States Constitution[1], implicitly for the promotion of the useful sciences. That is where congress derives its power to pass copyright laws, and to enforce copyright on the people of the United States. No other purpose is authorized by the US Constitution [1] https://www.law.cornell.edu/wex/intellectual_property_clause https://www.law.cornell.edu/wex/intellectual_property_clause
- maxboone 3y agoProbably because the contents are what's posted, i.e. if someone would post a link to an interesting video behind paywall / login and there was an easy mirror available that'd be posted too. If I could just buy one article for a coffee without entering a bunch of PII or go through a time-wasting process I would agree on the moral equivalence between the examples.
- jalapenos 3y agoIf ChatGPT is based on neural networks, with no actual save-and-replicate facsimile behaviour, it no more "copies" original work than I do when I tell you about the news article I read today. I'd say the only real reason the Piratebay links thing you mentioned is not the norm is purely because those media sources have done a better job of striking fear into people doing that, so it's gone more underground. I.e. they're better terrorists. There's no fundamental, moral reason why Piratebay links being posted and raised to the top would be wrong.
- octacat 3y agoSo, if someone applies a filter to a video/audio, it is no more "copies" of the original work (no, it is still protected). AI still could produce exact or extremely similar results of stuff it learned on.
- concordDance 3y ago> AI still could produce exact or extremely similar results of stuff it learned on. Can it do so more than a human can? I think that's the key here. If an AI is no more precise than a human telling you about the news article they read today then ChatGPT learning process probably can't be morally called copying.
- octacat 3y agoSo, if someone decompiles a program and compiles it again, it would look different. "It is not copying", we just did some data laundering. Feeding someone else data into your system is usually a violation of copyright. Even if you have a very "smart" system, trying to transform and obfuscate the original data.
- Matticus_Rex 3y ago> Feeding someone else data into your system is usually a violation of copyright In some circumstances, yes, but often it's not, especially if you're not continuing to store and use it (which OpenAI isn't).
- caeril 3y agoOh it's worse than that. The NYT is positing that any neural network that is trained on their data, and can summarize or very closely approximate an article's content on request, is in violation. This reasoning would presumably apply to any neural network, including one made of neurons, dendrites, and axons. So any human reader of the NYT who is capable of accurately summarizing what they read is an evil copyright violator, and must be "deleted". Effectively, the NYT legal department is setting the stage for mass murder.
- cycomanic 3y agoHyperbole much? There is a difference between a computer and a person. I'm not aware that people generally can be enticed to reproduce full articles verbatim just through questioning.
- ako 3y agoAs far as I know schools have to pay for the newspaper articles they use in class to educate students. Training an AI seems similar. Here’s a service for the UK providing paid access to copyrighted materials to schools: https://www.nlamediaaccess.com/newspapers-for-schools/ https://www.nlamediaaccess.com/newspapers-for-schools/
- mlindner 3y agoAt least in the US, copyright violation is a civil thing, it's handled by lawsuits. If the copyright violation is of such a small level that it's not worth the copyright owner to do anything about it then nothing's done. In this case it's worth a massive amount of money.
- davedx 3y agoI pay for multiple streaming services because I get a decent amount of value from their content. I do not pay for any news websites because I read very little of what they produce, and it tends to pop up more on aggregator sites like HN than me actually going to them. I actually did have a subscription to The Telegraph for a few months at one point because initially I wanted to read a full article (without cheating). But eventually I cancelled because so much of it is polemic trash. That's my justification: I pay for things that have value to me.
- ekianjo 3y ago> , a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link. Probably because most print media is garbage and nobody in their right mind would actually pay to read them
- davedx 3y agoI don't understand the downvotes - it's an extremely valid opinion. If people ask questions like that then they should be able to accept forthright answers? (It's the same reason for me. I have tried news site subs but eventually got so tired of the polemic that I cancelled. I won't sub again).
- iudqnolq 3y agoThe obvious response is that if you don't like news and think it has no value then you don't have to read it.
- sumedh 3y ago> Probably because most print media is garbage and nobody in their right mind would actually pay to read them NYTs revenue keeps growing though.
- octacat 3y agoAt least people do not obscure who is the original author of the content (so, if people like NYT articles - they could go and subscribe for more). Kinda "free advertising" (which still hurts the publisher in many cases, though). Same with search engines - as long as engine brings clicks - people are happy. If search engine just grabs the info and never redirects the user to the site - what is the point for the site to exist to begin with?
- initplus 3y agoI would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washingtonpost.com, or bloomberg.com, I want to read news.ycombinator.com. Paying an individual subscription to every possible underlying site that could be linked to from news.ycombinator.com is a non-starter.
- iudqnolq 3y agoThis is a common statement, but every attempt to sell that service has been a dismal failure. See for example blendle.
- deleted 3y ago[deleted]
- initplus 3y agoNearly every attempt at starting a new aggregation site like hackernews or reddit has been a failure. I’m not going to switch to a new website where no community exists just so I can pay for news articles. To work it needs to be integrated into an existing, successful aggregation website.
- jacquesm 3y agoBlendle failed because they went into competition with the papers whose content they reproduced.
- nithril 3y agoI would be happier to pay a small fee per article I want to read. But the norm seems a monthly subscription.
- elpocko 3y agoGood observation. I now wanna start commenting with pirate links to other media, but HN would tear me to shreds real quick I guess.
- anonfromsomewhe 3y agoit's similar to how easy it's to subscribe NY times and then how hard it's to unsubs. They require extra steps and it's well known. So They get what they deserve? Do you see the poínt. They are lie spreaders, nothing else
- cesarb 3y ago> Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. [...] And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link. If the story was linking directly to the "book, TV show, movie, video game, album, comic book, etc", and the link only worked for some people while others randomly got a login request or similar, you'd also see the top comment being a link to an archived version which avoids the login screen. That is: the main difference is that the archive link has the exact same content as the link submitted in the story, only bypassing the login screen that some people see. And the only reason the archive site has the content is that it didn't get the login screen; if everyone always got the login screen, what you would see on the archive site would be the same login screen.
- some1else 3y agoOkay. https://www.netflix.com/browse?jbv=81714181 https://www.netflix.com/browse?jbv=81714181
- infecto 3y agoi don’t believe that is fully correct. The general policy here is that you cannot link to something that is paywalled unless that site plays the game of allowing crawlers but not actual human eyeballs. In the latter case the link is allowable because there are ways around it that the site owners allow.
- tzs 3y ago> I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP. In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people explaining why they pirate (usually with ridiculous justifications like "subscribing is not an option because even though this paid service does exactly what I want now at a price that is trivial for me they might someday later change"). The impression I've gotten is that piracy of nearly everything is widely felt to be OK here. Information wants to be free, yada yada. About the only piracy that is consistently frowned upon here is piracy of open source software. When some company sells an embedded device that uses GPL code without releasing the corresponding source that's viewed as just a little short of a crime against humanity.
- alfiedotwtf 3y ago> About the only piracy that is consistently frowned upon here is piracy of open source software. When some company sells an embedded device that uses GPL code without releasing the corresponding source that's viewed as just a little short of a crime against humanity. Like what you said... > Information wants to be free
- tomComb 3y agoYeah, I don’t judge people for pirating or ad blocking, but the ludicrous justifications do get me - quite the entitled mental gymnastics. They remind me of bitcoin people trying to explain how mining is good for the environment.
- _jal 3y agoThere's a "polite society" thing going on. Briefly, something like: 1) Ycombinator could not tolerate HN becoming a site known for sharing IP-law-violating content. And the people who come here by and large are smart and socialized enough to implicitly understand why. 2) At the same time, a large number of folks here mostly wink and nod at that sort of consumer infringement. And there's a society-wide bias towards "things like news are less protected", so that gets to slide. 3) But people also have a need to tell consistent-seeming stories about how things work, thus the mental gymnastics. It ends up being similar to trying to explain why people pretend to be prudish innocents about sex. It largely reduces to "a small subset of the population goes sufficiently ballistic about what I consider to be relatively trivial stuff as to make it not worth fighting over, even if I find that to be ridiculous." There are a lot of different versions of this that become so normalized it can be hard to notice.
- Alex3917 3y ago> why we feel it's OK to pirate news articles, but not other IP. Once the NYT pays reparations for the Iraq war, I'll be the first to stop pirating it.
- bnralt 3y agoThis tendency at Hacker News are also much more of a threat to The New York Times than what Open AI is doing. Even the places like blogs/Reddit/social media submissions that summarize the article and post the relevant quotes. Unlike the summary of a movie, summarizing all of the relevant parts of a news article is extracting almost all the value from it, and giving it away for free. And the vast majority of people read news for it's breaking content, not for its archived content from years before (and I say this as someone who has often recommended the latter, but has gotten very few people to do so). So giving people that free breaking content (either in its entirety like on Hacker News, or summaries like you see all over social media) is actually a direct competition to the news business in a way that training an LLM on an article from months/years back isn't.
- skybrian 3y agoYes, and for nonfiction, it's also true that it usually depends on the original article for credibility. (If it were an anonymous poster making up a news story, most people wouldn't believe it.)
- tw1984 3y ago> why we feel it's OK to pirate news articles, but not other IP. Because those who own & produce such news articles asked to make them different. People listened and accepted their requests. When you make a TV show or a video game, you don't get any protection from the Geneva Conventions and a long list of other international treaties for your rights on stuff other than the content you are producing. The same can't be said when you are producing news.
- billywhizz 3y agothere's quite a big difference between "pirating" digital content and making it available to anyone for free and taking that content and building a for-profit service on top of it, which is what OpenAI are doing, no?
- pcmaffey 3y agoI was just going to post this. Seems quite an obvious and significant distinction, that doesn’t need to provoke all the existential hand wringing. Making money off someone else’s content is a totally different moral and legal case.
- ks2048 3y agobut what if they were also scraping, for example, Netflix content to use as part of their training set? There were some tweets the other day about how Midjourney could be prompted almost-exactly reproduce some frames of the film Dune. It wouldn't be shocking if these companies were using large databases of movies, with questionable legal status.
- j-bos 3y agoI see this a lot, and they very well may be. But, watch any behind the scenes documentary about any artsy movie and 9 out of 10, the director's will be waxing poetic about their inspirations, often include older movies or paintings which have uncannily similar scenes/frames. So it also wouldn't be shocking if a model trained on the same inspirations as the filmakers generates almost-exact frames as the movie makers.
- wilsynet 3y agoThe NYT and other newspapers don’t go after the archived link providers. Probably because the newspapers scholarly mission includes things like preservation. But they also have a profit motive or they can’t stay in business. This implicit permission for the archive links to exist, gives some of us the implicit permission to pirate the content. Disclaimer: I am a happy subscriber to the NYT (and other digital newspapers).
- detourdog 3y agoThe difference is that an individual pirating news is simply reading the article. OpenAI intends to digest news articles to the point of packaging them and reselling. My uncle used to distribute daily newspapers and his saying was "News ages like a fish". OpenAI is allegedly using NYTimes articles to train a computer and sell its services. I see different use scenarios. I guess another way to look at it is that human just reads the pirated material. A computer makes a verbatim copy and analyzes it to the point to mimicry and sells fuzzy versions.
- batch12 3y agoI believe it's tolerated here based on the site guidelines. I have always thought this was the case because otherwise these posts would all be pay to play which would limit who could participate and turn HN into more of a subscription farm. Maybe the way to make everyone feel ok about it is to disallow links to paywalled content.
- StanislavPetrov 3y agoThere are two fundamental differences. First, Open AI is the one doing the pirating here. Hacker News is the host, they aren't doing any pirating or posting any archival links to the copyrighted information themselves. Second, Open AI charges subscription fees and profits off of the copyrighted material they have pirated, whereas Hackers News does not, nor do the people who post the links.
- DennisP 3y agoI wouldn't say OpenAI has exactly the same attitude, since they also pulled in thousands of books. Their position has been that it's not piracy, since they don't republish the books; effectively the AI just reads them and learns from them. If GPT can be made to reproduce the original articles, that's a more difficult argument to make.
- Matticus_Rex 3y agoIt turns out you can reproduce articles with next-token prediction when the articles are quoted all over the dataset. The articles themselves are indisputably not a part of the model, because it doesn't store text at all. OpenAI's position is correct; people just underestimated how well the AI learns from reading, especially when it reads the same text in a bunch of different places because it's being quoted/excerpted.
- eigenket 3y agoIf it can and does reproduce a piece of text verbatim then the text is indisputably stored somehow in the model.
- Matticus_Rex 3y agoThat's just not true. There's no search and retrieval involved. It just associates the words so strongly in that context because they were in the training data so often that next-token prediction can (sometimes, in some limited circumstances) reproduce chunks of it. It's like if a human had read pieces of an article so many times and knew NYT style so well that they could spit out chunks of an article verbatim, but using more efficient hardware and with no actual self-understanding of what it's doing.
- briansm 3y agosort of like the idea of practice - repetition of something concentrates more brain space to that thing so the compression ratio of it can decrease and become less abstracted / more exact.
- rich_sasha 3y agoNot quite what parent means, but an interesting angle is: what if you scraped ChatGPT instead. NYT, or someone's blog? Meh, fair use, and if you say no, you're in the way of progress. But if you wanted to scrape ChatGPT answers to tweak your network, uh oh, violation of T&C!
- bitlax 3y agoBecause I'm not interested in the medium itself, as I would be with a Netflix show; I'm not even interested really in the article or the New York Times as an institution. I'm interested in discussing the supposed real-life phenomenon being covered, and the posted content is the primer for that discussion. I think if you get rid of the archive links on HN you need to ban the paywalled content as well. If you want to discuss paywalled content I'm sure you can do that in the article's comment section.
- jtc331 3y agoA book, TV show, movie, video game, album, or comic book is not available on the internet served by the copyright holder’s own servers with no authentication or authorization checks. But the NYT is available in that way.
- CamelCaseName 3y agoBut some are? I believe The Atlantic and The Economist are hard paywalled.
- cesarb 3y agoIf they're hard paywalled (everyone gets the same login prompt), they won't be available on archive sites.
- Zenst 3y agoWe are also happy to use open source, yet what open source alternatives are there for news that don't get shot down by the media or besmirched?
- orbisvicis 3y agoIf I can't read about it, it didn't happen.
- edude03 3y agoI think the intent is really different. For LLMs you're essentially teaching them language by showing them lots of examples of written language - newspapers are of course a great example of written language. The goal of OpenAI is not to reproduce newspaper articles verbatim when asked questions (even if the answer could be a newspaper article) and the fact that it can happen is a side effect of how LLMs work. When a HN participant shares a (pay walled) link to a NYT article, I do want to read the exact article linked verbatim because while the facts of the article may be reproduced elsewhere in a form that's free, specific word choices or whatever might be a focal point of the discussion on HN, and therefore I can't realistically participate in a discussion without having read the article being discussed. And as an aside, I have no problem with paying to read news, or whatever media, however it's impractical for me to subscribe to every news source HN participants link to, and therefore I gravitate to archiving services instead. I do wish there was a better solution - for example Blendle with more sources.
- rickydroll 3y ago> The goal of OpenAI is not to reproduce newspaper articles verbatim when asked questions (even if the answer could be a newspaper article), and the fact that it can happen is a side effect of how LLMs work. This is an excellent point. A properly functioning LLM should not return the original content it was trained on. When they return original content, I believe the prompt is tightly constrained and designed to extract or re-create original content. Another reason that occurred to me recently is that maybe the training set is too small, and more general prompts will re-create source material. Another question would be, are LLMs regurgitating what they were trained on, or are they synthesizing something very close to the original content? (Infinite Monkeys, Shakespeare). Court cases like this increase the need for understanding the "thinking processes" in an LLM.
- adolph 3y ago> The goal of OpenAI is not to reproduce newspaper articles verbatim when asked questions (even if the answer could be a newspaper article) and the fact that it can happen is a side effect of how LLMs work. Seems like a nice split-the-baby resolution would be to send the NYT Corp a single article read amount anytime GPT plagiarizes more than what’s allowed at an academic institution.
- FrustratedMonky 3y agoIs this really copywrite? Or is it "you can't talk to someone about an article they read". This is really saying you can't call up your buddy and have them tell you a summary of what they just read. Maybe my buddy has a good memory and some of the text is actually nearly duplicate. But I wouldn't know because I didn't read the original, I just asked for a summary from someone else that read it.
- chmod775 3y ago>I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP. A lot of of that is going to stem from the fact that respect for "journalism" is pretty low. More than 99% of news articles are copies of the <1% of original work that happens in that field. In news, everyone is already lifting content from everyone else.
- breck 3y ago> why we feel it's OK to pirate news articles, but not other IP Who thinks this? I don't. I think copyright is wrong across the board. I would love if the same pattern of posting archive'd articles held for books, movies, et cetera. I would love to change my mind on this, as it is a very unpopular opinion to have. But I have _never_ seen a morally or scientifically sound argument in favor of copyright law, and I've spent decades looking. I think it subsidizes the creation of junk food content (superhero movies and clickbait news for example) while not contributing anything to the progress of science (paywalled scientific journals and textbooks). I shudder how much time I have wasted in my life consuming crap attention grabbing media and advertisements. I like to think if we lived in a world where everyone could be a publisher if they wanted to, the quality filters would be better, and information reaching us all would be more likely to be in our best interests.
- metabagel 3y agoYou can self-publish. Oh, you want to be able to publish other people’s work, and without their permission? How does that benefit the author?
- breck 3y ago> How does that benefit the author? You speak of "the author". But the current system does not benefit "the author". 1% of authors profit off copyright. 99% lose money on copyright (they pay more for copyrighted media than they earn from it). Your question should be "How does that benefit monopolist authors"? I agree, my idea would not benefit monopolist authors. They would lose the bulk of their revenue stream. But it would benefit the average author whose cost of living would fall and information would start serving them more than serving business. I am not downplaying the talent and hard work of successful monopolist authors. But I do not think the works they create are worth everyone giving up their rights to reshare and remix information. I believe the world would look very different post-IP. You'd probably have a new profession--small independent librarians (similar to data hoarders today)--who would help their local communities maximize the value they got from humanity's best information. Maybe I'm wrong! Maybe the information ecosystem is better controlled and the genetic differences of monopolist authors are so stark that without the subsidies to this gifted class we'd all be worse off. But that's an argument based on outcomes and not principles. > without their permission The oxygen I'm breathing right now mostly was created by trees on land owned by others. But I don't ask for their permission to breath. Some things are just not natural. I am not saying plagiarize. It is always the right thing to do to link back and/or credit the source. But needing to ask permission to republish something seems to go against natural laws.
- u32480932048 3y agoAs a supporter of piracy in the general case, I tend to agree with your observations, including that pirating NYT (FT, NPR, ...) articles is somehow some kind of different class of offense as, say, stealing a movie or mp3. (Books, to me, are separate still, in that I like to have a physical copy (and generally see the authors as humans who deserve compensation, rather than mega-orgs that deserve eternal torment), so I'll frequently use the digital copy as a kind of preview, then purchase it once I see it's a good book I want to read.) I've only been reflecting on this difference for a few minutes, but, to me, I think the major difference boils down to: 1. Netflix series (movies, albums, etc) are non-essential, fictional works that take a long time to produce - think: fancy chocolates and caviar. 2. News, generally, contains timely, important information - more meat and potatoes. 3. While much of the super-critical news is not paywalled (e.g., product recalls, election dates, COVID stats, etc), a lot of information that is advantageous to know (discussions on interest rates, details on legislation, etc) is paywalled, compounding information asymmetries. Sure, "stealing bad", but, IMO, someone stealing rice and beans from WalMart to feed their family is a different class of offense than someone robbing a boutique bakery because they can't get enough chocolate cake.
- observationist 3y agoFirst and foremost, and please repeat after me: Copying is not stealing. You're not depriving anyone of anything. Unauthorized copying is not theft. There's no equivalency. You can't copy and paste a cake. If you take a cake from a bakery, you're depriving the bakery of a thing. If you take a picture of the trademarked bakery's sign, copy its the copyrighted text from its website, and print them out, you haven't stolen anything. Nobody has lost anything. Nothing was damaged. No person, place, or thing was harmed. Current copyright law is offensively absurd. Patenting of software, effectively eternal content copyrights, ridiculously broken DMCA, music publishers taking 99 cents of every artist's dollar, and so on and so forth. If you support the dissolution of archaic institutions and broken laws favoring those with entrenched wealth over individual rights, you support piracy. There is a legitimate case for laws respecting and protecting intellectual property rights. Such laws do not currently exist. These laws do not deserve to be followed or respected, and should be broken as a matter of course. Civil disobedience is called for. Refuse to participate in an exploitative market immovably entrenched in governments all over the world. Pay artists directly and commensurately if you feel they've brought value to your life. Copy whatever you want. Share those copies with whomever you want. Nobody gets hurt. Only conglomerates of already wealthy individuals and corporations are "deprived" of the potential transaction with you that they feel they are entitled to, as a matter of course. The NYT is just as complicit as any other legacy media institution in the enshittification of journalism and laying waste to the potential value of their content. The "Gray Lady" is not a person, or a valuable institution. It's a soulless corporate construct not deserving of our empathy or high regard simply because of the reputation of human individuals who previously produced quality content. Stop pretending these institutions serve some higher purpose than to fatten the wallets of shareholders. The good journalists have left. The ones left behind are naive, or are desperately clinging to an illusion of legacy and institutional legitimacy that no longer exists. All that is left for these media dinosaurs is to leech off the success of others, to use their reserves of wealth and influence to arbitrarily insert themselves into the market, with no regard to the fact that they no longer have value or prestige or purpose in the context of modern technology and communication. Anyway. Copying isn't theft. Don't give them the linguistic territory. Call a spade a spade, and media companies the desperate corporate leeches that they are.
- jasoneckert 3y agoI believe the reason many of us tolerate links to news articles and other content is because we believe in equality when it comes to information access. In other words, many of us believe that those who cannot afford a subscription to a paywalled site should still be able to read the articles, in much the same way public libraries allow those who cannot afford to purchase a book the ability to read it. However, this doesn't apply to organizations that freely share copyrighted information while making money in the process, or to organizations that share copyrighted information in a way that specifically disadvantages or does harm to the original creator of that information.
- raldi 3y agoI would broaden the question beyond HN to society as a whole. In 1990 it would have been considered normal and appropriate to clip an article out of a newspaper and post it on a communal corkboard. What are the key differences between that form of IP and others, and that analogy and the present situation of HN allowing archive links?
- layer8 3y agoReach, and ease of distribution.
- raldi 3y agoMakes sense. If you mail a friend a clipping, or post it on the corkboard, only so many people are going to see it, but then even though posting the "clipping" to HN may feel like the same thing, it's hard to appreciate the massive change in scale. As for ease of distribution, that might address OP's original question: It's easy to make and click an archive link, but it's a lot more effort to make or find a Pirate Bay link to another form of media, and for someone else to download and view it.
- cantSpellSober 3y agoIt's not just tolerated, it's encouraged because "the alternatives suck worse" https://news.ycombinator.com/item?id=23735026 https://news.ycombinator.com/item?id=23735026 Even talking about it will get you scolded for talking about something "off topic"
- zzzeek 3y agoit's different reading an NYT article on an archive site vs. putting copies of it at the core of your $100B for-profit content delivery enterprise.
- cwmma 3y agoI think one of the key differences is something pointed out in the article, in that what the Open AI is doing is a substitute for reading the new york times and possibly a rival to it. On the other hand having an archive link to a times article in order to discus it is not really a substitute for a times subscription as a news paper has to walk a line of letting some of it's articles be read while requiring payment for others (the times actually allows you to create a "gift link" to do exactly what the archive links do).
- kjkjadksj 3y agoBecause historically this is how news were shared. People would pick up a paper in a grocery store or cafe, read some of it, and leave it behind. They might rip out a page and take it home. Only one person paid and tens or hundreds gleam for free. This idea of sharing the story to nonsubcribers is as old as printed news itself. Instead news agencies prefer we forget that aspect of history, insist on being the “paper of record” while charging more money for easier to distribute media that gets sold globally. Yes, I think we are certainly not in the wrong here when we read the news for free.
- deleted 3y ago[deleted]
- at_a_remove 3y agoI think we're suffering from an excluded middle when it comes to this kind of intellectual property. Naturally, most readers want to pay zero. Naturally, owners of the publication think it is probably worth a couple hundred dollars a year to be this well-informed. The current arms race got us scrapers, and then paywalls, and then ad-blocking archivers ... But in reality, I might drop a penny to read a NYT article. Maybe a nickel. There's no reasonable way of performing microtransactions right now. Everything is still in hefty increments, so nobody can work out what the market would bear.