8 ms·
Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplex
by skilled 2y ago
Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it,
https://stackdiary.com/perplexity-has-a-plagiarism-problem/ https://stackdiary.com/perplexity-has-a-plagiarism-problem/
The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away.
[0]: https://www.semafor.com/article/06/12/2024/perplexity-was-planning-revenue-sharing-deals-with-publishers https://www.semafor.com/article/06/12/2024/perplexity-was-pl...
- Mathnerd314 2y agoIt's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).
- skilled 2y agoCan’t wait for OpenAI to settle with The New York Times. For a billion dollars no less.
- brookst 2y agoOnly reason OpenAI would do that would be to create a barrier for smaller entrants.
- JumpCrisscross 2y ago> Only reason OpenAI would do that would be to create a barrier for smaller entrants Only? No. Not even main. The main reason would be to halt discovery and setting a precedent that would fuel not only further litigation but also, potentially, legislation. That said, OpenAI should spin it as that master-of-the-universe take.
- monocasa 2y agoA billion dollar settlement is more than enough to fuel further litigation.
- JumpCrisscross 2y ago> billion dollar settlement is more than enough to fuel further litigation The choice isn’t between a settlement and no settlement. It’s between settlement and fighting in court. Binding precedent and a public right increase the risks and costs to OpenAI, particularly if it looks like they’ll lose.
- monocasa 2y agoRight, but a billion dollars to a relative small fry in the publishing industry (even online only) like the ny times is chum in the water. The next six publishers are going to be looking for $100B and probably have the funds for better lawyers. At some point these are going to hit the courts, an NY Times probably makes sense as the plaintiff as opposed to one of the larger publishing houses.
- JumpCrisscross 2y ago> ny times is chum in the water The Times has a lauder litigation team. Their finances are good and their revenue sources diverse. They’re not aching to strike a deal. > NY Times probably makes sense as the plaintiff as opposed to one of the larger publishing houses Why? Especially if this goes to a jury.
- sebzim4500 2y agoSettling for a billion dollars would be insane. They'd immediately get sued by everyone who ever posted anything on the internet.
- insane_dreamer 2y agoI, on the other hand, hope NYT refuses a settlement and OpenAI loses in court.
- skilled 2y agoSame, for sure!
- int_19h 2y agoBe careful what you wish for, because, depending on how broad the reasoning in such a decision would be, it is not impossible that the precedent would be used to then target ad blockers and similar software.
- insane_dreamer 2y agoFair point, but it's a risk I'd be willing to take.
- brookst 2y agoIf using copyrighted material to train an LLM is theft, so is reading a book.
- bakugo 2y agoHow is a human reading a book in any way related or comparable to a machine ingesting millions of books per day with the goal of stealing their content and replacing them?
- ysofunny 2y agoit's comparable exactly in the way 0.001% can be compared to 10^100 humans learning is the old-school digital copying. computers simply do it much faster, but it's the same basic phenomenon consider one teacher and one student. first there is one idea in one head but then the idea is in two heads. now add book technology1 the teacher writes the book once, a thousand students read it. the idea has gone from being in one head (book author) onto most of the book readers!
- somenameforme 2y ago> humans learning is the old-school digital copying. computers simply do it much faster, but it's the same basic phenomenon Train an LLM on the state of human knowledge 100,000 years ago - language had yet to be invented and bleeding edge technology was 'poke them with the pointy side.' It's not going to be able to do or output much of anything, and it's going to be stuck in that state for perpetuity until somebody gives it something new to parrot. Yet somehow humans went from that exact starting to state to putting a man on the Moon. Human intelligence, and elaborate auto-complete systems, are not the same thing, or even remotely close to the same thing.
- Terr_ 2y ago> bleeding edge technology was 'poke them with the pointy side.' Relevant: https://www.smbc-comics.com/comic/rise-of-the-machines https://www.smbc-comics.com/comic/rise-of-the-machines
- JumpCrisscross 2y ago> pending lawsuits notwithstanding That’s a hell of a caveat!
- AlienRobot 2y agoI'd believe it if they were targeting entities that could fight back, like stock photo companies and disney, instead of some guy with an artstation account, or some guy with a blog. To me it sounds like these products can't exist without exploiting someone and they're too coward to ask for permission because they know the answer is going to be "no." Imagine how many things I could create if I just stole assets from others instead of having to deal with pesky things like copyright!
- Pannoniae 2y ago...which is a great argument for abolishing copyright:P
- AlienRobot 2y ago...which is a great argument for how unjust is a law that only protects those that can afford it. Cheaper processes to protect smaller creators in cases like these is what is really needed.
- lolinder 2y ago> so far getting away with it (pending lawsuits notwithstanding). I know it feels like it's been longer, but it's not even been 2 years since ChatGPT was released. "So far" is in fact a very short amount of time in a world where important lawsuits like this can take 11 years to work their way through the courts [0]. [0] https://en.m.wikipedia.org/wiki/Oracle_v_Google https://en.m.wikipedia.org/wiki/Oracle_v_Google
- emporas 2y agoIn 9 years time, robots will publish articles on the web, and they will put a humans.txt file at their root index to govern what humans are allowed to read the content. Jokes aside, given how models become better, cheaper and smaller, RAG classification and filtering engines like Perplexity will become so ubiquitous that i don't see any way for a website owner to force anyone to visit the website anymore.
- twinge 2y agoAereo, Napster, Grokster, Grooveshark, Megaupload, and TVEyes: they all thought the same thing. Where are they now?
- losvedir 2y agoHeh, you're right, of course, but as someone who came of age on the internet around that era, it still seems strange to me that people these days are making the arguments the RIAA did. They were the big bad guys in my day.
- lofaszvanitt 2y agoThey were massacred by well funded corps. Who is on the side of single joes?
- s3r3nity 2y ago[flagged]
- oaththrowaway 2y agoWhat indie game dev shut down because of piracy?
- FireInsight 2y agoSomething not being stealing isn't the same as it not being able to hurt people or companies financially. Revenue lost due to copyright breach is not money stolen from you. I pay my indie creators fairly; big companies is when I stop caring.
- cyanydeez 2y agoRight, it's ironic we spent 30 years fighting piracy and then suddenly corporations start doing it and now it's suddenly ok.
- ben_w 2y agoFor me, the irony is the opposite side of the same coin, 30 years of "information wants to be free" and "copyright infringement isn't piracy" and "if you don't want to be indexed, use robots.txt"… …and then suddenly OpenAI are evil villains, and at least some of the people denounced them for copyright infringement are, in the same post, adamant that the solution is to force the model weights to become public domain.
- bee_rider 2y agoThe deal of the internet has always been: send me what you want and I’ll render it however I want. This includes feeding it into AI bots now. I don’t love being on the same side as these “AI” snakeoil salesmen, but they are following the rules of the road. Robots.txt is just a voluntary thing. We’re going to see more and more of the internet shut off by technical means instead, which is a bummer. But on the bright side it might kill off the ad based model. Silver linings and all that.
- int_19h 2y agoI broadly agree with you, but I don't see what's contradictory about the solution of model weights becoming public domain. When it comes to piracy, the people who have viewed it as ethical on the grounds that "information wants to be free" generally also drew the line at profiting from it: copying an MP3 and giving it to your friend or even a complete stranger is ethical, charging a fee for that (above and beyond what it costs you to make a copy) is not. From that perspective, what OpenAI is doing is evil not because they are infringing on everyone's copyright, but that they are profiting from it.
- ben_w 2y agoTo me, it's like trying to "solve The Pirate Bay" by making all the stuff they share public domain. But thank you for sharing your perspective, I appreciate that.
- more_corn 2y agoI hate to argue this side of the fence, but when ai companies are taking the work of writers and artists en mass (replacing creative livelihoods with a machine trained on the artists stolen work) and achieving billion dollar valuations that’s actual stealing. The key here is that creative content producers are being driven out of business through non consensual taking of their work. Maybe it’s a new thing, but if it is, it’s worse than stealing.
- bongodongobob 2y agoI cannot imagine how viewing/scraping a public website could ever be illegal, wrong, immoral etc. I just don't see the argument for it.
- ronsor 2y agoAI hysteria has made everyone lose their minds over normal things.
- tucnak 2y agoI guess people just LOVE twisting themselves in knots over some "ethical scandals" or whatnot. Maybe there's a statement on American puritanism hiding somewhere here...
- insane_dreamer 2y agoIt's scraping content to then serve up that content to users who can now get that content from you (via a paid subscription service, or maybe ad-sponsored) instead of visiting the content creator and paying them (i.e., via ads on their website) It's the same reason I can't just take NYT archives or the Britannica and sell an app that gives people access to their content through my app. It totally undercuts content creators, in the same way that music piracy -- as beloved as it was, and yeah, I used Napster back in the day -- took revenue away from artists, as CD sales cratered. That gave birth to all-you-can-eat streaming, which does remunerate artists but nowhere near what they got with record sales.
- insane_dreamer 2y agoOne more point on this, lest some people think, "hey Kanye, or Taylor Swift, don't need any more money!" I 100% agree. But the problem with streaming is that is disproportionately rewards the biggest artists at the expense of the smaller ones. It's the small artist, barely making a living from their craft, who were most hurt by the switch from albums to streaming, not those making millions.
- 2y ago
- dspillett 2y ago> piracy is not theft Correct, but it is often a licensing breach (though sometimes depending upon the reading of some licenses, again these things are yet to be tested in any sort of court) and the companies doing it would be very quick to send a threatening legal letter if we used some of their output outside the stated licensing terms.
- losvedir 2y agoYou wouldn't train an LLM on a car.
- insane_dreamer 2y ago> piracy is not theft it was when Napster was doing it; but there's no entity like the RIAA to stop the AI bots
- readyman 2y ago>and thinking they can get away with it Can they not? I think that remains to be seen.
- jhbadger 2y agoExactly. It's like when Uber started and flaunted the medallion taxi system of many cities. People said "These Uber people are idiots! They are going to get shut down! Don't they know the laws for taxis?" While a small number of cities did ban Uber (and even that generally only temporarily), in the end Uber basically won. I think a lot of people confuse what they want to happen versus what will happen.
- readyman 2y agoAmericans are incredibly ignorant of how the world actually works because the American living memory only knows the peak of the empire from the inside.
- seanhunter 2y agoIn London, uber did not succeed. Uber drivers have to be licensed like minicab drivers.
- jhbadger 2y agoPerhaps. But a reasonable license requiring you to pass a test isn't the same as a medallion in the traditional American taxi system. Medallions (often costing tens or even hundreds of thousands of dollars) were a way of artificially reducing the number of taxis (and thus raising the price).
- itissid 2y agoThis. Medallion systems in NYC were gamed by a guy who let people literally bet on its as if it were an asset. The prices went to a million per until the bubble burst. True story
- 2y ago
- zxxh 2y ago[dead]