10 ms·
It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly
by putlake 2y ago
It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy.
But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content (via Google, for example). This is unacceptable. A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content.
There are many benefits to having people visit your content on a property that you own. e.g., say you are a SaaS company and you have a bunch of Help docs. You can analyze traffic in this section of your website to get insights to improve your business: what are the top search queries from my users, this might indicate to me where they are struggling or what new features I could build. In a world where users ask Perplexity these Help questions about my SaaS, Perplexity may answer them and I would lose all the insights because I never get any traffic.
- epolanski 2y ago> they are decreasing the probability that this user would come to by content (via Google, for example). Google has been providing summaries of stuff and hijacking traffic for ages. I kid you not, in the tourism sector this has been a HUGE issue, we have seen 50%+ decrease in views when they started doing it. We paid gazzilions to write quality content for tourists about the most different places just so Google could put it on their homepage. It's just depressing. I'm more and more convinced that the age of regulations and competition is gone, US does want to have unkillable monopolies in the tech sector and we are all peons.
- itsoktocry 2y ago>We paid gazzilions to write quality content for tourists about the most different places just so Google could put it on their homepage. It's just depressing It's a legitimate complaint, and it sucks for your business. But I think this demonstrates that the sort of quality content you were producing doesn't actually have much value.
- OrigamiPastrami 2y agoI'd argue it only demonstrates that it doesn't produce much value for the creator.
- lolinder 2y agoThe Google summaries (before whatever LLM stuff they're doing now) are 2-3 sentences tops. The content on most of these websites is much, much longer than that for SEO reasons. It sucks that Google created the problem on both ends, but the content OP is referring to costs way more to produce than it adds value to the world because it has to be padded out to show up in search. Then Google comes along and extracts the actual answer that the page is built around and the user skips both the padding and the site as a whole. Google is terrible, the attention economy that Google created is terrible. This was all true before LLMs and tools like Perplexity are a reaction to the terrible content world that Google created.
- newaccount74 2y agoIt would be a lot better if Google just prioritised concise websites. If Google preferred websites that cut the fluff, then website operators would have an incentive to make useful websites, and Google wouldn't have as much of an incentive to provide the answer in a snippet, and everyone wins. I guess it's hard to rank website quality, so Google just prefers verbose websites.
- TeMPOraL 2y ago> Google wouldn't have as much of an incentive to provide the answer in a snippet, and everyone wins. Google has at least two incentives to provide that answer, both of which wouldn't change. The bad one: they want to keep you on their page too, for usual bullshit attention economy reasons. The good one: users prefer the snippets too. The user searching for information usually isn't there to marvel at beauty of random websites hiding that information in piles of noise surrounded by ads. They don't care about websites in the first place. They want an answer to the question, so they can get on with whatever it is they're doing. When Google can give them an answer, and this stops them from going from SERP to any website, then that's just few seconds or minutes of life that user doesn't have to waste. Lifespans are finite.
- CobrastanJorji 2y agoI'm curious about the tourism sector problem. In tourism, I would think the goal would be to promote a location. You want people to be able to easily discover the location, get information about it, and presumably arrange to travel to those locations. If Google gets the information to the users, but doesn't send the tourist to the website, is that harmful? Is it a problem of ads on the tourism website? Or is more of problem of the site creator demonstrating to the site purchaser that the purchase was worthwhile?
- SamBam 2y agoPresumably the issue is more the travel guides/Time Out/Tripadvisor type websites. They make money by you reading their stuff, not by you actually spending money in the place.
- epolanski 2y agoWe would employ local guides all around the world to craft itinerary plans to visit places, give tips, tricks, recommend experiences and places (we made money by selling some of those through our website) and it was a success. Customers liked the in depth value of that content and it converted to buys (we sold experiences and other stuff, sort of like getyourguide). One day all of our content ended up on Google "what time is best to visit the Sagrada Familia" and you would have a copy pasted answer by Google. This killed a lot of traffic. Anyway, I just wanted to point out that the previous user was a bit naive taking his fight to LLMs when search engines and OSs have been leeching and hijacking content for ages.
- popalchemist 2y agoIf your content has a yes/no or otherwise simple, factual answer that can be conveyed in a 1-2 sentence summary, then I don't see this as a problem. You need to adapt your content strategy, as we all do from time to time. There was never a guarantee -- for anyone in any industry at all -- that what worked in the past will always continue to work. That is a regressive attitude. However I do have concerns about Google and other monopolies replacing large swaths of people who make their livings doing things that can now be automated. I am not against automation but I don't think the disruption of our entire societal structure and economy should be in the hands of the sociopaths that run these companies. I expect regulation to come into play once the shit hits the fan for more people.
- Brybry 2y agoWhile I personally believe it should be opt-in, Google does have multiple ways to opt out of snippets while still being indexed. [1] [1] https://developers.google.com/search/docs/appearance/snippet#nosnippet https://developers.google.com/search/docs/appearance/snippet...
- jcynix 2y ago> Google has been providing summaries of stuff and hijacking traffic for ages. Yes, Google hijacked images for some time. But in general there has "always" been the option to tell Google not to display summaries etc with these meta tags: <meta name="googlebot" content="noarchive"> <meta name="googlebot" content="nosnippet">
- bigethan 2y agoThe problem with opt-out is that if the whole industry doesn't do it, google will show someone's summaries and the traffic still stays with them.
- Too 2y agoGoogle has been in trouble for doing so several times in the past and removed key features because of it. Examples: Viewing cached pages, linking directly to images, summarized news articles.
- briantakita 2y ago> But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content (via Google, for example). Perplexity has source references. I find myself visiting the source references. Especially to validate the LLM output. And to learn more about the subject. Perplexity uses a Google search API to generate the reference links. I think a better strategy is to treat this as a new channel to receive visitors. The browsing experience should be improved. Mozilla had a pilot called Context Graph. Perhaps Context Graph should be revisited? > In a world where users ask Perplexity these Help questions about my SaaS, Perplexity may answer them and I would lose all the insights because I never get any traffic. This seems like a missing feature for analytics products & the LLMs/RAGs. I don't think searching via an LLM/RAG is going away. It's too effective for the end user. We have to learn to work with it the best we can.
- TeMPOraL 2y ago>> In a world where users ask Perplexity these Help questions about my SaaS, Perplexity may answer them and I would lose all the insights because I never get any traffic. Alternative take: Perplexity is protecting users' privacy by not exposing them to be turned into "insights" by the SaaS. My general impression is that the subset of complaints discussed in this thread and in the article, boils down to a simple conflict of interest: information supplier wants to exploit the visitor through advertising, upsells, and other time/sanity-wasting things; for that, they need to have the visitor on their site. Meanwhile, the visitors want just the information without the surveillance, advertising and other attention economy dark/abuse patterns. The content is the bait, and ad-blockers, Google's instant results, and Perplexity, are pulling that bait off the hook for the fish to eat. No surprise fishermen are unhappy. But, as a fish, I find it hard to sympathize.
- lolinder 2y ago> A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content. If I visit your site from Google with my browser configured to go straight to Reader Mode whenever possible, is my visit more useful to you than a summary and a link to your site provided by Perplexity? Why does it matter so much that visitors be directly on your content?
- gpm 2y agoWell for one thing you visiting his site and displaying it via reader mode doesn't remove his ability to sell paid licenses for his content to companies that would like to redistribute his content. Meanwhile having those companies do so for free without a license obviously does.
- lolinder 2y agoShould OP be allowed to demand a license for redistribution from Orion Browser [0]? They make money selling a browser with a built-in ad blocker. Is that substantially different than what Perplexity is doing here? [0] https://kagi.com/orion/ https://kagi.com/orion/
- gpm 2y agoOrion browser presuming it does what does what it's name says it does doesn't redistribute anything... so presumably not.
- lolinder 2y agoI asked you this in the other subthread, but what exactly is the moral distinction (I'm not especially interested in the legal one here because our copyright law is horribly broken) between these two scenarios? * User asks proprietary web browser to fetch content and render it a specific way, which it does * User asks proprietary web service to fetch content and render it a specific way, which it does The technical distinction is that there's a network involved in the second scenario. What is the moral distinction?
- insane_dreamer 2y agoThis is why media publishers went behind paywalls to get away from Google News
- stqism 2y agoIronically, I’ve just started asking LLMs to summarize paywalled content, and if it doesn’t answer my question I’ll check web archives or ask it for the full articles text.
- antoniojtorres 2y agoThis hits the point exactly, it’s an extension of stuff like Google’s zero click results, they are regurgitating a website’s content with no benefit to the website. I would say though, it feels like the training argument may ultimately lead to a similar outcome, though it’s a bit more ideological and less tangible than regurgitating the results of a query. Services like chatgpt are already being used a google replacement by many people, so long term it may reduce clicks from search as well.
- SpaghettiCthulu 2y agoYou're missing the part where Perplexity still makes a request each time it's asked about the URL. You still get the traffic!
- danlitt 2y agoI'm not sure what you mean exactly. If Perplexity is actually doing something with your article in-band (e.g. downloading it, processing it, and present that processed article to the user) then they're just breaking the law. I've never used that tool (and don't plan to) so I don't know. If they just embed the content in an iframe or something then there's no issue (but then there's no need or point in scraping). If they're just scraping to train then I think you also imply there's no issue. If they're just copying your content (even if the prompt is "Hey Perplexity, summarise this article <ARTICLE_TEXT>") then that's vanilla infringement, whether they lie about their UA or not.
- lompad 2y agoSure it is, but which of the many small websites are going to be able to fight them legally? Most companies would go broke before getting a ruling. Reality is, the law doesn't matter if you're big enough. As long as they're not stealing content from the big ones, they're going to be fine.
- danlitt 2y agoWell, I guess what I mean is if the situation is as I describe in my previous comment, then anyone who did have the money to fight it would be a shoe-in. It's a much stronger case than, for example, the ongoing lawsuits by Matthew Butterick and others (https://llmlitigation.com/ https://llmlitigation.com/).
- lompad 2y agoThanks for the link, that's fantastic to hear! I'm seriously sick of that whole "laundering copyright via AI"-grift - and the destruction of the creative industry is already pretty noticable. All the creatives who brought us all those wonderful masterworks with lots of thought and talent behind, they're all going bankrupt and getting fired right now. It's truly a tragedy - the loss of art is so much more serious than people seem to think it is, considering how integral all kinds of creative works are to a modern human live. Just imagine all of that being without any thought, just statistically optimized for enjoyment... ugh.
- rcthompson 2y agoI don't know what the typical usage pattern is, but when I've used Perplexity, I generally do click the relevant links instead of just trusting Perplexity's summary. I've seen plenty of cases where Perplexity's summary says exactly the opposite of the source.
- anileated 2y ago> I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content (via Google, for example). This is unacceptable. This appears to be self-contradictory. If you let an LLM to be trained* on “all the books” (posts, articles, etc.) in the world, the implication is that your potential readers will now simply ask that LLM. Not only will they pay Microsoft for that privilege, while you would get zilch, but you would not even know they ever read the fruits of your research. * Incidentally, thinking of information acquisition by an ML model as if it was similar to human reading is a problematic fallacy.
- richardatlarge 2y agoI’m not sure if this is relevant but i go to a lot of sites because perplexity has it noted in its answer