5 ms·
A lot of comments here are confusing the two use cases for crawling: training and summarization. Perplexity's utility as an answer engine is RAG (retrieval aug
by putlake 2y ago
A lot of comments here are confusing the two use cases for crawling: training and summarization.
Perplexity's utility as an answer engine is RAG (retrieval augmented generation). In response to your question, they search the web, crawl relevant URLs and summarize them. They do include citations in their response to the user, but in practice no one clicks through on the tiny (1), (2) links to go to the source. So if you are one of those sources, you lose out on traffic that you would otherwise get in the old model from say a Google or Bing. When Perplexity crawls your web page in this context, they are hiding their identity according to OP, and there seems to be no way for publishers to opt out of this.
It is possible that when they crawl the web for the second use case -- to collect data for training their model -- they use the right user agent and identify themselves. A publisher may be OK with allowing their data to be crawled for use in training a model, because that use case does not directly "steal" any traffic.
- LeifCarrotson 2y agoGoogle and Bing increasingly do the same thing with their answer box featured snippets.
- int_19h 2y agoThe real question here is whether websites are entitled to that traffic, or even more specifically, to human eyes - and to what extent that should allow them to override users' preferences (which are made fairly clear by the very act of using Perplexity in the first place; the reason why you'd do it instead of doing a Google Search and then manually sifting through the links yourself is because most of what you see is garbage). I would even argue that the whole conversation about AI is a distraction here. Imagine if, instead of using an LLM, Perplexity actually assigned a human agent to your query who'd do the same thing that the model does: write the search queries based on your high-level question, read through the pages that come up, and condense it all into a summary with references to the original sources. That would, of course, be a lot more expensive, but the output would be the same, and so would be the consequences: the person who asked the original high-level question does not get exposed to all the content that had to be waded through to answer it. Is that unethical? If not, then why does replacing the human agent with an AI in this scenario becomes unethical? And if the answer is "scale", that gets uncomfortably close to saying that it's okay for the rich but not for the plebs.
- aspenmayer 2y agoI like your comment a lot, so much so that I replied to it on the top-level in hopes of promoting wider discussion of the points you have raised: https://news.ycombinator.com/item?id=40693140 https://news.ycombinator.com/item?id=40693140
- 627467 2y ago> in practice no one clicks through on the tiny (1), (2) links to go to the source I offer my self as specimen of someone who clicks on those citations ALL the time because thats how I can - most of the time - find download links, and other details faster than asking again