4 ms·
> What exactly is Perplexity doing here that isn't okay that people don't already do with their local user agents? It's in the title of TFA: they're being dish
by denton-scratch 2y ago
> What exactly is Perplexity doing here that isn't okay that people don't already do with their local user agents?
It's in the title of TFA: they're being dishonest about who they are. PerplexityBot seems to understand that robots.txt is addressed to it.
It's understood that site operators have a right to use the User-Agent to discriminate among visitors; that's why robots.txt is a standard. Crawlers that disrespect the standard have for many years been considered beyond the pale; thieves and snoopers. TFA's complaint is entirely justified.
- lolinder 2y ago> It's in the title of TFA: they're being dishonest about who they are. PerplexityBot seems to understand that robots.txt is addressed to it. First, I'm ignoring the output of Perplexity. I have no reason to believe that they gave the LLM any knowledge about its internal operations, it's just riffing off of what OP is saying. Second, PerplexityBot is the user agent that they use when crawling and indexing. They never claimed to use that user agent for ad hoc HTTP requests (which are notably not the same as crawling). Third, I disagree that anyone has an obligation to be honest in their User-Agent. Have you ever looked at Chrome's user agent? They're spoofing just about everyone, as is every browser. Crawlers should respect robots.txt, but I'd be totally content if we just got rid of the User-Agent string entirely.
- denton-scratch 2y ago> (which are notably not the same as crawling) Is that a distinction without a difference? I think the robots.txt RFC was addressed specifically to crawlers; so technically "ad hoc" requests generated automatically (i.e. by robots) aren't included. But the distinction operators would like to make is between humans and automata. Whether some automaton is a crawler or not isn't relevant.
- lolinder 2y agoActually, no, the fact that it's a crawler is the most important fact. The reason why website operators care at all about robots accessing their site (as distinct from humans controlling a browser) is historically one of two reasons: * The pattern of requests can be very problematic. Impolite crawlers are totally capable of taking down a website by hitting it over and over and over again for hours in a way that humans won't. * Crawlers are generally used to build search indexes, so instructing them about URLs that would be inappropriate to have show up in a search is relevant. The behavior that OP is complaining about is that when the user pastes a URL into Perplexity, Perplexity fetches that URL. Neither the traffic pattern nor the persistence profile are remotely similar to typical crawler behavior. As far as I can see there's almost nothing to distinguish it from someone using Edge and then using Edge's built-in summarizer.
- Dylan16807 2y agoIf explicitly telling it to access a URL is an access by automaton, then isn't every web browser load an access by automaton?
- BeefWellington 2y agoThe flaw with that example is your web browser isn't between other users and the website, turning 500 views into one. And if we took the analogy to the other end, one could argue that all crawlers have to be kicked off manually at some point... The problem is here in reality the differentiation is somewhat more understood. The honor system web is going away, that's for sure.
- lolinder 2y ago> your web browser isn't between other users and the website, turning 500 views into one. There are a lot of people making this assumption about the way Perplexity is working, but there is no evidence in TFA that Perplexity is caching its ad hoc requests. And even if they were, what's left unsaid is why it even would matter if 500 views turned into one. It matters either because of lost ad revenue or lost ability to track the users' behavior. Personally, I'm okay with moving past that phase of the internet's life and look forward to new business models that aren't built around getting large numbers of "views".
- Dylan16807 2y ago> The flaw with that example is your web browser isn't between other users and the website, turning 500 views into one. So, a caching proxy? That has its own issues, but it's the opposite of access by automaton. One button press causes less than one access to the server. Though one button press still results in one user view, so it's only reducing loads in some ways. But also is that happening here? > And if we took the analogy to the other end, one could argue that all crawlers have to be kicked off manually at some point... One button press causing a million page loads is access by automaton. The distinction seems pretty simple to me.