9 ms·
The author has misunderstood when the perplexity user agent applies. Web site owners shouldn’t dictate what browser users can access their site with - whether
by maxrmk 2y ago
The author has misunderstood when the perplexity user agent applies.
Web site owners shouldn’t dictate what browser users can access their site with - whether that’s chrome, firefox, or something totally different like perplexity.
When retrieving a web page _for the user_ it’s appropriate to use a UA string that looks like a browser client.
If perplexity is collecting training data in bulk without using their UA that’s a different thing, and they should stop. But this article doesn’t show that.
- rknightuk 2y agoIt’s not retrieving a web page though is it? It’s retrieving the content then manipulating it. Perplexity isn’t a web browser.
- dewey 2y ago> It’s retrieving the content then manipulating it. Perplexity isn’t a web browser. So a browser with an ad-blocker that's removing / manipulating elements on the page isn't a browser? What about reader mode?
- cdme 2y agoHow a user views a page isn't the same as a startup scraping the internet wholesale for financial gain.
- ulrikrasmussen 2y agoBut it's not scraping, it's retrieving the page on request from the user.
- cdme 2y agoWith no benefit provided to the creator — they're not directing users out, they're pulling data in.
- deleted 2y ago[deleted]
- threecheese 2y agoThey are directing users __in__ in some cases though, no? I’m a perplexity user, and their summaries are often way off which drives me to the references (attribution). The ratio of fetches to clickthroughs is what’s important now though; this new model (which we’ve not negotiated or really asked for) is driving that upward from 1, and not only are you paying more as a provider but your consumer is paying more ($ to perplexity and/or via ad backend) and you aren’t seeing any of it. And you pay those extra costs to indirectly finance the competitor who put you in this situation, who intends to drive that ratio as high as it can in order to get more money from more of your customers tomorrow. Yay.
- JumpCrisscross 2y ago> it's not scraping, it's retrieving the page on request from the user Search engines already tried it. It’s not retrieving on request because the user didn’t request the page, they requested a bot find specific content on any page.
- alexey-salmin 2y agoBut it's not what happened here. It WAS retrieving on request. > I went into Perplexity and asked "What's on this page rknight.me/PerplexityBot?". Immediately I could see the log and just like Lewis, the user agent didn't include their custom user agent
- JumpCrisscross 2y agoThat was to test the user-agent hiding. The broader problem—Perplexity laundering attribution—is where the scraping vs retrieval question comes into play.
- Dylan16807 2y agoWell the example in the post doesn't show any laundering. Do you have an example of it? Unless you mean the entire concept of training launders attribution, but that's basically unrelated to this post and the complaints inside it.
- threecheese 2y agoIn this case you are 100% correct, but I think it’s reasonable to assume that the “read me this web page” use case constitutes a small minority of perplexity’s fetches. I find it useful because of the attribution - more so its references - which I almost always navigate to because its summaries are frequently crap.
- SamBam 2y agoThis is why this conversation is making me insane. How are people saying straight-faced that the user is requesting a specific page? They aren't, they're doing a search of the web. That's not at all the same as a browser visiting a page.
- gruez 2y agoThat's not a relevant factor in most legal regimes. At best it's a moral argument.
- deleted 2y ago[deleted]
- manojlds 2y agoSo if you have a browser that has Greasemonkey like scripts running on it, then it's not a browser? What about AI summary feature available on Edge now?
- LeifCarrotson 2y agoRetrieving the content of a web page then manipulating it is basically the definition of a web browser.
- maxrmk 2y agoI’d consider it a web browser but that’s a vague enough term that I can understand seeing it differently. I’d be disappointed if it became common to block clients like this though. To me this feels like blocking google chrome because you don’t want to show up in google search (which is totally fine to want, for the record). Unnecessarily user hostile because you don’t approve of the company behind the client.
- TeMPOraL 2y agoYes, that's literally why "user agent" is called "user agent". It's a program that acts in place and in the interest of its user, and this in particular always included allowing the user to choose what will or won't be rendered, and how. It's not up to the server what the client does with the response they get.
- JoosToopit 2y agoUA is just a signature a client sends. It's up to the client to use the signature they want to use.
- mattigames 2y agoAnd its up to the client to send as many requests as they see fit, it still called a DDOS attack when overdone regardless of the available freedom that the client has to do it.
- wonnage 2y agoSetting a correct user agent isn't required anyway, you just do it to not be an asshole. Robots.txt is an optional standard. The article is just calling Perplexity out for some asshole behavior, it's not that complicated It's clear they know they're engaging in poor behavior too, they could've documented some alternative UA for user-initiated requests instead of spoofing Chrome. Folks who trust them could've then blocked the training UA but allowed the alternative
- kuschkufan 2y ago[flagged]
- JimDabell 2y agoJust to go a little bit more into detail on this, because the article and most of the conversation here is based on a big misunderstanding: robots.txt governs crawlers. Fetching a single user-specified URL is not crawling. Crawling is when you automatically follow links to continue fetching subsequent pages. Perplexity’s documentation that the article links to describes how their crawler works. That is not the piece of software that fetches individual web pages when a user asks for them. That’s just a regular user-agent, because it’s acting as an agent for the user. The distinction between crawling and not crawling has been very firmly established for decades. You can see it in action with wget. If you fetch a specific URL with `wget https://www.example.com https://www.example.com` then wget will just fetch that URL. It will not fetch robots.txt at all. If you tell wget to act recursively with `wget --recursive https://www.example.com https://www.example.com` to crawl that website, then wget will fetch `https://www.example.com https://www.example.com`, look for links on the page, then if it finds any links to other pages, it will fetch `https://www.example.com/robots.txt https://www.example.com/robots.txt` to check if it is permitted to fetch any subsequent links. This is the difference between fetching a web page and crawling a website. Perplexity is following the very well established norms here.
- mattigames 2y agoIts fairly logical to assume that robots.txt governs robots (empahsis in "bots") not just crawlers, if they are only intended to block crawlers why aren't they called crawlers.txt instead and remove all ambiguity?
- bluish29 2y agoThat's a historical question. At this time, most if not all the bots were either search engines or archival. The name was even "RobotsNotWanted.txt" at the beginning but made "robots.txt" for simplicity. To give another example, Internet Archive stopped respecting it a couple of years ago, and they discuss this point (crawlers vs other bots) here [1]. [1] https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...