3 ms·
Then they should just say that outright instead of pretending they right thing.
by rknightuk 2y ago
Then they should just say that outright instead of pretending they right thing.
- lolinder 2y agoThey're not lying, you just misunderstood their docs [0]. > To provide the best search experience, we need to collect data. We use web crawlers to gather information from the internet and index it for our search engine. > You can identify our web crawler by its user agent To anyone who's familiar with web crawling and indexing, these paragraphs have an obvious meaning: Perplexity has a search engine which needs a crawler which crawls the internet. That crawler can be identified by the User-Agent PerplexityBot and will respect robots.txt. Separately, if you give Perplexity a specific URL then it will go fetch the contents of that URL with a one-off request. That one-off request does not respect robots.txt any more than curl does, and that's 100% normal and ethical. The one-off request handler isn't PerplexityBot, it's a separate part of the application that's probably just a regular Chrome browser that issues the request. [0] https://docs.perplexity.ai/docs/perplexitybot https://docs.perplexity.ai/docs/perplexitybot