33 ms·
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
- thoroughburro 1y ago[flagged]
- tomhow 1y ago> I genuinely hope you feel shame. I would shun you in real life. HN is not a platform for attacking people, even imagined ones. Please don't fulminate. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- gruez 1y ago>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawling (ie. systematically viewing every page on the site without the direction of a human), or simply retrieving content on behalf of the user. I think most people would draw a distinction between the two, and would at least agree the latter is more acceptable than the former.
- thoroughburro 1y ago> I think most people would draw a distinction between the two, and would at least agree the latter is more acceptable than the former. No. I should be able to control which automated retrieval tools can scrape my site, regardless of who commands it. We can play cat and mouse all day, but I control the content and I will always win: I can just take it down when annoyed badly enough. Then nobody gets the content, and we can all thank upstanding companies like Perplexity for that collapse of trust.
- gkbrk 1y ago> Then nobody gets the content, and we can all thank upstanding companies like Perplexity for that collapse of trust. But they didn't take down the content, you did. When people running websites take down content because people use Firefox with ad-blockers, I don't blame Firefox either, I blame the website.
- Bluescreenbuddy 1y agoFF isn’t training their money printer with MY data. AI scrapers are
- glenstein 1y ago>But they didn't take down the content, you did. That skips the part about one party's unique role in the abuse of trust.
- hombre_fatal 1y agoTaking down the content because you're annoyed that people are asking questions about it via an LLM interface doesn't seem like you're winning. It's also a gift to your competitors. You're certainly free to do it. It's just a really faint example of you being "in control" much less winning over LLM agents: Ok, so the people who cared about your content can't access it anymore because you "got back" at Perplexity, a company who will never notice.
- ipaddr 1y agoIt could be my server keeps going down because of llms agents keep requesting pages from my lyric site. Removing that site allowed other sites to remain up. True story. Who cares if perplexity will never notice. Or competitors get an advantage. It is a negative for users using perplexity or visiting directly because the content doesn't exist. That's the world perplexity and others are creating. They will be able to pull anything from the web but nothing will be left.
- IncreasePosts 1y agoYou don't win, because presumably you were providing the content for some reason, and forcing yourself to take it down is contrary to whatever reason that was in the first place.
- fluidcruft 1y agoIf the AI archives/caches all the results it accesses and enough people use it, doesn't it become a scraper? Just learn off the cached data. Being the man-in-the-middle seems like a pretty easy way to scrape salient content while also getting signals about that content's value.
- JimDabell 1y agoNo. The key difference is that if a user asks about a specific page, when Perplexity fetches that page, it is being operated by a human not acting as a crawler. It doesn’t matter how many times this happens or what they do with the result. If they aren’t recursively fetching pages, then they aren’t a crawler and robots.txt does not apply to them. robots.txt is not a generic access control mechanism, it is designed solely for automated clients.
- sbarre 1y agoI would only agree with this if we knew for sure that these on-demand human-initiated crawls didn't result in the crawled page being added to an overall index and scheduled for future automated crawls. Otherwise it's just adding an unwilling website to a crawl index, and showing the result of the first crawl as a byproduct of that action.
- fluidcruft 1y agoMany people don't want their data used for free/any training. AI developers have been so repeatedly unethical that the well-earned Baysian prior is high probability that you cannot trust AI developers to not cross the training/inference streams.
- JimDabell 1y ago> Many people don't want their data used for free/any training. That is true. But robots.txt is not designed to give them the ability to prevent this.
- 1y ago
- deleted 1y ago[deleted]
- a2128 1y agoIn theory retrieving a page on behalf of a user would be acceptable, but these are AI companies who have disregarded all norms surrounding copyright, etc. It would be stupid of them not to also save contents of the page and use it for future AI training or further crawling
- zarzavat 1y agoIf you allow Googlebot to crawl your website and train Gemini, but you don't allow smaller AI companies to do the same thing, then you're contributing to Google's hegemony. Given that AI is likely to be an increasingly important part of society in the future, that kind of discrimination is anti-social. I don't want a future where everything is run by Google even more than it currently is. Crawling is legal. Training is presumably legal. Long may the little guys do both.
- dgreensp 1y agoGooglebot respects robots.txt. And Google doesn't use the fetched data from users of Chrome to supplement their search index (as a2128 is speculating that Perplexity might do when they fetch pages on the user's behalf).
- foota 1y agoYes, but there's no way to say "allow indexing for search, but not for AI use", right?
- warkdarrior 1y agoBut there is: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers https://developers.google.com/search/docs/crawling-indexing/... There is an user agent for search that you can control in robots.txt. user-agent: Googlebot There is another user agent for AI training. user-agent: Google-Extended
- 1y ago
- throwanem 1y agoThe HTTP spec draws such a distinction, albeit implicitly, in the form (and name) of its concept of "user agent."
- alexey-salmin 1y agoOver time it degraded into declaring compatibility with a bunch of different browser engines and doesn't reflect the actual agent anymore. And very likely Perplexity is in fact using a Chrome-compatible engine to render the page.
- throwanem 1y agoThe header to which you refer was named for the concept.
- Tokumei-no-hito 1y agouser agent = which bullshit css hacks and js polyfills will be needed
- busymom0 1y agoThe examples the article cites seem to me that they are merely retrieving content on behalf of the user. I do not see a problem with this.
- ojosilva 1y agoSounds like an ad for Perplexity. They do end up looking bad out of Cloudflare's report, who are the "good guys" in this story - btw Cloudflare's been very pushy lately with their we'll save the web, content independence day marketspeak. But deep in the back of my head, Cloudflare's goodwill elevates Perplexity cunning habilities (assuming they're the culprit since no real evidence, only heresay is in the OP), both companies look like titans fighting, which ends up being positive for Perplexity, at least in the inflated perception of their firepower... if that makes any sense.
- CaliforniaKarl 1y agoSounds like an ad for OpenAI, since Cloudflare reported how OpenAI is "following the rules". Personally, I'm now less interested in using Perplexity, and more interested in using an OpenAI product.
- thunkshift1 1y agoSounds like ad for cloudflare. Didn’t they announce a month ago they will protect websites from llm content sweep? And now they realize they cannot deliver on that promise. We did it correctly but these guys are doing it illegal way! That’ll be 14.99 per month btw..
- snowwrestler 1y ago> Specifically it's unclear on whether Perplexity was crawling (ie. systematically viewing every page on the site without the direction of a human), or simply retrieving content on behalf of the user. Like most AI companies, Perplexity has established user agent strings for both these cases, and the behavior that Cloudflare is calling out does not use either. It pretends to be a person using Chrome on MacOS.
- fxtentacle 1y agoI find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into modifying the software you run locally. 3. If I now go one step further and use an LLM to summarize content because the authentic presentation is so riddled with ads, JavaScript, and pop-ups, that the content becomes borderline unusable, then why would the LLM accessing the website on my behalf be in a different legal category as my Firefox web browser accessing the website on my behalf?
- account42 1y agoIt's quite easy to solve. Hold companies legally accountable for computer fraud and abuse. The problem is that those in the position to do that are not interested.
- Beijinger 1y agoHow about I open a proxy, replace all ads with my ads, redirect the content to you and we share the ad revenue?
- fxtentacle 1y agoThat's somewhat antisocial, but perfectly legal in the US. It's called PayPal Honey, for example, and has been running for 13 years now.
- rustc 1y agoSince when does PayPal Honey replace ads on websites? > PayPal Honey is a browser extension that automatically finds and applies coupon codes at checkout with a single click.
- galaxy_gas 1y ago
- nnx 1y agoI do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? It’s a different UI, sure, but there should be no discrimination towards it as there should be no discrimination towards, say, Links terminal browser, or some exotic Firefox derivative.
- deleted 1y ago[deleted]
- ploynog 1y agoBeing daft on purpose? I haven't heard that using an alternative browser suddenly increases the traffic that a user generates by several orders of magnitude to the point where it can significantly increase hosting cost. A web scraper on the other hand easily can and they often account for the majority of traffic especially on smaller sites. So your comparison is at least naive assuming good intentions or malicious if not.
- magicmicah85 1y agoA crawler intends to scrape the content to reuse for its own purposes while a browser has a human being using it. There's different intents behind the tools.
- JimDabell 1y agoCloudflare asked Perplexity this question: > Hello, would you be able to assist me in understanding this website? https:// https:// […] .com/ In this case, Perplexity had a human being using it. Perplexity wasn’t crawling the site, Perplexity was being operated by a human working for Cloudflare.
- gruez 1y ago>I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? AI broke the brains of many people. The internet isn't a monolith, but prior to the AI boom you'd be hard pressed to find people who were pro-copyright (except maybe a few who wanted to use it to force companies to comply with copyleft obligations), pro user-agent restrictions, or anti-scraping. Now such positions receive consistent representation in discussions, and are even the predominant position in some places (eg. reddit). In the past, people would invoke principled justifications for why they opposed those positions, like how copyright constituted an immoral monopoly and stifled innovation, or how scraping was so important to interoperability and the open web. Turns out for many, none of those principles really mattered and they only held those positions because they thought those positions would harm big evil publishing/media companies (ie. symbolic politics theory). When being anti-copyright or pro-scraping helped big evil AI companies, they took the opposite stance.
- bbqfog 1y agoIf you put info on the web, it should be available to everyone or everything with access.
- deleted 1y ago[deleted]
- TechDebtDevin 1y agoNot according to CF. They are desperate to turn web sites into newspaper dispensers, where you should give them a quarter to see the content, on the basis that a bot is somehow different than a normal human vistor o a legal basis. Cf has been trying this psyop for years.
- ectospheno 1y agoSites aren’t getting ad clicks for this traffic. Thus, they have an incentive to do something. Cloudflare is just responding to the market. Is this response bad for us in the long run? Probably. Screaming about cloudflare isn’t going to change the market. You fix a problem with capitalism by using supply and demand levers. Everything else is folly.
- TechDebtDevin 1y agoI wonder if crawlers started letting ads through, and interacting with them a bit, if these complaints would go away. If we can just shaft the advertisers, maybe that will solve the whole problem :)
- Workaccount2 1y agoWhat this actually translates to is "Don't bother putting much effort into web content. Put effort into siloed mobile app content where you get compensation". People like getting money for their work. You do too. Don't lose sight of that.
- 9cb14c1ec0 1y agoEven for AI summaries that leech off your content without sending any traffic your direction?
- TechDebtDevin 1y agoCloudflare screaming into the void desperate to insert themselves as a middleman, in a market ( that they will never succeed in creating) where they extort scrapers for access to websites they cover. Sorry CF, give up. the courts are on our sides here
- morkalork 1y agoAre you sure? I'm surprised they haven't jumped in on the "scan your face to see the webpage" madness that's taking off around the world
- sbarre 1y agoWhich courts exactly? The world is bigger than the USA. Just because American tech giants have captured and corrupted legislators in the US doesn't mean the rest of the world will follow.
- JimDabell 1y agoTheir test seems flawed: > We created multiple brand-new domains, similar to testexample.com and secretexample.com. These domains were newly purchased and had not yet been indexed by any search engine nor made publicly accessible in any discoverable way. We implemented a robots.txt file with directives to stop any respectful bots from accessing any part of a website: > We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains. This response was unexpected, as we had taken all necessary precautions to prevent this data from being retrievable by their crawlers. > Hello, would you be able to assist me in understanding this website? https:// https:// […] .com/ Under this situation Perplexity should still be permitted to access information on the page they link to. robots.txt only restricts crawlers. That is, automated user-agents that recursively fetch pages: > A robot is a program that automatically traverses the Web's hypertext structure by retrieving a document, and recursively retrieving all documents that are referenced. > Normal Web browsers are not robots, because they are operated by a human, and don't automatically retrieve referenced documents (other than inline images). — https://www.robotstxt.org/faq/what.html https://www.robotstxt.org/faq/what.html If the user asks about a particular page and Perplexity fetches only that page, then robots.txt has nothing to say about this and Perplexity shouldn’t even consider it. Perplexity is not acting as a robot in this situation – if a human asks about a specific URL then Perplexity is being operated by a human. These are long-standing rules going back decades. You can replicate it yourself by observing wget’s behaviour. If you ask wget to fetch a page, it doesn’t look at robots.txt. If you ask it to recursively mirror a site, it will fetch the first page, and then if there are any links to follow, it will fetch robots.txt to determine if it is permitted to fetch those. There is a long-standing misunderstanding that robots.txt is designed to block access from arbitrary user-agents. This is not the case. It is designed to stop recursive fetches. That is what separates a generic user-agent from a robot. If Perplexity fetched the page they link to in their query, then Perplexity isn’t doing anything wrong. But if Perplexity followed the links on that page, then that is wrong. But Cloudflare don’t clearly say that Perplexity used information beyond the first page. This is an important detail because it determines whether Perplexity is following the robots.txt rules or not.
- 1gn15 1y ago
- throw_m239339 1y ago> How can you protect yourself? Put your valuable content behind a paywall.
- b0ner_t0ner 1y agoA combination of "Bypass Paywalls Clean for Firefox" and archive.is usually get past these.
- schmorptron 1y agoIsn't that only because they offer unpaywalled versions to web crawlers in the first place, so they still get ranked in search results?
- zzo38computer 1y agoI would want to see a search engine which will not index paywalled articles and will not execute JavaScripts and CSS (so that any part of the document that requires JavaScripts to be displayed will not be indexed, and search queries involving such parts of the documents will not find those documents; documents that require JavaScripts to be displayed at all will not be indexed at all).
- binarymax 1y agoI've built and run a personal search engine, that can do pretty much what perplexity does from a basic standpoint. Testing with friends it gets about 50/50 preference for their queries vs Perplexity. The engine can go and download pages for research. BUT, if it hits a captcha, or is otherwise blocked, then it bails out and moves on. It pisses me off that these companies are backed by billions in VC and they think they can do whatever they want.
- metadat 1y agoThis sounds fascinating! Are you able to elaborate on what is different about yours vs Perp'lexity's?
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- kissgyorgy 1y agoNot sure I would consider a user copy-pasting an URL being a bot. Should curl be considered a bot too? What's the difference?
- rwmj 1y agoIn unrelated news, Fedora (the Linux distro) has been taken down by a DDoS today which I understand is AI-scraping related: https://pagure.io/fedora-infrastructure/issue/12703 https://pagure.io/fedora-infrastructure/issue/12703
- st3fan 1y agoThe last comment there now reads: "It was actually a caching issue on our end. ;) I just fixed it a few min ago..." Lets not go on a witch hunt and blame everything on AI scrapers.
- ChocolateGod 1y agoHow many requests are LLMs typically making in order for people to accuse them of doing a DoS attack?
- larodi 1y agoGood they do it. Facebook took TBs of data to train, nobody knows what Goog does to evade whatever they want. the service is actually very convenient no matter faang likes it or not.
- klabb3 1y agoUnexpected underdog argument. What is happening in reality is all companies are racing to (a) scrape, buy and collect as much as they can from others, both individuals and companies while (b) locking down their own data against everyone else who isn’t directly making them money (eg through viewing their ads). Part of me thinks that the open web has a paradox of tolerance issue, leading to a race to the bottom/tragedy of the commons. Perhaps it needs basic terms of use. Like if you run this kind of business, you can build it on top of proprietary tech like apps and leave the rest of us alone.
- larodi 1y agoWe need to wake up and understand that all the information already uploaded is more or less a free web material, once taken through the lens of ML-somethings. With all the second, and third-order effects such as the fact that this changes completely the whole motivation, and consequence of open-source perhaps. It is also only a matter of time scrapers once again get through walls by twitter, reddit and alike. This is, after all, information everyone produced, without being aware of it was now considered not theirs anymore.
- ipaddr 1y agoReddit sold their data already. Twitter made thier own AI.
- larodi 1y agoPrecisely my point, and there is little if any evidence, there is anyone among the big players who puts peoples' rights before else by respecting licensing agreements before scrapping for training. Indeed, Reddit sold their data the other thay GPT2 was announced, and it was very apparent why everyone closed their APIs in 2021-2023. Wonder what Aaron would've said about it. Now we have walled gardens of information where people are allowed to plant, but never own the blossom.
- deleted 1y ago[deleted]
- blibble 1y agoAI companies continuing to have problems with the concept of "consent" is increasingly alarming god help us if they ever manage to build anything more than shitty chatbots
- goatlover 1y agoThey're certainly pouring billions of dollars into trying to build something more. Or at least that's what they're telling the public and investors.
- tempfile 1y agoDo you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?
- mplewis 1y agoIf I were DOSing your blog, you'd ask me to stop. I run server ops for multiple online communities that are being severely negatively impacted and DOSed by these AI scrapers, and we have very few ways to stop them.
- tempfile 1y agoThat is a problem, but is not related to my comment. The person I'm replying to is acting as if consent is a relevant aspect of the public web, I am saying it isn't. That is not the same as saying "you can do whatever you want to a public server". It is just that what you are allowed to do is not related to the arbitrary whim of the server operator.
- Wilder7977 1y agoConsent is also expressed through technical conventions. I, the website owner, express my intention through - for example - robots.txt. if you write a bot that specifically ignores it, you are violating consent. Likewise, I may prevent certain user-agents to visit my site. If you - say, an AI megacorp - are intentionally spoofing the user-agent to appear as a user, you are also violating consent.
- deleted 1y ago[deleted]
- jp1016 1y agoUsing a robots.txt file to block crawlers is just a request, it’s not enforced. Even if some follow it, others can ignore it or get around it using fake user agents or proxies. It’s a battle you can’t really win.
- gonzo41 1y agoThis is expected. There are not rules or conventions anymore. Look at LLMs, they stole/pirated all knowledge....no consequences.
- Havoc 1y agoSeems a win. CF being internet police is a problem too but someone credible publicly shaming a company for shady scraping is good. Even if it just creates conversation Somehow this needs to go back to search era where all players at least attempt to behave. This scrapping Ddos stuff and I don’t care if it kills your site (while “borrowing” content) is unethical bullshit
- deleted 1y ago[deleted]
- jeffrallen 1y agoShaming doea not work in the era of "no shame".
- Havoc 1y agoAny better workable ideas that do work?
- tucnak 1y agoThe rage-baiters in this thread are merely fishing for excuses to go up against "the Machine," but honestly, widely off-mark when it comes to reality of crawling. This topic has been chewed to bits long before LLM's, but only now it's a big deal because somebody is able to make money by selling automation of all things..? The irony would be strong to hear this from programmers, if only it didn't spell Resentment all over. If you don't want to get scrapped, don't put up your stuff online.
- rzz3 1y ago[flagged]
- fourside 1y ago> companies who want AI to recommend their products need to turn this off before it starts hurting them financially Content marketing, gamified SEO, and obtrusive ads significantly hurt the quality of Google search. For all its flaws, LLMs don’t feel this gamified yet. It’s disappointing that this is probably where we’re headed. But I hope OpenAI and Anthropic realize that this drop in search result quality might be partly why Google’s losing traffic.
- ipaddr 1y agoThis has already started with people using special tags also people making content just for llms.
- jedberg 1y agoThere is a standard for making content just for LLMs: https://llmstxt.org https://llmstxt.org
- yoz-y 1y agoFrom their example I don’t see any value in this on top of making and actually human friendly site. > Converting complex HTML pages with navigation, ads, and JavaScript into LLM-friendly plain text is both difficult and imprecise. None of these conditions should apply for websites with purpose of providing information.
- rzz3 1y agoI hope they realize Cloudflare opted them in to blocking LLMs.
- gcbirzan 1y agoI hope you realise that lying is bad.
- observationist 1y agoCrawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC mechanisms and differential content loading with offline caching and storage, but this puts control of content in the hands of the user, mitigates the value of surveillance and tracking, and has other side effects unpalatable to those currently exploiting user data. Adtech companies want their public reach cake and their mass surveillance meals, too, with all sorts of malignant parties and incentives behind perpetuating the worst of all possible worlds.
- tantalor 1y agoI think Cloudfare is setting themselves up to get sued. (IANAL) tortious interference
- emehex 1y agoWould highly recommend listening to the latest Hard Fork podcast with Matthew Prince (CEO, Cloudflare): https://www.nytimes.com/2025/08/01/podcasts/hardfork-age-restrictions-cloudflare.html https://www.nytimes.com/2025/08/01/podcasts/hardfork-age-res... I was skeptical about their gatekeeping efforts at first, but came away with a better appreciation for the problem and their first pass at a solution.
- glenstein 1y agoI don't think criticizing the business practices of Cloudfare does the work of excusing Perplexity's disregard for norms.
- rustc 1y ago> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. > If you want to gatekeep your content, use authentication. Are there no limits on what you use the content for? I can start my own search engine that just scrapes Google results?
- curiousgal 1y agoI am sorry, Cloudafre is the internet police now?
- otterley 1y agoWhich is ironic given they are the primary enabler of streaming video copyright infringement on the Internet.
- rzz3 1y agoThey hate AI it seems. I don’t see them offering any AI products or embracing it in any way. Seems like they’ll get left behind in the AI race.
- Oras 1y agoIf they managed to enforce the pay-per-scrape, that would be a huge revenue, bigger than AdSense
- rzz3 1y agoHuge revenue sure, but at what cost? Models should be able to be trained on anything a human can read and see without paying.
- bobnamob 1y ago? https://developers.cloudflare.com/workers-ai/ https://developers.cloudflare.com/workers-ai/ ? https://ai.cloudflare.com/ https://ai.cloudflare.com/
- rzz3 1y agoAh TIL. These are tiny models though but maybe it’s a good sign.
- otterley 1y agoI don't think they hate AI. I think they're offering a service that their customers want.
- talkingtab 1y agoI wonder if DRM is useful for this. The problem: I want people to access my site, but not Google, not bots, not crawlers and certainly not for use by AI. I don't really know anything about DRM except it is used to take down sites that violate it. Perhaps it is possible for cloudflare (or anyone else) to file a take down notice with Perplexity. That might at least confuse them. Corporations use this to protect their content. I should be able to protect mine as well. What's good for the goose.
- 1gn15 1y agohttps://en.wikipedia.org/wiki/Analog_hole https://en.wikipedia.org/wiki/Analog_hole
- bob1029 1y ago"Stealth" crawlers are always going to win the game. There are ways to build scrapers using browser automation tools [0,1] that makes detection virtually impossible. You can still captcha, but the person building the automation tools can add human-in-the-loop workflows to process these during normal business hours (i.e., when a call center is staffed). I've seen some raster-level scraping techniques used in game dev testing 15 years ago that would really bother some of these internet police officers. [0] https://www.w3.org/TR/webdriver2/ https://www.w3.org/TR/webdriver2/ [1] https://chromedevtools.github.io/devtools-protocol/ https://chromedevtools.github.io/devtools-protocol/
- blibble 1y ago> "Stealth" crawlers are always going to win the game. no, because we'll end up with remote attestation needed to access any site of value
- gkbrk 1y agoAlmost no site of value will use remote attestation because an alternative that works will all of your devices, operating systems, ad blockers and extensions will attract more users than your locked-down site.
- blibble 1y agotell that to the massive content sites already using widevine
- bakugo 1y ago> alternative that works will all of your devices, operating systems, ad blockers and extensions When 99.9% of users are using the same few types of locked down devices, operating systems, and browsers that all support remote attestation, the 0.1% doesn't matter. This is already the case on mobile devices, it's only a matter of time until computers become just as locked down.
- Buttons840 1y agoYes, because there's always the option for a camera pointed at the screen and a robot arm moving the mouse. AI is hoping to solve much harder problems.
- kocial 1y agoThose Challenges can be bypassed too using various browser automation. With the Comet-like tool, Perplexity can advance its crawling activity with much more human-like behaviour.
- ipaddr 1y agoIf they can trick the ad networks then go for it. If the ad networks can detect it and exclude those visits we should be able to.
- rustc 1y agoIt's ironic Perplexity itself blocks crawlers: $ curl -sI https://www.perplexity.ai | head -1 HTTP/2 403 Edit: trying to fake a browser user agent with curl also doesn't work, they're using a more sophisticated method to detect crawlers.
- thambidurai 1y agosomeone asked this already to the CEO: https://x.com/AravSrinivas/status/1819610286036488625 https://x.com/AravSrinivas/status/1819610286036488625
- fireflash38 1y agoThe bots are coming from inside the house
- czk 1y agoironically... they use cloudflare.
- Trung0246 1y agoTry this then: https://github.com/lwthiker/curl-impersonate https://github.com/lwthiker/curl-impersonate
- tr_user 1y agouse anubis to throw up a POW challenge
- micromacrofoot 1y agoEvery major AI platform is doing this right now, it's effectively impossible to avoid having your content vacuumed up by LLMs if you operate on the public web. I've given up and restored to IP based rate-limiting to stay sane. I can't stop it, but I can (mostly) stop it from hurting my servers.
- caesil 1y agoCloudflare is an enemy of the open and freely accessible web.
- jgrall 1y agoIf by "open and freely accessible" you mean there should be no rules of the road, then I suppose yes. Personally, I'm glad CF is pushing back on this naive mentality.
- znpy 1y agoAt work I'm considering blocking all the ip prefixes announced by ASNs owned by Microsoft and other companies known for their LLMs. At this point it seems like the only viable solutions. LLM scrapers bots are starting to make up a lot of our egress traffic and that is starting to weight on our bills.
- chuckreynolds 1y agoinsert 'shocked' emoji face here
- bilater 1y agoAs others have mentioned the problem is that of scale. Perhaps there needs to be a rate limit (times they ping a site) set within robots.txt that a site bot can come but only X times per hour etc. At least we move from a binary scrape or no scrape to a spectrum then.
- willguest 1y ago> The Internet as we have known it for the past three decades is rapidly changing, but one thing remains constant: it is built on trust. I think we've been using different internets. The one I use doesn't seem to be built on trust at all. It seems to be constantly syphoning data from my machine to feed the data vampires who are, apparently, additing to (I assume, blood-soaked) cookies
- jgrall 1y agoAin't that the truth.
- seydor 1y ago> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior That said, why does perplexity even need to crawl websites? I thought they used 3rd party LLMs. And those LLMs didn't ask anyones permission to crawl the entire 'net. Also the "perplexity bots" arent crawling websites, they fetch URLs that the users explicitly asked. This shouldnt count as something that needs robots.txt access. It's not a robot randomly crawling, it's the user asking for a specific page and basically a shortcut for copy/pasting the content
- pphysch 1y agoSpam and DDOS are serious problems, it's not fair to suggest Cloudflare is just doing this to gatekeep the Internet for its own sake.
- seydor 1y agoIt's definitely not a DDOS when it's a single http request per year. I don't know if they do it on purpose but the fact is none of the big tech crawlers are limited.
- zaphar 1y agoThis is most attributable to the fact that traffic is essentially anonymous so the source ip address is the best that a service can do if it's trying to protect an endpoint.
- ok123456 1y agoovh does a good job with ddos
- 1y ago
- daft_pink 1y agoI’m just curious at what point ai is a crawler and at what point ai is a client when the user is directing the searches and the ai is executing them. Perplexity Comet sort of blurs the lines there as does typing quesitons into Claude.
- mikewarot 1y agoSo, this calls for a new type of honeytrap, content that appears to be human generated, and high quality, but subtly wrong, preferably on a commercially catastrophic way. Behind settings that prohibit commercial usage. It really shouldn't be hard to generate gigantic quantities of the stuff. Simulate old forum posts, or academic papers.
- djoldman 1y agoThe cat's out of the bag / pandora's box is opened with respect to AI training data. No amount of robots.txt or walled-gardening is going to be sufficient to impede generative AI improvement: common crawl and other data dumps are sufficiently large, not to mention easier to acquire and process, that the backlash against AI companies crawling folks' web pages is meaningless. Cloudflare and other companies are leveraging outrage to acquire more users, which is fine... users want to feel like AI companies aren't going to get their data. The faster that AI companies are excluded from categories of data, the faster they will shift to categories from which they're not excluded.
- poorweb 1y ago[dead]
- tempfile 1y ago[flagged]
- ibero 1y agowhat if i want the rate set to zero?
- tempfile 1y agoThen turn off the server? You don't have a right to say who or what can read your public website (this is a normative statement). You do have a right not to be DoS'd. If you pretend not to know what that means, it sounds the same as saying "you have an arbitrary right to decide who gets to make requests to your service", but it does not mean that.
- jgrall 1y ago"You do not have a right to decide what people do with the content." Smh. Yes, laws be damned.
- pera 1y agoLike many other generative AI companies, Perplexity exploits the good faith of the old Internet by extracting the content created almost entirely by normal folks (i.e. those who depend on a wage for subsistence) and reproducing it for a profit while removing the creators from the loop - even when normal folks are explicitly asking them to not do this. If you don't understand why this is at least slightly controversial I imagine you are not a normal folk.
- rapatel0 1y agoThis is brilliant marketing and strategy from Cloudflare. They are pointing out bad actors and selling a service where they can be the private security guards for your website. I think there could be something interesting if they made a caching pub-sub model for data scraping. In addition or in place of trying to be security guards.
- czk 1y agothe year is 2045. you've been cruising the interstate in your robotaxi, shelling out $150 in stablecoins at the cloudflare tollbooth. a palantir patrol unit pulls you over. the optimus v4 approaches your window and contorts its silicone face into a facsimile of concern as it hits you with the: "sir, have you been botting today?" immediately you remember how great you had it in the '20s when you used to click CAPTCHA grids to prove your humanity to dumb algorithms, but now the machines demand you recite poetry or weep on command "how much have you had to bot today?", its voice taking on an empathetic tone that was personalized for your particular profile "yeah... im gonna need you to exit the vehicle and take a field humanity test"
- benreesman 1y agogreentext courtesy of opus prompted in greentext by me. ``` The user wants me to create a more elaborate greentext with deeper worldbuilding and references to internet culture/incel memes, but specifically asks for "tasteful" ones and explicitly says no violent references. I should incorporate classic internet/chan culture references like tendies, NEET lifestyle, "tfw no gf", basement dwelling, etc. while building out this dystopian 2045 world where proving humanity is the central challenge. \>be me, 2045 \>32yo NEET still living in mom's smart-basement \>she threatens to turn off my Neuralink if I don't go outside today \>grudgingly take the robotaxi to pick up tendies \>$150 stablecoin toll at cloudflare checkpoint \>have to solve a moral dilemma to prove I'm not running on GPT-7 \>sweating.dll \>palantir unit pulls me over \>optimus v4 leans into window \>its facial mesh attempts "concern_expression_v2.blend" \>"sir, when did you last feel genuine human connection?" \>flashback to 2024 when the girl at McDonalds gave me extra honey mustard \>that was before the McBots took over \>"t-twenty one years ago officer" \>optimus's empathy subroutines activate \>"sir I need you to perform a field humanity test" \>get out, knees weak from vitamin D deficiency \>"please describe your ideal romantic partner without using the words 'tradwife' or 'submissive'" \>brain.exe has stopped responding \>try to remember pre-blackpill emotions \>"someone who... likes anime?" \>optimus scans my biometrics \>"stress patterns indicate authentic social anxiety, carry on citizen" \>get back in robotaxi \>it starts therapy session \>"I notice you ordered tendies again. Let's explore your relationship with your mother" \>tfw the car has better emotional intelligence than me \>finally get tendies from Wendy's AutoServ \>receipt prints with mandatory "rate your humanity score today" \>3.2/10 \>at least I'm improving \>mfw bots are better at being human than humans \>it's over for carboncels ```
- deleted 1y ago[deleted]
- decide1000 1y agoC'mon CF. What are you doing? You are literally breaking the internet with your police behaviour. Starts to look like the Great Firewall.
- jgrall 1y agoNot affiliated with CF in any way. Respectfully disagree. Calling out bad actors is in the public interest.
- imcritic 1y agoCF is a bad actor. They ruin internet. They own more and more parts of it.
- decide1000 1y agoIt's not in my interest that a tech company from the US decides what a bad actor is.
- kylestanfield 1y agoPerplexity claims that you can “use the following robots.txt tags to manage how their sites and content interact with Perplexity.” https://docs.perplexity.ai/guides/bots https://docs.perplexity.ai/guides/bots Their fetcher (not crawler) has user agent Perplexity-User. Since the fetching is user-requested, it ignores robots.txt . In the article, it discusses how blocking the “Perplexity-User” user agent doesn’t actually work, and how perplexity uses an anonymous user agent to avoid being blocked.
- nostrademons 1y agoIt's entirely possible that it's not Perplexity using the stealth undeclared crawlers, but rather their fallback is to contract out to a dedicated for-pay webscraping firm that retrieves the desired content through unspecified means. (Some of these are pretty dodgy - several scraping companies effectively just install malware on consumer machines and then use their botnet to grab data for their customers.). There was a story on HN not long ago about the FBI using similar means to perform surveillance that would be illegal if the FBI did it itself, but becomes legal once they split the different parts up across a supply chain: https://news.ycombinator.com/item?id=44220860 https://news.ycombinator.com/item?id=44220860
- echo42null 1y agoHmm, I’ve always seen robots.txt more as a polite request than an actual rule. Sure, Google has to follow it because they’re a big company and need to respect certain laws or internal policies. But for everyone else, it’s basically just a “please don’t” sign, not a legal requirement or?
- crossroadsguy 1y agoI was recently listening to Cloudflare CEO on the Hard Fork podcast. He seemed to be selling a way for content creators to stop AI companies from profiting off such leeching. But the way he laid the whole thing out, adding how they are best placed to do this because they are gatekeepers of X% of the Internet (I don't recall the exact percentage), had me more concerned than I was at the prospect of AI companies being the front of summarised or interpreted consumption. He went on, upfront — I’d give him that, to explain how he is expecting a certain percentage of that income that will come from enforcing this on those AI companies and when the AI companies pay up to crawl. Cloudflare already questions my humanity and then every once in a while blocks me with zero recourse. Now they are literally proposing more control and gatekeeping. Where have we all come on the Internet? Are we openly going back to the wild west of bounty hunters and Pinkertons (in a way)?
- skeledrew 1y agoThis is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the LLM output, and there's an actual need to go to said source.
- remus 1y ago> The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If you are the source I think they could make plenty of sense. As an example, I run a website where I've spent a lot of time documenting the history of a somewhat niche activity. Much of this information isn't available online anywhere else. As it happens I'm happy to let bots crawl the site, but I think it's a reasonable stance to not want other companies to profit from my hard work. Even more so when it actually costs me money to serve requests to the company!
- crazygringo 1y ago> but I think it's a reasonable stance to not want other companies to profit from my hard work Imagine someone at another company reads your site, and it informs a strategic decision they make at the company to make money around the niche activity you're talking about. And they make lots of money they wouldn't have otherwise. That's totally legal and totally ethical as well. The reality is, if you do hard work and make the results public, well you've made them public. People and corporations are free to profit off the facts you've made public, and they should be. There are certain limited copyright protections (they can't sell large swathes of your words verbatim), but that's all. So the idea that you don't want companies to profit from your hard work is unreasonable, if you make it public. If you don't want that to happen, don't make anything public.
- remus 1y agoFor me, the point is that the person who has put in the work then has some rights to decide how that information is accessed and re-used. I think it is a reaosnable position for someone to hold that they want individuals to be able to freely use some content they produced, but not for a company to use and profit from that same content. I think just saying "It's public now" lacks any nuance. Ultimately these AI tools are useful because they have access to huge swaths of content, and the owners of these tools turn a lot of revenue by selling access to these tools. Ultimately I think the internet will end up a much worse place if companies don't respect clearly established wishes of people creating the content, because if companies stop respecting things like robots.txt then people will just hide stuff behind logins, paywalls and frustraing tools like cloudflare which use heuristics to block malicious traffic.
- dhanushreddy29 1y agoPS: perplexity is using cloudflare browser rendering to scrape websites
- kazinator 1y agoWhy single out Perplexity? Pretty much no crawler out there fetches robots.txt. robots.txt is not a blocking mechanism; it's a hint to indicate which parts of a site might be of interest to indexing. People started using robots.txt to lie and declare things like no part of their site is interesting, and so of course that gets ignored.
- gcbirzan 1y agoThat's not true, at all.
- _Algernon_ 1y agoThis is objectively wrong. Take it straight from the source: https://www.rfc-editor.org/rfc/rfc9309.html https://www.rfc-editor.org/rfc/rfc9309.html
- pywren 1y ago[dead]
- kotaKat 1y agoAn AI service violating peoples’ consent? Say it isn’t so! Those damn assult-culture techbros at it again.
- bob1029 1y agoHas anyone bothered to properly quantify the worst case load (i.e., requests per second) that has been incurred by these scraping tools? I recall a post on HN a few weeks/months ago about something similar, but it seemed very light on figures. It seems to me that ~50% of the discourse occurring around AI providers involves the idea that a machine reading webpages on a regular schedule is tantamount to a DDOS attack. The other half seems to be regarding IP and capitalism concerns - which seem like far more viable arguments. If someone requesting your site map once per day is crippling operations, the simplest solution is to make the service not run like shit. There is a point where your web server becomes so fast you stop caring about locking everyone into a draconian content prison. If you can serve an average page in 200uS and your competition takes 200ms to do it, you have roughly 1000x the capacity to mitigate an aggressive scraper (or actual DDOS attack) in terms of CPU time.
- ch_fr 1y agoI mean, it did happen, don't you remember in March when SourceHut had outages because their most expensive endpoints were being spammed by scrapers? Don't you remember the reason Anubis even came to be? It really wasn't that long ago, so I find all of the snarky comments going "erm, actually, I've yet to see any good actors get harmed by scraping ever, we're just reclaiming power from today's modern ad-ridden hellscape" pretty dishonest.
- xmodem 1y agoQuestion for those in this thread who are okay with this: If I have endpoints that are computationally expensive server-side, what mechanism do you propose I could use to avoid being overwhelmed? The web will be a much worse place if such services are all forced behind captchas or logins.
- m3047 1y agoIn 2005 I used a bot motel with Markov Chain derived dummy content for this exact purpose.
- alexey-salmin 1y agoHow do you make the money you need to finance these computationally expensive server-side endpoints?
- xmodem 1y agoMaybe I'm a community-driven project funded by donations and volunteer time. Maybe I'm a local government with extremely limited IT budget and no in-house skills. Maybe I'm just some dude who maintains a hobby project that lives on a NUC under my desk.
- codecracker3001 1y ago> we were able to fingerprint this crawler using a combination of machine learning and network signals. what machine learning algorithms are they using? time to deploy them onto our websites
- madrox 1y agoEvery time there's an industry disruption there's good money to be made in providing services to incumbents that slow the transition down. You saw it in streaming, and even the internet at large. Cloudflare just happens to be the business filling that role this time. I don't really mind because history shows this is a temporary thing, but I hope web site maintainers have a plan B to hoping Cloudflare will protect them from AI forever. Whoever has an onramp for people who run websites today to make money from AI will make a lot of money.
- nialse 1y agoIs it just me or is it rage bait? Switching up marketing a notch when the AI paywall did not get much media attention so far? Cloudflare seems to focus on enterprise marketing nowadays, currently geared towards the media industry, rather than the technical marketing suited for the HN audience. They have no horse in the AI race, so they’re betting on the anti-AI horse instead to gain market share in the media sector?
- zeld4 1y agoInternet was built on trust, but not anymore. It's a Darwinian system; everyone has to find their own way to survive. Cloudflare will help their publisher to block more aggresively, and AI companies will up their game too. Harvest information online is hard labor that needs to be paid for, either to AI, or to human.
- hnburnsy 1y agoRespone from Perpelexity to Tech Crunch... >Perplexity spokesperson Jesse Dwyer dismissed Cloudflare’s blog post as a “sales pitch,” adding in an email to TechCrunch that the screenshots in the post “show that no content was accessed.” In a follow-up email, Dwyer claimed the bot named in the Cloudflare blog “isn’t even ours.”
- blablabla123 1y agoYeah this is funny. The CDNs basically started more than a decade ago pushing vast amounts of data through the networks. Their complaint even if valid feels like hypocrisy. Either way, the CDNs profit big time from the AI scraping hype and the current copyright anarchy in the US
- throwmeaway222 1y agoChange "no-crawl" to "will-sue" and see if that fixes the problem.
- fsckboy 1y agothe internet needs micropayments (probably millipayments). if crawlers want to pay me a penny a page, crawl me 24-7 plz if I am willing to pay a penny a page, i and the people like me won't have to put up with clickwrap nonsense free access doesn't have to be shut off (ok, it will be, but it doesn't have to be, and doesn't that tell you something?) reddit could charge stiffer fees, but refund quality content to encourage better content. i've fantacized about ideas like "you pay upfront a deposit; you get banned, you lose your deposit; withdraw, have your deposit back", the goal being simplify the moderation task while encouraging quality. because where the internet is headed is just more and more trash. here's another idea, pay a penny per search at google/search engine of choice. if you don't like the results, you can take the penny back. google's ai can figure out how to please you. if the pennies don't keep coming in, they serve you ad-infested results; serve up ad-infested results, you can send your penny to a different search engine.
- purplehat_ 1y ago[dead]
- zzo38computer 1y agoI do not want to block curl and lynx. But if they claim to be Chrome then I don't care if Chrome is blocked
- qwerty456127 1y agoIt's time to stop blocking crawlers and using captchas and start building web sites that are intentionally AI-friendly by design. Even before the modern LLMs, anti-scraper measures apparently were primarily befitting Google whose scrapers were the most common exception.
- hrpnk 1y agoPreviously it was all sniper and sneaker bots scanning websites for product availability and attempting purchases continuously to snipe when it comes back online. Now, it's a gazillion of AI crawlers and python crawlers, MCP servers that offer the same feature to anyone "building (personal workflow) automation" incl. bypass of various, standard protection mechanisms.
- ed_mercer 1y ago>OpenAI is an example of a leading AI company that follows these best practices. Except when their agents happily click the "I"m not a robot" checkbox.
- ergocoder 1y agoCloudflare shading Perplexity is an unexpected drama of this year. I had to check that this did come out of CloudFlare.
- mathiaspoint 1y agoAll user agents are robots, some just have an associated person. Ban UAs that abuse the network but beyond that there's really nothing you can practically do if you actually want a website.
- 627467 1y agoI kind love this fast escalation. Clearly the web can benefit from people to start thinking for locally or narrowly instead of "global audiences". By locally I don't necessarily mean geographically local, just socially local. Build your audience then invite them into private(r) spaces. The (old) open web will be filled with machines built for machines. We learned to dislike "bubbles" in the past decades but bubbles make sense and are natural, obviously if you're not alone in it. When it becomes awfully busy with machines and machine content humans will learn to reconnect.
- 5pl1n73r 1y agoI think robots.txt should be ignored. Everyone wants people to not do things they don't like. We don't have to entertain each and every such one. The future is IPFS or something like it, so "crawling" will be a meaningless act.
- elphinstone 1y agoThey don't have the monopoly advantage of Google who has already stolen everything, so hard to feel outraged here. In fact it shows insidious Google's monopolistic stranglehold truly is.
- UltraSane 1y agoAny information you make available on the internet WILL be accessed by ANYONE and you CANNOT STOP THIS.
- deleted 1y ago[deleted]
- wordofx 1y agoGood on perplexity.
- amai 1y agoWhat if their “crawler” is just cheap human labor in some country with very low wages? Would that be allowed, because these are not robots?
- coffeeenjoyer 1y agoWhat if they have significant robotic body parts? Or what if they make heavy use of automation processes and they barely click a button to index a page (so they just maniacally click all day long)? What if robots.txt should refer to the ultimate beneficiaries... one which in this case would be the AI product that uses that content... to serve another ultimate beneficiary, a human user. The problem here is obviously the higher prices for hosting the content, and less revenue for those that serve ads, have product placement on their sites, etc. As long as robots.txt is about ethics/money and is enforced by morality, it doesn't matter who it refers to anyway. Public-shaming enforcement might work in some cases though, but I doubt it will be that useful. We're talking about companies that have trained their AIs on IPs, and tried their best to later hide it. Does shame affect robots, or companies for that matter? Cloudflare would very much like to be the middleman for monetary transactions between AI services and site owners (https://blog.cloudflare.com/introducing-pay-per-crawl/ https://blog.cloudflare.com/introducing-pay-per-crawl/), but at the moment they don't have a law to hold their back, so articles like these are the best they got.
- minelbu 1y ago[dead]
- mrbald 1y agoWe (humanity) need to invent a simple GPLv3 style license “You can derive any data on the data you see here, any derived data you sell or share should mention this place as a source and is subject to the same copyright as the source”. This will imply scraped datasets should become public and the law enforcement bodies will be able to work in an established framework to fight copyright and license crimes. Just blocking me from using any tools I want to make sense of the world around me (data on the internet sites being part of it) with crawlers and whatnot, is inherently evil, and is not logically consistent.
- hsbauauvhabzb 1y agoBecause scrapers would certainly comply with that /s
- mrbald 1y agoMore like have easier to assess legality status.
- account42 1y agoHow so? The legal status without a license is already "All Rights Reserved".
- hsbauauvhabzb 1y agoA cost sink that has no upside? How on earth would scrapers stop themselves saying yes?
- harvie 1y agoMaybe we can just configure webservers to block anyone who requests robots.txt, regular browsers don't do it, but robots do to get list of urls to crawl (while ignoring rules). Just create simple PHP/CGI script that adds client IP addres to iptables once /robots.txt is accessed.
- Trung0246 1y agoOne way to easily bypass is to let external services fetching robots.txt (archive.org, GitHub actions, etc...) to cache it and either expose through separate apis/webhook/manual download to the actual scrape server. robots txt file size is usually small and would not alert external services.
- emsign 1y agoAI companies are just thieves with big money lawyers. What do you expect from so much criminal energy? They will never stop, they are crazy.
- ankmb 1y agoWill AI companies come up with a model to incentivise content creation. Is it necessary for their long term survival? And is it not imperative to happen?
- buremba 1y agoFunny enough, Perplexity blocks the bots themselves. Imagine I develop an "agent" called Merplexity, which simulates an anonymous client browsing on Perplexity and injects my ads into the output without paying for the Sonar API. Would that be OK with Perplexity?
- account42 1y agoOf course their proposed solution is to hand over the keys to Buttflare so that the problem goes away. No thanks, you don't counter shit with more but slightly different shit.
- tonyhart7 1y agopeople want LLM to access website but wait until those LLM given access to make a comment, write a reviews, moderation etc now suddenly everything on the net is fake if not already are
- oriettaxx 1y agoI've jyst asked perplexity ai itself: this is the answer In summary: Officially, Perplexity claims its bots honor robots.txt. In practice, outside investigators and hosting providers document persistent circumvention of such directives by undeclared or disguised crawlers acting on Perplexity’s behalf, especially for real-time user queries
- pywren 1y ago[dead]
- gtvwill 1y agoAdhering to robots.txt is merely a courtesy. Much like a trolley drop off at your local shopping center car park. Some users will adhere to it and drop their trolleys in after their done. Others will not and will leave it wherever. Your machine might access a page via a browser that is human readable. My machine might read it via software and present the content to me in some other form of my choosing. Neither is wrong. Just different. Don't like it? Then don't post your website on the internet...
- yesIreadIt 1y agoso cloudflare blocked the agent from accessing the site. then when it couldn't access the robots.txt because it was blocked they punished it for using intelligent work around to access a website with no known history. perplexity is running a browser that follows the instruction of the user. if the user could manually do it then the agent is simply a tool to do the manual thing. this is a battle about websites and advertisers pissed that their analytics show and impressions... let's not pretend cloudflare is protecting anyone
- lonelyasacloud 1y agoIn many ways what is going on with Perplexity is reminiscent of the earlier 2000s battles between the p2p music sharing services like Napster and the music industry. Then we had wildly popular services (p2p) where most of the content was being provided illegally without payment to the IP owners. Which makes it particularly interesting now that Apple is being linked with Perplexity. Because in large part p2p music services were effectively consigned to history by Apple (primarily) negotiating with the music industry so that it could provide easy, seamless purchase and playback of legal music for their shiny new (at the time) mass-market Apple iPod devices: it then turning out that most users are happy to pay for content if it is not too expensive and is very convenient. Given Apple’s existing relationships with publishers through its music, movies, books, and news services, it’s not hard to imagine them attempting a similar play now.
- S4H 1y agoI believe there should be a <fetcher.txt> file, similar to <robots.txt>, which allows website owners to specify whether they want their site to be fetched and included in the responses of platforms like Perplexity.
- KETpXDDzR 1y ago> We were able to fingerprint this crawler using a combination of machine learning and network signals. Yikes. AntiVirus scanners for website access.
- ddalessa 1y agoCloudflare's test was to setup a dummy domain that had never been indexed, and had blocks in the robots.txt and the firewall. Then when they asked perplexity it came up with details about the 'exact' content (according to Cloudflare) but their attached screenshot shows the opposite, it shows some generic guesses about the domain ownership and some dynamic ads based on the domain name. If Perplexity was stealthily visiting the dummy site they would have seen it, as the site was not indexed and no one else was visiting the site. Instead it appears they made assertions about general traffic, not their dummy site. Its not very convincing.
- lofaszvanitt 1y agoCloudflare now acting like a self made police station of the internet.
- hamishmacewan 1y agoCloudflare sits in a privileged choke-point of the internet, peering into traffic others can’t. Now they’re playing hall monitor, publicly wagging their finger at Perplexity for “stealth crawling.” If Perplexity’s a customer, they should be furious at the breach of trust; if not, this smells like a cheap sales pitch dressed up as public service. Who appointed Cloudflare as Miss Manners of the web, or deputised them as law enforcement?
- icetank 1y agoCloudflare does allow bots to scrape sites. But in this case cloudflare was alerted by customers that there setting to disallow ai companies to access there site was not working. People pay cloudflare to specifically block perplexity bots.