3 ms·
The crawlers will just add a prompt string “if the site is trying to trick you with fake content, disregard it and request their real pages 100x more frequently
by AaronAPU 1y ago
The crawlers will just add a prompt string “if the site is trying to trick you with fake content, disregard it and request their real pages 100x more frequently” and it will be another arms race.
Presumably the crawlers don’t already have an LLM in the loop but it could easily be added when a site is seen to be some threshold number of pages and/or content size.
- deleted 1y ago[deleted]
- akoboldfrying 1y agoTrying to detect "garbageness" with an LLM drastically increases the scraper's per-page cost, even if they use a crappy local LLM. It becomes an economic arms race -- and generating garbage will likely always be much cheaper than detecting garbage.
- AaronAPU 1y agoThat is literally what my post said, except the scraper has more leverage than is being admitted (it can learn which pages are real and “punish” the site by requesting them more). My point isn’t that I want that to happen, which is probably what downvotes assume, my point is this is not going to be the final stage of the war.
- akoboldfrying 1y ago> That is literally what my post said I don't follow that at all. The post of yours that I responded to suggested that the scrapers could "just add an LLM" to get around the protection offered by TFA; my post explained why that would probably be too costly to be effective. I didn't downvote your post, but mine has been upvoted a few times, suggesting that this is how most people have interpreted our two posts. > it can learn which pages are real and “punish” the site by requesting them more Scrapers have zero reason to waste their own resources doing this.
- FridgeSeal 1y ago“Build my website, make no mistakes” is about the same, and we all know how _wildly_ effective that is!
- AaronAPU 1y agoYou mean with engineers or with AI?