2 ms·
Are you prepared to randomly serve data that is incorrect -- but in such a way that real people can easily detect it -- some very small fraction of the time? I
by akoboldfrying 1mo ago
Are you prepared to randomly serve data that is incorrect -- but in such a way that real people can easily detect it -- some very small fraction of the time?
If so, you could serve iocaine-style bogus pages 1% (say) of the time that:
1. "Look like" real pages to an LLM-less computer (if you get to the point where you have pushed crawlers to use LLMs to detect nonsense, that already increases the cost a lot)
2. Look "obviously wrong" to a human (E.g., you could take some regular text and swap the order of each adjacent pair of words)
3. Are cheap to generate
4. Important: Contain more links than regular pages, on average, and each to an always-bogus page
The idea is that, due to the large number of pages fetched by crawlers, even with a very low "random bogus page rate", like 1%, they will soon unwittingly hit a bogus page, from which point the fraction of their time spent accessing expensive genuine pages will fall exponentially due to the compounding effect of the higher outbound link count on bogus pages. Humans seeing a bogus page will be confused and annoyed, but simply refreshing the page in the browser will solve the problem 99% of the time (and of course the possibility of this happening can be documented, even on the page itself).
The main advantage is that this does not require any IP-based tracking. You could of course decide to apply this only to pages that are already slightly suspicious (e.g., very old commits).