3 ms·
> Can’t believe that: No one did it before me (to my knowledge) This has certainly been done before: http://en.wikipedia.org/wiki/Spider_trap http://en.wikiped
by solox3 13y ago
> Can’t believe that: No one did it before me (to my knowledge)
This has certainly been done before: http://en.wikipedia.org/wiki/Spider_trap http://en.wikipedia.org/wiki/Spider_trap
- Buetol 13y agoThanks, corrected it
- ozh 13y agoI had the same "fun" idea 6 years ago. In 3 months Yahoo Site Explorer reported 150 million pages indexed. As of today Google still shows 55 million pages (in comparison, https://www.google.com/search?q=site:http://en.wikipedia.org/ https://www.google.com/search?q=site:http://en.wikipedia.org... reports 34 million pages). I had to kill the experiment (no more new "pages" crawled) because of the CPU load and bandwidth costs, even throttling robots
- scoot 13y agoThe wikipedia article doesn't mention how long this has practice existed, but I know this goes back to at least the late 90s. Circa '98 IIRC, there was a module available for a webserver I worked with which generated a page with number of bogus email addresses per page, and a number of random urls per paged that when followed generated yet another page of bogus email addresses and links. You hid the link somewhere on a legitimate page, added the base path as an exclude in robots.txt, and any mail harvesting spam-bots would get sucked in. The idea may well have been around longer still.