4 ms·
When was the last time you looked at robots.txt to find a page that wasn't linked anywhere else?
by voidUpdate 2y ago
When was the last time you looked at robots.txt to find a page that wasn't linked anywhere else?
- sharlos201068 2y agoCrawlers aren't interested in fake pages that aren't linked to anywhere, they're crawling the same pages your users are viewing.
- danielheath 2y agoAdding a disallowed url to your robots.txt is a quick way to get a ton of crawlers to hit it, without linking to it from anywhere. Try it sometime.
- brookst 2y agorobots.txt is not a sitemap. If it worked that way you could just make a 5TB file linking to a billion pages that look like static links but are dynamically generated.
- ccgreg 2y agorobots.txt has a maximum relevant size of 500 kib.
- gadflyinyoureye 2y agoTuesday. But I have odd hobbies.
- zzo38computer 2y agoIt was a while ago, and it was not deliberate (wget downloaded robots.txt as well as the files I requested, and I was able to find many other files due to that, some of which could not be accessed due to requiring a password, but some were interesting (although I did not use wget to copy those other files; I only wanted to copy the files I originally requested)).