3 ms·
> how should they deal with robots.txt files that are hundreds of megabytes large? What do huge robots.txt files like that contain? I tried a couple domains ju
by rococode 7y ago
> how should they deal with robots.txt files that are hundreds of megabytes large?
What do huge robots.txt files like that contain? I tried a couple domains just now and the longest one I could find was GitHub's - https://github.com/robots.txt https://github.com/robots.txt - which is only about 30 kilobytes.
- jedberg 7y agoThey enumerate every page on the site sometimes specifically for different crawlers. Or they have a ton of auto generated pages they don’t want crawled and call them out individually because they don’t realize robots.txt supports globing.
- greglindahl 7y agoCan you given an example in the wild?
- jedberg 7y agoI was actually trying to find an example when I made my initial comment, but was unable to. It's been a long time since I did web scraping. Since then there are a lot more frameworks that help you build a website (and a correspondingly sane robots.txt), so there may not be as many as before.