3 ms·
It's terabytes of content and other than we're the host not really related to each other. However, for most projects, it's possible to download a zip file of al
by davidfischer 23d ago
It's terabytes of content and other than we're the host not really related to each other. However, for most projects, it's possible to download a zip file of all the HTML docs for that project. We have a lower rate limit to pull these, but a scraper can pull thousands of docs at once. We only host a few hundred thousand projects so pulling a zip of the latest docs for all of them could be done in a day or two at a very reasonable rate.
It's also possible to request the docs already processed into markdown[1]. Lastly, basically all of the docs come from Git. A smart scraper could just clone a project's repo.
[1] https://docs.readthedocs.com/platform/stable/reference/markdown-for-agents.html https://docs.readthedocs.com/platform/stable/reference/markd...