4 ms·
It would have to be aggregated data that's constantly updated if it were to be anywhere competitive to just simply scraping the page.
by MattDaEskimo 2y ago
It would have to be aggregated data that's constantly updated if it were to be anywhere competitive to just simply scraping the page.
- UncleEntity 2y agoPlus, the "biggest scrapers" have shown their unwillingness to pay for content.
- aftbit 2y agoYes, presumably the approach would be to publish both a current dump of all of the databases every week or month, as well as a daily diff. Even better if it were just given away for free. The challenge would be around monetization and licensing. Even if you only published e.g. the kernel mailing list archive in this way, I could see someone like OpenAI turning around and suing to claim that you had given them copyright over that content rather than just providing the content itself.