4 ms·
LLM training is difficult to audit, and as such good-faith barriers such as robots.txt or API endpoint scraping are all easily ignored. Even ignoring the query
by maven29 3y ago
LLM training is difficult to audit, and as such good-faith barriers such as robots.txt or API endpoint scraping are all easily ignored. Even ignoring the query side, and just from an indexing perspecrive, you see a considerable advantage.
LLMs (at runtime) and LLM mills (during training) have nothing to gain from restricting themselves to the open web. This is in contrast to Google, who is in the business of selling links to information and not information in itself.