3 ms·
Yes, but the LLM is roughly equivalent to lossy compression of the corpus it is trained on so why wouldn't you preserve that actual corpus so that it can be use
by bloak 2y ago
Yes, but the LLM is roughly equivalent to lossy compression of the corpus it is trained on so why wouldn't you preserve that actual corpus so that it can be used to train some better LLM, or something better than an LLM, in the future?
(There may be a good answer to that question: perhaps, for example, the corpus can't be preserved for data protection reasons but the LLM trained on it can be preserved? For various reasons that doesn't seem very plausible, however.)
- crackalamoo 2y agoYou can do both. Preserving the corpus and building the LLM probably gives the best chance for future generations.