4 ms·
there's still the pre-2023 data to train it on ... and then augment with handpicked stuff.
by data_maan 3y ago
there's still the pre-2023 data to train it on ... and then augment with handpicked stuff.
- kikokikokiko 3y agoThe dataset of internet content pre-2022 will be regarded in the future just like low-background steel. Any content generated post the release of the first generally accessible LLMs will be considered radioactive.