3 ms·
I think I'm slow. Can you explain this again, maybe with more words?
by monktastic1 2y ago
I think I'm slow. Can you explain this again, maybe with more words?
- kccqzy 2y agoLet's say we choose 1900 as the cutoff date. That means during training the model is only able to access material written before 1900. It would have a good knowledge about everything discovered in the 19th century and before. There's a great deal of mathematics, physics and chemistry available then. What if now we engage a discussion with that LLM on something discovered after 1900? Say transmutation and nuclear weapons, or general relativity, or the ZFC set theory.
- VierScar 2y agoWouldn't it be easier to cutoff pre-2020-ish, and ask it to create the transformer architecture of gpt? 1900 is so long ago I doubt most documents are good quality if they've been digitised at all. Most likely just low quality scanned images of inconsistent, half-illegible typewriter documents. Transcribed with OCR at best.
- cellis 2y agoAlso so little training data from that era. Like, exponentially more data was created after, say, <year when most records become digitized = 1970>
- kccqzy 2y agoThe problem I see with any date after the popularity of the internet is that you just can't be sure of the right date. A lot of traditional web forums now have backdated forum posts that are clearly made by LLM with an implausible date: https://hallofdreams.org/posts/physicsforums/ https://hallofdreams.org/posts/physicsforums/
- throwup238 2y agoYou can use CommonCrawl - which has massive datasets going back to 2008 - and the Internet Archive.