4 ms·
I suspect this is a large reason why OpenAI trains GPT-4 using a dataset that is mostly clamped at September 2021. It's almost certainly a problem for LLM deve
by mrshadowgoose 3y ago
I suspect this is a large reason why OpenAI trains GPT-4 using a dataset that is mostly clamped at September 2021.
It's almost certainly a problem for LLM development, just like it is for humans. Humans generate all sorts of stuff, some of it being absolute bullshit, and it does seem to cause problems for other humans. Yet with effort it still seems to be possible to cut through the bullshit in a lot of cases, so it's likely not an insurmountable problem for LLM development either.
- philipkglass 3y agoIt makes me wonder if long-neglected archives of old newspapers, internal corporate documentation, and government publications could have newfound economic value as training material. I know that the absolute volume is small compared to e.g. large scale document scraping from the Web, but intuitively I would guess that one old Bureau of Standards report written with a high degree of literacy has more value than 1000 SEO-chasing "how to make pancakes" guides. A lot of old documents were never digitized simply because people didn't think they had value. Is it time for a reinvigorated Google Books with a wider mission?
- thomashop 3y agoNot if you're trying to churn out "How to make pancakes" guides with your LLM
- pmoriarty 3y ago"A lot of old documents were never digitized simply because people didn't think they had value. Is it time for a reinvigorated Google Books with a wider mission?" We have to be very careful with the "factual" content of old non-fiction works. So much of that, from history, to medicine, to biology, etc, turned out to be straight from their writers' imaginations. If an LLM considers such books to be no different from modern books on these subjects it would get a very skewed view of the world. Imagine asking an LLM for a medical diagnosis and it responding with something about the humors.
- philipkglass 3y agoThat's a good point. I wonder how LLMs currently deal with the passage of time, if they can at all. Do they know that Satya Nadella is the current CEO of Microsoft and that Steve Ballmer no longer is simply because there are more training documents reflecting the present state of things, or is there an explicit time component that helps to resolve conflicting facts from different years? Based on what I've read so far about LLMs (and not actually working on any of these models myself), I wouldn't have thought they model time or facts in a way that you could expect them to resolve conflicting claims from a 2010 document and a 2020 document. Yet they're unreasonably effective at question answering anyway.
- ris 3y agoSomething I've been thinking a lot about lately is the potential historical usefulness of LLMs which have been selectively trained on material from before a certain date, presumably resulting in a model exhibiting attitudes of the time in question. It would be extremely interesting to be able to ask for a 1950s take on a modern idea. For this, large amounts of very old text would be extremely valuable.
- tyingq 3y agoI suppose the relative low value for low volumes of text makes this not a problem, but... What about data that's sitting around and isn't supposed to be public? If training data gets scarce, does a market for small-medium sized data emerge? Like old homework papers, internal company documents, etc?