4 ms·
Where is it getting the content from? My understanding is that, as a statistical model, it doesn’t contain a database of complete works.
by d13 3y ago
Where is it getting the content from? My understanding is that, as a statistical model, it doesn’t contain a database of complete works.
- robinduckett 3y agoI thought Wikipedia was included in its training corpus, in which case the full Litany of fear is available on Wikipedia. https://en.m.wikipedia.org/wiki/Bene_Gesserit https://en.m.wikipedia.org/wiki/Bene_Gesserit
- Sharlin 3y agoAnd thousands of other websites…
- viraptor 3y agoIf it's repeated enough times in the training data, it will become statistically the most likely continuation of that response. The difference between a database and statistical response is... not very clear. Just like you could learn a poem in a language you don't know by repeating various parts of it enough times.
- Sharlin 3y agoUh, obviously it can't recite the entirety of Dune, for example. But just as obviously, the Litany Against Fear is famous enough to likely exist verbatim thousands of times in its training corpus.
- jncfhnb 3y agoI don’t think this is obvious at all. Ignore the algorithm and temperature and all that for a bit and just have it show you the word probabilities. I wouldn’t bat an eye if the entirety of dune was in there.
- Sharlin 3y agoMaybe obvious was too strong a word, but I very much doubt that it contains a lossless representation of any copyrighted novel-length work. Just due to the pigeonhole principle if not anything else. The model is a vastly compressed approximation of the full training corpus.
- voxelghost 3y agoEven though vastly compressed, in terms of lossy vs. fidelity, it seems to be able to recite pretty much any bible verse you ask it for.
- jncfhnb 3y agoThe intersection space of 10^15 parameters is a lot of pigeon holes though.
- Maken 3y agoIn a sense, it contains a compressed database of its entire training data. That's where its knowledge comes from.
- dacryn 3y agoat some points it starts to resemble more of a lossy compression algorithm. The 'next most likely word' aspect of it all seem to encode entire sequences of text. Especially if it is a bit of a niche domain. Then these texts can be 'retrieved' from the 'compressed' set of examples.
- quickthrower2 3y agoLLMs have crazy good emergent properties. Even the stupidest model you use for learning NN, you chuck a transformer on it and it makes a huge difference, and I have a hard time knowing really, really, really ... WHY? Yes keys/values/queries, words talking to each other and stuff. But why!