7 ms·
Interesting. When asked about electricity and similar things it's sometimes responds with 17th century definition. And sometimes it spits out modern wikipedia-l
by severak_cz 3y ago
Interesting. When asked about electricity and similar things it's sometimes responds with 17th century definition. And sometimes it spits out modern wikipedia-like definition.
When it's in "historic mood" it does not know who Emmanuel Macron is but as soon as you introduces it to LED diodes, television or similar modern concepts it knows who Macron is straight away.
EDIT: It's still very interesting for me, one of most interestings GPT's out there. Also it does not hide how to make sulphuric acid from me. :-D
- thomasahle 3y agoI was hoping it was trained ground up on old texts only. Now we don't really know, when it says something archaic, if it's because it's "pretending" to be old school, or because that is what it truly believes.
- eternauta3k 3y agoProbably not enough old texts for it to learn how the language and the world work. Hence the fine tuning.
- aroo 3y agoSounds like something right up the domain of synthetic data.
- yyyk 3y agoThere's a sufficient number of old texts - far more than necessary - if we were willing to converse with it in Latin*! Perhaps we could make a translator LLM to converse with the 17th century LLM... * I'm not talking about Roman-era Latin, Latin was the scholarly/international language for Europe for a long long time. AFAIK, most European Latin literature is untranslated...
- jstanley 3y agoAre you sure? In Andrej Karpathy's intro to LLMs[1], he says pre-training uses about 10TB of text from the web. It's hard to believe that a comparable amount of text had been created, even across all languages, in the entire history of humanity up to the 17th century. [1] https://www.youtube.com/watch?v=zjkBMFhNj_g https://www.youtube.com/watch?v=zjkBMFhNj_g
- yyyk 3y agoIf data quality is good, one can reach results with much less data: https://arxiv.org/abs/2305.07759 https://arxiv.org/abs/2305.07759
- thomasahle 3y agoHow many stories/tokens did they actually train on? I can't find it in the paper.
- yyyk 3y agoMaybe this will help? https://huggingface.co/datasets/roneneldan/TinyStories https://huggingface.co/datasets/roneneldan/TinyStories Also note https://arxiv.org/abs/2309.05463 https://arxiv.org/abs/2309.05463 (which is larger - obviously size still does contribute to performance)
- Dorialexander 3y agoYes. I think we may have enough for "full finetuning" and erasing to a large extent the previous knowledge. But that's still very far off for pretraining. "RomeGPT" is next on my list of Monad successors and to give you a general idea, we have on the order of tens of millions of words in classical Latin (and biggest source will… Augustine). There was a BERT Latin project that was able to collect roughly 500 million words in all with mostly early modern and modern Latin. In comparison I'm currently part of a project to pretrain a French model and we need… 140 billion words.
- vorticalbox 3y agoIt's finetuned on [0] OpenHermes-2-Mistral-7B From the model description > OpenHermes was trained on 900,000 entries of primarily GPT-4 generated data, from open datasets across the AI landscape. https://huggingface.co/teknium/OpenHermes-2-Mistral-7B https://huggingface.co/teknium/OpenHermes-2-Mistral-7B
- Dorialexander 3y agoYes I needed that for the conversational/instructional capacities. I've made a lot of tests with base models and it would not listen to instruction very well…
- marci 3y agoIt's an LLM, "pretending" and "truly beleiving" are the same, or rather don't exist. It's like, when you use prompts like : "You're an helpful assistant", is it believing, pretending, or beleiving to be pretending to be an helpful assistant? It's as funny as disconcerting to see intent and will attributed to probabiltities. Feels sometimes like we're close to making a religion out of this. History of the human race, i guess (https://www.youtube.com/watch?v=xuCn8ux2gbs https://www.youtube.com/watch?v=xuCn8ux2gbs for the ref).
- cubefox 3y agoIf LLMs from an internal world model, there are definitely things they "believe", and things they can pretend to believe.
- marci 3y agoIf LLMs have models (tokens), there are definitely tokens inside said models, and a program can manipulate them.
- thomasahle 3y agoIt's more like when you prompt the model with "be an assistant from the time of Shakespeare", you don't know if it'll imitate actual Shakespearean data in its training, or the sea of modern human imitations or "fan fiction". There are various ways modern knowledge may have "snuck in" through data contamination. If we really want to know what a Shakespearean chatbot would have looked like, we need to cap the training data. Maybe in the future, using better mechanistic interpretability we can get around this, but not right now.
- mycall 3y agoAn ole English LLM would be a delight but is there enough source material to revive dead languages?
- Dorialexander 3y agoYes you're perfectly right. I've currently tried to maintain some kind of uneasy balance between good conversational capacities (so that it really is a "chatGPT") and cultural reset, which means it may revert from its 17th persona occasionally. Actually this issue is a good illustration that LLM really are latent space explorer. When you prompt a clearly contemporary concept, the default embedding position will shift back to contemporary associations. As a prompt engineering trick, I find it helps to use faux archaism (such as "pray tell" as an introductory phrase). This is basically a reinforcement anchor in the 17th century region of embedding space.
- Llamamoe 3y agoI wonder if you could get around this by penalising obviously modern words.
- Dorialexander 3y agoEither that or appending archaic expressions in the prompts (a bit like the prompt extension of Midjourney)