5 ms·
I’m AI stupid. Does anyone know if training on multiple languages provides “cross-over” — so training done in German can be utilized when answering a prompt in
by jorgesborges 2y ago
I’m AI stupid. Does anyone know if training on multiple languages provides “cross-over” — so training done in German can be utilized when answering a prompt in English? I once went through various Wikipedia articles in a couple languages and the differences were interesting. For some reason I thought they’d be almost verbatim (forgetting that’s not how Wikipedia works!) and while I can’t remember exactly I felt they were sometimes starkly different in tone and content.
- bernaferrari 2y agono, it is basically an 'auto-correct' spell checker from the phone. It only knows what it was trained on. But it has been shown that a coding LLM that has never seen a programming language or a library can "learn" a new one faster than, say, a generic LLM.
- StevenWaterman 2y agoThat's not true, LLMs can answer questions in one language even if they were only trained on that data in another language. IE you train an LLM on both English and French in general, but only teach it a specific fact in French, it can give you that fact in English
- hdhshdhshdjd 2y agoYou, you can write a prompt in English, give it French, and get an accurate answer in English even with the original Mistral. Still blows my mind we came so far so fast.
- miki123211 2y agoGenerally yes, with caveats. There was some research showing that training a model on facts like "the mother of John Smith is Alice" but in German allowed it to answer questions like "who's the mother of John Smith", but not questions like "what's the name of Alice's child", regardless of language. Not sure if this holds at larger model sizes though, it's the sort of problem that's usually fixable by throwing more parameters at it. Language models definitely do generalize to some extend and they're not "stochastic parrots" as previously thought, but there are some weird ways in which we expect them to generalize but they don't.
- planb 2y ago> Language models definitely do generalize to some extend and they're not "stochastic parrots" as previously thought, but there are some weird ways in which we expect them to generalize but they don't. Do you have any good sources that explain this? I was always thinking LLMs are indeed stochastic parrots, but language (that is the unified corpus of all languages in the training data) already inherently contains the „generalization“. So the intelligence is encoded in the language humans speak.
- michaelt 2y agoI don't have explanations but I can point you to one of the papers: https://arxiv.org/pdf/2309.12288 https://arxiv.org/pdf/2309.12288 which calls it "the reversal curse" and does a bunch of experiments showing models that are successful at questions like "Who is Tom Cruise’s mother?" (Mary Lee Pfeiffer) will not be equally successful at answering "Who is Mary Lee Pfeiffer’s son?"
- spookie 2y agoIsn't that specific case just a matter of not having enough data _explicitly_ stating the reverse? Seems as if they are indeed stochastic parrots from that perspective.
- moffkalast 2y ago> language already inherently contains the „generalization“ The mental gymnastics required to handwave language model capabilities are getting funnier and funnier every day.
- dannyw 2y agoAnecdata, but I did some continued pretraining on a toy LLM using machine-translated data; of the original dataset. Performance improved across all benchmarks; in English (the original language).
- benmanns 2y agoAm I understanding correctly? You look an English dataset, trained an LLM, machine translated the English dataset to e.g. Spanish, continued training the model, and performance for queries in English improved? That’s really interesting.
- bionhoward 2y agoThere is evidence code training helps with reasoning so if you count code as another language then, this makes sense https://openreview.net/forum?id=KIPJKST4gw https://openreview.net/forum?id=KIPJKST4gw Is symbolic language a fuzzy sort of code? Absolutely, because it conveys logic and information. TLDR: yes!