7 ms·
I had a related, but orthogonal question about multilingual LLMs. When I ask smaller models a question in English, the model does well. When I ask the same mod
by ozgune 1y ago
I had a related, but orthogonal question about multilingual LLMs.
When I ask smaller models a question in English, the model does well. When I ask the same model a question in Turkish, the answer is mediocre. When I ask the model to translate my question into English, get the answer, and translate the answer back to Turkish, the model again does well.
For example, I tried the above with Llama 3.3 70B, and asked it to plan me a 3-day trip to Istanbul. When I asked Llama to do the translations between English <> Turkish, the answer was notably better.
Anyone else observed a similar behavior?
- petesergeant 1y agoFascinating phenomenon. It's like a new Sapir–Whorf hypothesis. Do language models act differently in different languages due to those languages or the training materials?
- shaky-carrousel 1y agoThey absolutely do. They know more in English than in Spanish, I've seen that on all models, since the beginning.
- namaria 1y agoThey have more data in English than Spanish. LLMs don't know or reason or follow instructions. They merely render text continuations that are coherent with the expectations you set when prompting. The fact that they are not able to sustain the illusion in languages with less available training data than English should make that clear.
- shaky-carrousel 1y ago> They have more data in English than Spanish. Yep, that there seems like the definition of knowing. Don't worry, your humanity isn't at risk.
- namaria 1y agoNo, mental models matter. This has nothing to do with AGI doomerism. Knowing implies reasoning. LLMs don't "know" things. These statistical models continuate text. Having a mental model that they "know" things, that they can "reason" or "follow instructions" is driving all sorts of poor decisions. Software has an abstraction fetish. So much of the material available for learners is riddled with analogies and "you don't need to know that" attitude. That is counter productive and I think having accurate mental models matters.
- petesergeant 1y ago> Knowing implies reasoning That's not really clear-cut, that's simply a position you're taking. JTB could (I reckon) say that a model's "knowledge" is justified by the training process and reward functions. > LLMs don't "know" things. These statistical models continuate text. I don't think it's clear to anyone at this point whether or not the steps taken before token selection (eg: the journey through their dimensional knowledge space provided by attention) are close to or far from how our own thought processes work, but the description of LLMs as "simply" continuating text reduces them to their outputs. From my perspective, as someone on the other side of a text-based web-app from you, you also are an entity that simply continuates text. You have no way of knowing whether this comment was written by a sentient entity -- with thoughts and agency -- or an LLM.
- shaky-carrousel 1y agoI have to disagree. We've been using "knowing" for programs for decades without requiring it to imply reasoning. Just because the output now looks more realistic doesn't mean we need to suddenly get philosophical about it. That shift says more about us than about the software. And while accurate mental models can help in certain contexts, they're not always necessary. I don't need a detailed model of how my OS handles file operations to use it effectively. A high-level understanding is usually enough. Insisting on deep internal accuracy in every case seems more like gatekeeping than good practice.
- evgen 1y agoThis is one of those subtle clues that the LLM does not actually 'know' anything. It is providing you the best consensus answer to your prompt using the data upon which the weights rest, is that data was input primarily as english then you are going to get better results asking in english. It is still Searle's Chinese Room except you need to first go to the 'Language X -> English' room and then deliver its output to the general query room before delivering the next result to the 'English -> Language X' room.
- jug 1y agoAnthropic’s research did find that Claude seemed to have an inner language agnostic ”language” though. And that the larger a LLM got, the more it could realize the innate meaning of words between language barriers as well as expand upon its internal non-specific language representation. So, part of its improved performance as they grow in parameter count is probably not only due to expanded raw material that it is trained upon, but a greater ability to ultimately ”realize” and connect apparent meanings of words, so that a German speaker might benefit more and more from training material in Korean. > These results show that features at the beginning and end of models are highly language-specific (consistent with the {de, re}-tokenization hypothesis [31] ), while features in the middle are more language-agnostic. Moreover, we observe that compared to the smaller model, Claude 3.5 Haiku exhibits a higher degree of generalization, and displays an especially notable generalization improvement for language pairs that do not share an alphabet (English-Chinese, French-Chinese). Source: https://transformer-circuits.pub/2025/attribution-graphs/biology.html https://transformer-circuits.pub/2025/attribution-graphs/bio... However, they do see that Claude 3.5 Haiku seemed to have an English ”default” with more direct connections. It’s possible that a LLM needs to go a more roundabout way via generalizations to communicate in alternative languages and where this causes a dropoff in performance the smaller the model is?
- numpad0 1y agoThe modern Standard Chinese language is almost syntactically "identical" to English, for some reason. French was direct ancestor to medieval British language that came to be the modern English. My point is, those language pairs aren't random examples. Chinese isn't something completely foreign and new thing when it comes to difference between it and English.
- input_sh 1y agoBoth, but primarily due to the lack of training materials. 10 or so million native speakers of my language will never be able to generate the same amount of training material as over a billion English speakers do. There is a steep drop in quality in any non-English language, but in general less native speakers = worse results. They tend to have a certain "voice" which is extremely easy to spot and the accuracy of results goes out the window (way worse than in English).
- petesergeant 1y agoRight, but it’s interesting that means its reasoning abilities potentially drop off when it’s talking Thai, or its knowledge of WW2 history in the Eastern Theatre might drop off when speaking French, where the same model has no trouble with the same questions in English. My French and Thai are both rudimentary, but I’m working from the same set of facts and reasoning ability in both languages. Will it give different answers on what the greatest empire that ever existed was if you ask it in Mandarin vs Italian vs Mongolian?
- mrweasel 1y agoSomeone apparently did observe ChatGPT (I think it was ChatGPT) switch to Chinese for some parts of it's reasoning/calculations and then back to English for the final answer. That's somehow even weirder than the LLM giving different answers depending on the input.
- ApolloFortyNine 1y agoI've seen this happen as well with o3-mini, but I'm honestly not sure what triggered it. I use it all the time but have only had it switch to Chinese during reasoning maybe twice.
- Telemakhos 1y agoI've seen Grok sprinkle random Chinese characters into responses I asked for in ancient Greek and Latin.
- andai 1y agoI get strange languages sprinkled through my Gemini responses, including some very obscure ones. It just randomly changes language for one or two words.
- genewitch 1y agoIs it possible the "vector" is more accurate in another language? Like espirit d'esclair or schadenfreude, or any number of other things that are a single word in a language but paragraphs or more in others?
- sanxiyn 1y agoPossibly. I have seen Claude switching to Russian for a word or two when it is about revolution!
- numpad0 1y agoIsn't it just it getting increasingly incoherent as non-English data fraction increases? Last I checked, none of open weight LLMs has languages other than English as its sole dominant language represented in the dataset.
- spacebanana7 1y agoI suspect this also happens in programming languages. Subjectively I get the feeling that LLMs prefer to write in Python or JS. Would be interesting to see whether they actually score better in leetcode questions when using python.
- beAbU 1y agoBased on my very very limited understanding of how LLMs work, surely they don't "prefer" anything, and just use what they have been trained on? Presumably there is a lot more public info about, and code in Javascript and Python, hence this "preference" Maybe the LLM preferring English is because of a similar phenomenon - it has been trained on mostly western, English speaking internet?
- spacebanana7 1y ago> Presumably there is a lot more public info about, and code in Javascript and Python, hence this "preference" This likely plays a major - probably dominant - role. It's interesting to think of other factors too though. The relatively concise syntax of those languages might make them easier for LLMs to work with. If resources are in any way token limited then reading and writing Spring Boot apps is going to be burdensome. Those languages also have a lot of single file applications, which might make them easier for LLMs to learn. So much of iOS development for example is split across many files and I wonder if that affects the quality of the training data.
- idle_zealot 1y agoAlso worth considering: there's a wider range of "acceptable" output programs when dealing with such forgiving scripting languages. If asked to output C then there are loads of finicky bits it could mess up, pointer accesses, writing past the end of an array, using uninitialized memory, using a value it already freed, missing a free, etc. All things that the language runtime handles in Python or JS. There's a higher cognitive load it needs to take on.
- wongarsu 1y ago
- hnfong 1y agoI'd mentally put this in the same box as "chain of thought", where models perform better when explicitly describing the reasoning steps. The only difference in your case being that the model is undertrained in non-English data, so it's "next token prediction" of non-English prompts is less robust, and thus explicitly converting to English and then back makes it better. This is probably the case for the "deep reasoning" models as well. If you for example try DeepSeek R1, it will likely reason in either English or Chinese (where it presumably is well trained) even if the prompt is in other languages.
- laurent_du 1y agoChatGPT is very informal and talks like a millennial when I ask questions in French. I hate it.
- dingnuts 1y agosorry u hate a whole generation
- mdp2021 1y agoThat's not a "generation", that is a "portrait" (a characterization).
- Der_Einzige 1y agoThe most french response of all... "Euh, en fait.."
- mdp2021 1y agoBritons have "actually", and German speakers spawned the great wave of analytics in the past century (that Britons dominated).
- flir 1y agoOut of curiosity, does vous/tu change its behaviour?
- bee_rider 1y agoIs there a phenomenon where middle-aged people are very informal or slang-y in France? Usually the kids are the ones creating new lingo in English.
- laurentlb 1y agoIn ChatGPT settings, you can set your preferences, e.g. choose between tu/vous, and ask it to be more formal. This should fix your issue, right?
- 1y ago
- mdp2021 1y agoSome studies are trying to ensure that the model reasons through abstractions instead of linguistic representations. (Of course the phenomenon of reasoning in substantially different quality depending on input language signals a fault - reasoning is beyond "spoken" language.) In the past hours a related, seemingly important article appeared - see https://www.quantamagazine.org/to-make-language-models-work-better-researchers-sidestep-language-20250414/ https://www.quantamagazine.org/to-make-language-models-work-...
- jmmcd 1y agoThis important paper from Anthropic includes evidence that part (but only part) of reasoning is cross-lingual: https://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language...
- omneity 1y agoFor most low-resource languages, support in LLMs is trained through translation pairs between english and the other languages, because translation data is easier to come across than say, conversations about coding, history, physics, basically the kind of data that is usually used for instruct training. This kind of training data typically looks like ChatGPT style conversations where all the prompts are all templated like “Translate the following text from X to Y: [text]” and the LLM’s expected answer is the translated text. LLMs can generalize through transfer learning (to a certain extent) from these translation pairs to some understanding (strong) and even answering (weak) in the target language. It also means that the LLM’s actual sweet spot is in translation itself since that’s what was trained in, not just a generalization.
- anon291 1y agoI have observed this and this is what I would expect to have happened thinking from first principles.
- n49o7 1y agoI sometimes dream that they would internally reason in Ithkuil and gain amazing precision.
- quonn 1y agoGiven the fact that LLMs like most neural networks work by passing their input through layers, wouldn't this be expected? There's no going back to an earlier layer and if the first layers are in some sense needed for "translating" [0] to English, any other functionality in those layers cannot be used. [0] I am simplifying here, but it would make sense for an LLM to learn this, even though the intermediate representation is not exactly English, given the fact that much of the internet in English and the empirical fact that they are good at translating.
- dingdingdang 1y agoIndeed. I've thought from the beginning that LLMs should focus specifically on ONE language for this exact reason (i.e. mediocre/bad duplication of data in multiple languages). All other languages than English essentially "syphon" off capacity/layers/weights that could otherwise have held more genuine data/knowledge. Other languages should not come into the picture afaics - dedicated translation LLMs/existing-solutions can handle this aspect just fine and there's just no salient reason to fold partial-multi-language-capacity in through fuzzy/unorganised training.