3 ms·
> As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of thes
by jaggederest 11d ago
> As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar.
I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be interested in hearing more if anyone has direct experience.
- mncharity 11d agoFwiw, it's been very easy to nudge Qwen 27B/35B to think in Chinese. No need for <think> prefills like "思考:" ("Think:") or such. Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US means one (dysfunctional:) thing, which similarly distracts. Perhaps if AIs become increasingly multilingual, but remain weak at deep conceptual reasoning, it may be fruitful to have multilingual thesauruses, to use language-associated cultural conceptual differences as a way to convey conceptual nuances with which LLMs otherwise struggle?
- eru 11d ago> In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". For that particular attractor, even German or Spanish or so might help avoid it? Or perhaps even just using British English?
- mncharity 10d agoOh, good point! Something to check. I've been thinking of Chinese as a strength of Chinese models, but other languages might be sufficient here.
- normie3000 11d agoThis is an interesting post, but as a non-American it's hard to understand what's it's alluding to in some parts. The one criticism of NGSS I found from a quick skim of https://en.wikipedia.org/wiki/Next_Generation_Science_Standards https://en.wikipedia.org/wiki/Next_Generation_Science_Standa... is that it teaches evolution. What does "estimation" mean in the US?!
- mncharity 10d ago> What does "estimation" mean in the US?! "Estimation" as a skill in US education, primary school on, is pervasively about point estimates - single numbers. Bounding, a range of safe/reasonable/likely/possible numbers, is very rare. In Fermi questions/problems also. Even in such goodness as Sanjoy Mahajan's Art of Insight and Street-Fighting[1]... the only only "bound" mentioned is "Printed and bound in [USA]", and there is no "round".[2] In contrast, IIUC, China is pervasively about "Da Gu" and "Xiao Gu", "round up" and "round down". Point estimates are secondary. Lower and upper bounds are apparently even a cultural concept: an individual's degree vs capability/fate; policy equity vs excellence; economic targets. In education, my understanding is point estimates have much poorer conversational/collaborative dynamics than range estimates. Bounding lends itself to incremental collaboration. "Can anyone suggest another low or high bound?". Discussions around point estimates seem usually merely about the collection of individual estimates. > NGSS Ah, NGSS isn't bad particularly. I meant that discussion/mention of NGSS appears so very much in English training data, that AI chat on topics related to NGSS seem to draw regurgitations of NGSS. AI "slop" in the sense that "ignore NGSS, just reason from a superset of its underlying concepts"... isn't an LLM strength. > hard to understand Sorry - Thanks for asking! [1] open access: https://direct.mit.edu/books/oa-monograph/5345/The-Art-of-Insight-in-Science-and https://direct.mit.edu/books/oa-monograph/5345/The-Art-of-In... https://direct.mit.edu/books/oa-monograph/5339/Street-FightingMathematicsThe-Art-of-Educated https://direct.mit.edu/books/oa-monograph/5339/Street-Fighti... [2] search "bound": https://www.google.com/books/edition/The_Art_of_Insight_in_Science_and_Engine/xRgeBQAAQBAJ?hl=en&gbpv=1&bsq=%20bound https://www.google.com/books/edition/The_Art_of_Insight_in_S... https://www.google.com/books/edition/Street_Fighting_Mathematics/VrkZN0T0GaUC?hl=en&gbpv=1&bsq=bound https://www.google.com/books/edition/Street_Fighting_Mathema...
- tomaskafka 10d agoThe newly opened possibility to learn“how non-English cultures do X” is criminally underutilized.
- mncharity 10d agoDoes European public dialog do much causal comparisons among states? Does somewhere else? Or do you mean professionally... Instead of merely searching for "ways to do X", one might translate the query into a couple of languages, and translate the results, to get a broader culture sampling. Hmm... That seems something an AI-based search could integrate smoothly - cross cultural insights as a normal part of search results? While there's a lot of diversity in policy and culture among US regions and states, contrasting them isn't leveraged much. Perhaps failing a "that's analysis, not news" threshold for being part of public dialog? So maybe it's not "non-English cultures" neglected as much as comparative discussion in general? One "criminally underutilized" which really struck me, was Japan having a far more sensible and powerful COVID safety one-liner than the US managed - "Three C's" (Closed spaces, Crowded places, and Close-range contact - poorly-ventilated / crowd / conversation-contact - avoid overlapping all three) vs "6-foot rule".