5 ms·
> Marius Husnes, the Head of IT Platform at the library (Nasjonlbiblioteket) discussed the project at Huawei’s ID Forum 2026 in Paris, saying that no commercial
by rafram 5mo ago
> Marius Husnes, the Head of IT Platform at the library (Nasjonlbiblioteket) discussed the project at Huawei’s ID Forum 2026 in Paris, saying that no commercial LLM provider was developing a local (Norwegian) language LLM. He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language.
I am not overly confident that Marius Husnes knows what he’s talking about here.
- spiderfarmer 5mo agoIt sounds plausible enough to get subsidies.
- idiotsecant 5mo agoYou're making the mistake of thinking whether he knows what he is talking about matters. He is brewing a potion. It's ingredients are a trendy term, a vaguely spooky threat and a clear, overly simplistic solution that of course he will graciously assume control of, for the good of the motherland. This potion is potent and you'd think it would stop working from frequent misuse but you'd be wrong!
- vintermann 5mo agoHe won't have control over it.
- chvid 5mo agoWhen I am chatting with ChatGPT - it is fairly obvious that it is American - its native language, its style, its attitude is American - even if we chat in Danish. Just as we cannot rely on Netflix and HBO to produce Scandinavian TV-shows even though they might do at the moment, we need to make our own stuff in this area too. And over time, the technology to do this will become cheap and readily available for us to do so.
- anal_reactor 5mo ago> And over time, the technology to do this will become cheap and readily available for us to do so. But then the English models will be even better and you'll be back to square one. My guess is that things are going to become more and more American. If you assume that "culture" is a resource like "microchips", then from economic point of view it makes sense to have one country specialize in producing it, and the rest just consume. This is why when you turn on the main radio station of a random country, you're so likely to hit American music.
- ikr678 5mo ago'Only one country should export culture, for economic efficiency' is the kind of take that the Norweigians (and everyone else) would like to protect themselves from.
- pjc50 5mo ago> then from economic point of view it makes sense to have one country specialize in producing it, and the rest just consume And, for exactly the same reasons as Europeans need to have sovereign compute to protect against economic imperialism, it is also essential to maintain local culture in order to avoid the great replacement of everything with Americanisms. Yes, it requires pushing against the economics. But you have to do that if you believe that culture has any value per se at all.
- wasmitnetzen 5mo ago> If you assume that "culture" is a resource like "microchips" I do not. American culture exports American values, which are not universal. Simplest examples being the attitudes towards violence and nudity, which are very different in Europe, and vary within Europe as well.
- anal_reactor 5mo agoWhich is already changing thanks to the American influence.
- 5mo ago
- fnordpiglet 5mo agoHe’s right though, although it’s not entirely about the training corpus. It’s about the tokenizer that tokenizes substrings more efficiently based on a necessary bias towards a target language. English oriented LLMs are more powerful for English than other languages because the token space is more parsimonious in English language. Try any online Anthropic tokenizer that calls their api with common English words (typically one or fewer tokens) and Norwegian words - you’ll often see 2-4 tokens instead sometimes more. Some languages like Thai are at a huge disadvantage. Likewise often the corpus selection also is heavily skewed towards the target language simply because more energy is applied to sourcing written works in that language. There will also be semantic biases in the vector space due to cross influence between semantically similar embeddings between languages that create a different than cultural baseline. Finally fine tuning greatly impacts cultural expression in the LLM. None of these are trivial effects. There are a lot of efforts to create LLMs for dying languages and others that use cross cultural models to boost, but if your language is well literate, there’s a good reason to build a heritage LLM specific to your language and culture. Expecting OpenAI or Anthropic to prioritize your language over their target audience when a tradeoff is to be made is absurd.
- deleted 5mo ago[deleted]
- YetAnotherNick 5mo agoDid you even try to verify your claims. I tested it on few translations on wikipedia articles using [1] and it takes 15-20% more tokens for Norwegian. English performs the best because there is more data in English and high quality sources are either only in English or there is a good translation in English. [1]: https://platform.openai.com/tokenizer https://platform.openai.com/tokenizer
- tecleandor 5mo agoTests I've done with NO and FI texts, for the same number of characters, with the GPT5 tokenizer I get around 2x the tokens than EN. With the older tokenizers it's more like 2x or even 3x.
- isawczuk 5mo agoPoland have its one LLM called Bielik. It's not only better in preserving Polish sounding wording, it's also better in writing government documents. Why better? They did arena and statistically it's just better.
- maxloh 5mo ago[dead]
- KaiserPro 5mo agocould you provide evidence to suggest he is wrong? It seems like you've made an assertion but not provided evidence. Why is it not a disadvantage to only have english LLMs? Can you get the nuance of Norwegian history/culture with present models?