6 ms·
I keep seeing these "sovereign" LMs time and time again. In Sweden we had GPT-SW3 (https://www.ai.se/en/project/gpt-sw3 https://www.ai.se/en/project/gpt-sw3) an
by armcat 4mo ago
I keep seeing these "sovereign" LMs time and time again. In Sweden we had GPT-SW3 (https://www.ai.se/en/project/gpt-sw3 https://www.ai.se/en/project/gpt-sw3) and same story there. Instead of burning money on "sovereign" claims, national research labs should instead focus on building on top of solid baselines (like Qwen/Kimi) and finetuning frontier models with real agentic utility that can be applied across actual use cases and can be widely used by its people, basically for free. Nations should mirror what Cursor has done with Composer 2.5 for example.
- thevinter 4mo agoAnd what happens once the "solid baselines" become unavailable for a reason or the other?
- zozbot234 4mo agoYou keep building on the last available version? Fine tuning is a whole lot cheaper, easier and more useful than pretraining a model from scratch. It's a complete no brainer.
- rapidfl 4mo ago> You keep building on the last available version? yes but a sovereign can allocate some resources and a few people to stay in the loop from a first principles level. No need to wait for a rug pull. Of course, it can not compete with the frontier labs. But good to have researchers and professors "in-house". LLMs are here for the long-term.
- michaelscott 4mo agoUnfortunately in this game first principles requires massive resources, not "some". Building in-house on top of existing open weights is a good way to bootstrap this process, especially since there's nothing inherently magical or particularly expertise-heavy when it comes to weights themselves
- GTP 4mo ago> But good to have researchers and professors "in-house". I'm not in this field, but I think we already have them. Probably the main difference is that we have most or all of them in academia and next to none insode private companies. But we do have them, and they could start working for private companies if the market moves in that direction in the EU as well
- ozim 4mo agoSeems like you don’t understand. You take current version and build on top of it. You have the weights. You might not get some n+1 version at some point but the n version you will have will be still most likely much better than whatever you come up with burning good will money of people believing in „sovereignty”. You are not getting ahead in this game by being „true to your local values” capital expenditure is insane in this game.
- oneshtein 4mo agoIt seems like you don't understand. For fine tuning, it's cheaper to fine tune an existing model. For massive changes, it's better to retrain from scratch. Otherwise, model will UNLEARN a lot first, and then you will train about twice longer to the same result. https://en.wikipedia.org/wiki/Catastrophic_interference https://en.wikipedia.org/wiki/Catastrophic_interference
- ozim 4mo agoOk this „you don’t understand” was uncalled for, I am sorry. I am speaking from very practical point of view. English and whatever frontier models are trained on is lingua Franca of software/tech/science currently - you don’t want to make massive changes because just exactly like you wrote model will unlearn a lot. Current models translate easily between the languages. So from my point of view even if we as a smaller country will have N-2 model that we slightly fine tune or just give it a harness with national RAG it will be better than wasting money on training a model from scratch only on „texts in country own language „ because that is loosing proposition. It’s usefulness is going to be really limited compared to model trained on body of knowledge in English. It is a lot like company CEO thinking they have to „train AI model” for their company on their company materials - well no, you just make RAG and eventually fine tune some models, give access to dat, give access to MCP. Because even if you are F1000 company you don’t have resources to train your own generic model, as a nation there is no that have resources to train „national model” on par with frontier labs.
- deleted 4mo ago[deleted]
- mschuster91 4mo agoKimi and Qwen come out of China, which means that their training material may be biased e.g. relating to Taiwan [1]. In addition, there is no way to determine what input went into the training, if it was properly licensed, if it was legal (e.g. not contaminated by CSAM), or how the human component of RLHF was sourced - in US models, for example, stories about exploitation like [2] have been floating for years. Assuming us Europeans finally get our act together, I think it is better for our long-term future (and the ethical problems) if we manage to get a baseline of training input and data ourselves, from scratch, with everything being ethically sourced. Oh and, while we're at it, the EU has 24 official languages plus a host of minority languages. Most LLMs focus on the English, German, French and Chinese languages, but everything else is... left behind at best. An European model with actual funding and proper data sources might be able to significantly reduce that. [1] https://www.taiwannews.com.tw/news/6245677 https://www.taiwannews.com.tw/news/6245677 [2] https://www.theguardian.com/technology/2024/apr/16/techscape-ai-gadgest-humane-ai-pin-chatgpt https://www.theguardian.com/technology/2024/apr/16/techscape...
- siva7 4mo agoUh, some would say it's easy to determine what input went into the training for kimi and qwen.. since they were caught stealing it from American labs. Some cultural cliches may never change.
- ignoramous 4mo ago> since they were caught stealing it from American labs. Some cultural cliches may never change. Has a formal lawsuit been brought to bear? Given, Anthropic & OpenAI are being dragged through courts for copyright violation (or stealing, as you'd call it, if the companies involved were culturally Chinese) by newspapers, publishing houses etc; one'd think they'd pass on some of that medicine to Alibaba, which does have business entities registered in the US.
- janc_ 4mo agoIt's well-known that all commercial models are based on stolen content. That doesn't mean there is no filtering/censoring, just that the censoring likely depends on where it's happening…
- TJSomething 4mo agoIf open frontier models start closing up and states start more export controls on AI services and hardware, it might be good to ensure the supply chain is there to reproduce the SotA, or even a couple generations behind it.
- 627467 4mo agoDo we know for sure how much national corpus of knowledge (like dutch) goes into these "global" models and how that affects "localized" model biases? What's wrong with specialized models?
- appplication 4mo agoDisagree, it’s in the country’s best interest to facilitate internal expertise on the full stack and own their “supply chain” so to speak and fight brain drain. The outcome isn’t just the model, it’s the expertise. Otherwise all their smartest folks will depart for countries where LLM development is strongest.
- bradhe 4mo agoGot it, so every country should focus on having a mediocre-at-best AI strategy by refusing to work together? Surely this will create a better future instead of pooling resources. This vaguely-nationalist world view around tech that’s emerging in Europe is dangerous, man. On the brain drain problem in particular, one way to ensure talent sticks around is to create a good environment for people to do their best work. In much of Europe, getting bureaucracy out of the way and encouraging real investment would go a long way. People leave because they can make more money and they want to be surrounded by the best people. People would trade some of that off to stick around their home countries, however if you go to California and talk to folks from e.g. NL or DE working on this stuff, they have a lot to say about innovation and working culture back home.
- Gud 4mo agoWhy would it have to be mediocre? Both Netherlands and Sweden produce highly competent researchers. Per capita on par with any other location on the planet, including California. This will be good for Europe as a whole and also the planet.
- colordrops 4mo agoSure, but there is still a more dense talent pool with SV companies, and training frontier models requires massive amounts of capital. Are Netherlands and Sweden prepared to invest 10s of billions?
- Gud 4mo agoHonestly I have a strong suspicion that “the talent pool” will shift towards Europe and China and I believe it’s already happening. And yes, I do believe Europe will invest in this technology.
- khafra 4mo agoToday, you keep seeing "sovereign" LMs that are subject to the sovereignty of some human-led state. Tomorrow, the "sovereign" LMs will be called that for a completely different reason.
- Scarblac 4mo agoTheir legality is very questionable given all the likely copyright infringement going on, and a state can't really ignore that.
- vintermann 4mo agoStates are the things which can ignore that, and I'm pretty sure US and China already do. No state is going to respect copyright if they think its future is at stake, and apparently even Netherlands thinks the future is at stake. (Of course states can ignore copyright in a legally polite manner, such as asserting that training on all published material in the National Library is fair game)
- entropyneur 4mo agoSounds like calling those model "open source" did its wicked job. You can not take an open weight model and build a next-generation model using that as a foundation. Once those companies decide it's no longer in their interest to release new open weights everything you've created this way becomes a pile of rapidly deprecating legacy.
- michaelscott 4mo agoOP probably means using the existing open weights as a base for further homegrown development and research, not that the homegrown models are always updated based on whatever US or China are doing in the moment
- teekert 4mo agoThere is something to be said for this "most cheapest" approach, there is also something to be said for making models that are entirely ethically sourced: 1. Free of controversy like unlicensed training materials 2. Free of exploitative rlfh loops by people in low-wages countries 3. The leasons learned (and published) from going through the entire training process on "European" hardware: "AI factories" (the term for Slurm HPC/HTC systems with lots of heavy GPU nodes, heavily subsidized by our government [0]) 1 and 2 are strong counter-LLM arguments at the moment, and hold back some groups of potential users. Another is energy/water use, so going for maximum green energy would be a nice boon as well. 3 is something I consider to be highly useful for our European identity and "way of the ninja" (for you Naruto fans out there). [0 https://hpc-portal.eu/funding-opportunities https://hpc-portal.eu/funding-opportunities]
- dijksterhuis 4mo agoone would also hope there'll be less pressure to "make line go up", i.e. not having to do attention-engineering via deliberate sycophancy to trap individuals into using it more and more and more and more and more and more. but in general, yes, as someone who is vehemently anti-ai GPT-NL has piqued my interest specifically because of the ethical protections / measures they're talking about. question is whether they stick to it.
- saidnooneever 4mo agoTNO doesn't serve nor represent global interest and hence does not care about global progress. It exists to enhance knowledge on things within Dutch society primarily with some ripple effect outward to EU because they have interests within the EU. Its purpose is not to become some kind of OpenAI or global foundation offering services/tools on that scale. There is a lot of critisism on this project, not invalid, but mostly based in lack of understanding of what the goals are of the organization as well as the people building the thing. The people building it, are well aware of how it will be less capable than other LLMs on a general reasoning aspect, not only due to having actually purchased _all_ licensed data that has been used as inputs. Not being a multi billion dollar corporation, this means having very little data and should be an obvious signal to observers that it has not the goal to outdo other models. In my opinion (personal) its a project that has a learning and demonstration value that is not 'look how well our model performs against others', but still offers value.
- enaaem 4mo agoSame reasons why every country, or close allies, build their own tanks or space program. You want to keep some level of capability within your control. Compared to weapon programs, AI research is very cheap.