23 ms·
Open Euro LLM: Open LLMs for Transparent AI in Europe
- NunoSempere 2y agoThey just shipped a frontpage; there is no model yet.
- podgorniy 2y agourl says "Press release" https://openeurollm.eu/launch-press-release https://openeurollm.eu/launch-press-release. They delivered what they promised.
- moffkalast 2y ago> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep relying on Google to do proper multilinguality in open models because Mistral can't be arsed to bother beyond French and German.
- htrp 2y ago>Mistral can't be arsed to bother beyond French and German. Any more details here or a writeup you can link to?
- moffkalast 2y agoMy own experience mainly, only Gemma seems to have been any good for Slavic languages so far, and only the 27B when unquantized is reliable enough to be in any way usable. Ravenwolf posts tests on his German benchmarks every so often in locallama and most models seem to do well enough, but I've heard some claims from people about Mistral's being their favorite models in German anyhow. And I think Mistral-Large scores higher than Llama-405B in French on lmsys and that's at least something one would expect from a French company.
- mhitza 2y agoIn my experience Mistral (at least Nemo) works well with other languages. Don't know about Slavic languages but it does Romanian, with apparent issues around the translation of technical terms.
- jampekka 2y agoWhat do you mean by relying on Google? Llama 3.1 and DeepSeek v3/R1 largest models are rather good at even a niche language like Finnish. The performance does plummet in the smaller versions, and even quantization may harm multilinguality disproportionally. Something like deliberately distilling specific languages from the largest models could work well. Starting from scratch with a "legal" dataset will most likely fail as you say. Silo AI (co-lead of this model) already tried Finnish and Scandinavian/Nordic models with the from-scratch strategy, and the results are not too encouraging. https://huggingface.co/LumiOpen https://huggingface.co/LumiOpen
- moffkalast 2y agoYes I think small languages which have a total corpus of maybe a few hundred million tokens total have no chance of producing a coherent model without synthetic data. And using synthetic data from existing models trained on all public (and less public) data is enough of a legally gray area that I wouldn't expect this project to consider it, so it's doomed before it even starts. Something like 4o is so perfect in most languages that one could just make an infinite dataset from it and be done with it. I'm not sure how OAI managed it tbh.
- simion314 2y ago>As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. Train an LLM with text books and other legal books, you do not need to train it on pop culture to make it intelligent. For face generations you might need to be more creative, you should not need milions of images stolen from social media to train your model. But makes sense that tech giants do not want to share their data set and be transparent about stuff.
- jampekka 2y ago> Train an LLM with text books and other legal books Without licenses to the books, they are just as illegal (and maybe even moreso) than web content.
- troyvit 2y agoIf LLM organizations are free to throw billions at hardware they can spare a paltry €50 million for 10 million e-books though, right?
- jampekka 2y ago€50 million is about 130% of OpenEuroLLM's budget. And I'm very sceptical publishers will give training licenses for €5 per book. Especially as OpenEuroLLM intends to have an openly available training set. Copyright sucks.
- simion314 2y ago>Without licenses to the books, they are just as illegal (and maybe even moreso) than web content. There are books that are out of copyright, and also free books.
- idunnoman1222 2y agoOpen AI fed their original model Anna’s archive for breakfast.
- Fnoord 2y agoI've been using Mistral past week due to changes in geopolitics, and Mistral works absolutely great in English. I haven't bothered in my native language yet, but in English it worked great. Better than my first experience with ChatGPT (GPT 3.5), actually.
- lnauta 2y agoI've been using mistral for most of January at the same rate as chatgpt before. I decided to pay for it as its per token (in and out) and the bill came yesterday... A whopping 1 cent. Thats probably rounded up.
- ur-whale 2y ago> I decided to pay for it as its per token (in and out) and the bill came yesterday... A whopping 1 cent. Doesn't sound too good wrt their eventual profitability.
- Fnoord 2y agoUpdate: tried a couple of Dutch (my native language) queries, and it worked well. No issues whatsoever. Which is no surprise, given Dutch <-> English and vice versa translations often work very well.
- moffkalast 2y agoOk I see we're very far from being on the same page. Multilingualism in context of language models means something more than English, because that's what every model trained on the internet already knows. There aren't any I'm aware of that don't, since it would be exceedingly hard to exclude it from the dataset even if you wanted to for some reason. This is like the "what about men's rights" when talking about women's rights... yes we know, they're already entirely ubiquitous. But more properly I would consider LLM multilingualism straight up knowing all languages. We benchmark models on the MMLU and similar collections that contain all fields of knowledge known to man, so I would say it's reasonable to expect fluency of all languages as well.
- sarusso 2y agoThey allocated €37.4 million [1]. As an European, I truly don’t understand why they keep ignoring that the money required for such projects is at least an order of magnitude more. [1] https://digital-strategy.ec.europa.eu/en/news/pioneering-ai-project-awarded-opening-large-language-models-european-languages https://digital-strategy.ec.europa.eu/en/news/pioneering-ai-...
- everyone 2y agoNot anymore, with Deepseek's stuff right? Which is open.
- blackeyeblitzar 2y agoDeepSeek probably spent closer to two billion on hardware. And then there’s the energy cost of numerous runs, staff costs, all of that. The 5.5m cost was basically misleading info, maybe used strategically to create doubt in the US tech industry or for DeepSeek’s parent hedge fund to make money off shorts. https://semianalysis.com/2025/01/31/deepseek-debates/ https://semianalysis.com/2025/01/31/deepseek-debates/
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- sarusso 2y agoDeepSeek had plenty of R&D expertise which were not included in the (declared) model training cost. Here we are talking about building something nearly from scratch, even if there is an open source starting point you still need the infrastructure, expertise and people to make it work, which with that budget are going to be hard to secure. Moreover these projects take months and months to get approved, meaning that this one was conceived long before DeepSeek, thus highlighting the original disalignment between the goal and the budget. DeepSeek might have changed the scenario (I hope so) but it would be just a lucky ex-post event… not a conscious choice behind that budget.
- jpdus 2y agoAs someone who is in general skeptical of programs like this (and an European) there are 2 remarkable / timely things about this: - This project doesn't just allocate money to universities or one large company, but includes top research institutions as well as startups and GPU time on supercomputing clusters. The participants are very well connected (e.g. also supported by HF, Together and the likes with European roots) - Deepseek has just shown that you probably can't beat the big labs with these resources, but you can stay sufficient close to the frontier to make a dent. Europe needs to try this. Will this close the Gap to the US/China? Probably not. But it could be a catalyst for competitive Open source models and partially revitalize AI in Europe. let's see.. PS: on Twitter there was a screenshot yesterday that in a new EU draft, "accelerate" was used six times. Maybe times are changing a little bit. Disclaimer: Our company is part of this project, so I might be biased.
- FanaHOVA 2y agoThe problem is that: - These are not really super computing cluster in LLM terms. Leonardo is a 250 PFlops cluster. That is really not much at all. - If people in charge of this project actually believe R1 costs $5.5M to build from scratch, it's already over.
- whimsicalism 2y ago> If people in charge of this project actually believe R1 costs $5.5M to build from scratch, it's already over. wdym?
- jpdus 2y agoI think no one believes that R1 costs $5.5m from scratch. People in this project (most, not all) are very aware of the realities in training and are very well connected in the US as well. Besides Leonardo there are JUWELS, LUMI & other which can be used for ablations and so on. This will never compete with what the frontier labs have (+ are building) but might be just enough for something, that is close enough to be a useful alternative :). PS: Huge fan of Latent Space :)
- 2y ago
- mt_ 2y agomaybe 5 million for the html frontend
- Havoc 2y agoIm all for more models. Go for it
- Xelbair 2y agoNo model, so irrelevant because EU is extremely late to the party, as i say that as someone from EU.
- m3kw9 2y agoThey want transparency but they give small peanuts, are they thinking they can just take llama and then distill other models into it can call it transparent? I’m not sure if they understand how that works
- ur-whale 2y ago> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence. translation: weakest competitor in the contest enters the fight with both hands tied behind its back and a budget akin to what OpenAI spends in a week on compute. But hey, more power to you Europe: the more models around, the better. Eventually, we'll be able to bind all those censored models worldwide into one giant mixture of expert to get rid of the built-in censorship of each individual component.
- openrisk 2y agoJust the extensive list of academic partners is a major difference with any existing effort on LLM's.
- phatnguyen 2y ago[flagged]
- isoprophlex 2y agoIs this slop? You're posting slop, aren't you? It's bad. Stop doing that. You're not doing anyone a service.
- calmoo 2y agoNobody wants to read your LLM generated garbage.
- wg0 2y agoUnited we stand, divided we fall. More than LLMs, Europe needs chip autonomy. ASAP. Own fabs, own IP.
- lodovic 2y agoEurope has ARM, AMD and ASML, but everything is produced elsewhere. Time to build giant automated chip fabs.
- thorianus 2y ago"automated" is the issue when producing chips. You can automate some to a certain level of nm. Have you seen the labor needed for maintaing a single modern ASML machine? Also we got a lot, but Japan has the tools to verify results.
- dailykoder 2y agoThere are companies like infineon and elmos in germany (and i think a handful more). So it is happening, just not (yet) with consumer products
- preisschild 2y agoWe dont need "total autonomy" when we have friends that already produce chips, for example. Rather, invest in defense and protect trade with Taiwan. We can do what we are best at and produce lithography equipment at ASML and the Taiwanese at TSMC produce the chips with it and send them back to us.
- flanked-evergl 2y agoEurope as a project is over, it's all downhill from here. You can't regulate yourself into economic growth, and there is nothing else the Eurocrats care to do. There is still lots of fat in the land that they can leech off before Europe kicks the bucket.
- kubb 2y agoIf I got an Euro for every time someone has expressed this opinion, I'd be a rich man. Save it until the EU really doesn't exist.
- fforflo 2y agoIn the EU, we need some social contracts for those things. Multiple EU-funded projects are launched, consortiums between Unis and Private sector companies are created, deliverables are delivered, and grants are allocated, but the cumulative results haven't been that great, have they? That has happened for every "trend," from nuclear physics to "expert systems" in the 1990s to green tech, now AI, etc.
- websap 2y agoHere's the simple question - who gets fired when the deliverables aren't met? This is the single greatest motivators for American companies in our exhausting capitalistic society.
- thatguymike 2y agoSo for €52mn you'll get... a worse Llama? But don't worry, it'll be "transparent and compliant" which will make people want to use it. Very European.
- dailykoder 2y agoThat'd be very very good actually. I'd be happy if institutions would use that where one could TECHNICALLY (maybe just a miniscule amount of people would do that) verify data from end to end, instead of some "open" model that is actually not open at all. A little worse performance is a good trade-off imo
- jamil7 2y ago> a worse Llama? As someone who lives here, I'd actually be surprised if we even got that. I expect lots of taxpayer funded websites, manifestos, PowerPoints and numerous discussions and ultimately nothing.
- malthaus 2y agoyou get money to sustain a bunch of academics and startups past their good-by date the eu gets some publicity and the public gets nothing but another bite out of their taxes
- askonomm 2y agoWhat's with this American mentality that everything needs to always be the best, and if it isn't, it should't even exist? I know USA is alright with breaking the law, invading people's privacy and lobbying its government to the point where it's really the corporations that elect politicians into power, but why do you also need Europe to be the same way? I thought us Europeans have made it pretty clear we don't like your way of governing, so stop forcing it on us. I'd much rather use a less capable LLM if it meant that the LLM isn't driven on top of mountains of illegally collected data.
- seydor 2y agotranslation: a few finetunes of llama plus a lot of travel grants
- cynicalsecurity 2y ago> will build a family of performant, multilingual, large language foundation models for commercial, industrial and public services. The commercial aspect must be implemented ASAP, or it's going to flop.
- isoprophlex 2y agoI'm HIGHLY sceptical. The academics will love it, because they get money. But look at that list of parties involved. More than twenty parties supplying people; none of them will have this initiative on the top of their list of loyalties and priorities. Meaning, everyone will talk, noone will take charge, some millions change hands and we continue with business as usual. Instead this should have been a single new non-profit or whatever with deep pockets that convinces smart people to give their 100% for a while. Death by committee. And I say this as someone who was in a multi-million research program across ~8 universities, that was going to do "groundbreaking" research. After a few months everyone was back to pushing their own lines of research, there was almost zero collaboration let alone common language or goal setting.
- tmikaeld 2y ago” The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence.” They may release something, but i doubt it will be more useful than what already exists.
- bayindirh 2y ago> They may release something, but i doubt it will be more useful than what already exists. I wouldn't put such prejudice in this thing. I'm not implying that you're wrong, but I'm highly skeptical that the model will be incompetent or inferior. Also, don't forget. They'll open source it end to end. From data to training/testing code and everything in between.
- deleted 2y ago[deleted]
- miohtama 2y agoThe model itself might be useful in the end. But it's terrible industrial policy to kill your startups with regulation and then the state needs to step in, because private companies no longer want to work with you, or no new companies are created.
- rixed 2y agoThis article stayed on the front page of HN for a couple of days: https://timsh.org/tracking-myself-down-through-in-app-ads/ https://timsh.org/tracking-myself-down-through-in-app-ads/ The author was in Europe. Aparently, all the rules protecting the privacy of european citizens make no difference in practice. I wonder why, but I believe the EU will look into this soon, since it would be so unconfortable if the king were bad.
- microtonal 2y agoOr it just takes time to enforce the regulations. As an EU citizen, the recent regulations have already helped me a lot - a lot of companies provide data takeout now, it has become easier to remove accounts, many more websites ask specific consent, etc. Or even small things, our daughter's school has to ask specific consent if they can make photos and where they can post them (of group activities, etc.). Does everyone play according the rules? Not yet, but we will get there.
- 627467 2y ago> our daughter's school has to ask specific consent if they can make photos and where they can post them The result of this is we don't see anything out daughter does in school because school decides to comply with draconian regulation by saying "fuck it". The same applies to having parents present daily: we don't touch the grounds of the school unless we make a formal request, we don't see the teacher everyday, we don't hear how the day went from professionals who actually spent time with them. This is all 100% the opposite of our experience outside of Europe before moving and I'm comparing public school system in a third world country to an European one. It's just an anecdote but it hasn't been more clear to me how much in a death spiral the EU is than the experience we currently are having
- probably_wrong 2y agoMaybe it's because I'm not a parent, but what you describe seems to me less "privacy gone wild" and more "Europe vs Non-Europe". The German parents I know wouldn't consider going to the school without reason (kids go either alone or as a group), nor would they expect their teachers to give daily reports. Not because of privacy rules, but rather because you're expected to grow up independent. There are of course regular reports, but talking to the kid's teacher every day would, I believe, get you classified as "oh, that parent". And then there's also the problem of vindictive divorcing parents who take their children away before the other parent shows up. > we don't see anything our daughter does in school If you're talking about photos of the children then I can't imagine a cost-effective way to ensure that photos of your children end up on the Internet while photos of my (hypothetical) children hugging yours do not. But perhaps you have a more precise example in mind.
- glooglork 2y agoEven without the problem of insufficient financing, you probably need to bend some rules to be at the top of LLM game (not just how you collect the training data, we all sort of forgot that Altman was kicked out of OpenAI because the board thought he was prioritizing features over security). With all the regulations and paperwork around EU projects, I can't really see them competing against private sector.
- sirsinsalot 2y agoThe losers in the "rule bending" are humans, whose creative works are turned into weights, whose livelihoods are diminished by indiscriminate greed. I'm all for the advancement of AI, but not at the cost of humility and compassion for those enabling the models to be built. Model building is already a community project, you just weren't asked if you wanted to contribute. You just did. Without compensation. "rule bending" is putting it lightly.
- glooglork 2y agoData is getting scraped, models are being built, nobody is really going to stop that now. You may wish it wasn't like that (not you, but all of us), but there's no way China or USA block their companies in development of key technology like this, and I think we (EU countries) should act in the same way.
- dailykoder 2y ago> will ensure that the models, software, data and evaluation will be fully open I am VERY curious about that. Will they open up ALL the training data? Which would be a massive amount I'd guess, but I'd be curious where they got it and how they got it. Inb4 they just take some meta model and retrain it and then only publish the data from the fine tuned training.
- intellectronica 2y agoHistory repeats itself as farce.
- tobyhinloopen 2y agoBetter late than never. I can't wait to get my hands on a mediocre AI that's 2 generations behind!
- CrimsonRain 2y ago2 is generous. 5 is more likely. Also can't wait to get bombarded with cookie popup, ai bias popup, then ai accuracy popup etc.
- Lionga 2y agoAnything more then a powerpoint coming out of this would be a generous expectation.
- sirsinsalot 2y agoThe alternative is businesses are not held to account. I'd much rather have a cookie pop-up and GDPR notices than businesses have no guard rails against moves that are not in the interest of the user/customer.
- CrimsonRain 2y agoOr hear me out...mandate cookie setting option in browser (no cookie, essential cookies, tracking cookies) instead of fuking prompt in every single site. That also allows me to not let sites ask and force essential cookies everywhere. or block tracking altogether. these cookie banners collectively wasted billions of hours for no gain.
- sirsinsalot 2y agoThank Google. There is no way they'd implement blanket config. Therefore neither will Firefox. Lack of competition for you. Very American, not very EU.
- CrimsonRain 2y ago
- dcreater 2y agoNo Mistral?
- voidr 2y ago> The models will be developed within Europe's robust regulatory framework I'm sure that all AI research needs is "robust regulation". As a European, it annoys me to no end that Brussels bureaucrats think they know and understand everything and they can regulate everything, the only thing they are achieving is making sure that AI companies will avoid forming in the EU, because nobody wants to be at a disadvantage compared to the rest of the world, sure eventually they will provide service to the EU countries, but we will never have our own industry. The EU needs to stop having pencil pushers make decisions on things they have no clue about and somehow get people who know what they are talking about to make the choices.
- ein0p 2y agoJust take R1 base model and fine tune it "for Europe". Done.
- flanked-evergl 2y agoAcademia in Europe is a black hole which turns money into pipe dreams that has no practical application. It's really sad and frustrating, because so many of the pipe dreams sound cool, but it's clear that the people working in academia has no sense for what is needed for industry. The quality of the software they write is shockingly bad, and they have no interest in improving it because they only care about using it for academia anyway. It's like a walled garden where people get fat of poor people who keep everything going, just how it has always been.
- octacat 2y agoShow us results, please
- ranguna 2y agoWith this kind of money, I have a feeling we'll only come up with a good dataset. LLM will be a completely different beast with completely different requirements.
- DocTomoe 2y ago> Europe's leading AI companies Followed by a list of five EU AI companies I have never heard of before and which seem to have little to no market penetration.
- KolmogorovComp 2y agoWill they release the training data? That would really set them apart from other models.
- karol 2y agoNonsense, got something then just share a link to eurogpt. Fed up with this academic BS.
- topazas 2y agoNo name scientists will create an actual product. Plain funny or just sad?...
- zero_k 2y ago"Europe's leading AI companies and research institutions" -- order should be reversed, if the EU is serious about AI. It should say: "Europe's leading research institutions and AI companies..."
- alecco 2y agoThis money would be better spent in producing good training data. But quality datasets don't generate the same administrative overhead or consulting fees.
- Frieren 2y agoAnswering some questions: - Good data is available already for the project. - There are already previously existing models, it is not starting from scratch. - Companies like Red Hat, Volvo, SAAB are part of the initiative thru partners like AI Sweden. So, this is a initiative to support the public sector and universities but also will have commercial output. All that is public information. > The project, which has been awarded the STEP (Strategic Technologies for Europe Platform) seal, leverages support from previous European projects and the experience of the partners and their results, including large repositories of high-quality data and pilot LLMs developed previously. The consortium commences its work on February 1st, 2025, with funding from the European Commission under the Digital Europe Programme. - https://sciencebusiness.net/network-updates/charles-university-coordinate-large-scale-project-large-language-models-transparent https://sciencebusiness.net/network-updates/charles-universi...
- beernet 2y agoThe actual top EU AI labs like Mistral, Black Forest Labs, or Stability AI are nowhere to be seen. Same goes for potent, established companies like SAP, Schwarz Group and the like. They likely made the right move here as this is doomed to fail, as correctly elaborated by the top comments.
- machiaweliczny 2y agoDo they have something tangible now? At least PWr (from Poland) have fine tuned Mistral and a lot of language data AFAIK - https://bielik.ai/ https://bielik.ai/ Anyway I think some novel training method is needed and I’m with Yann on this. They could luck into good idea but with this budget chances are < 1%
- tessitore 2y agoI see plenty of pessimism in the comments, & talk about how unqualified the organizations people are expected to be, without addressing how important this initiative is to EU & how necessary it is they succeed. USA has just proven they are economically unpredictable & so unstable they have become fiscally volatile with control in the hands of the lobbiests. This is why Open LLM has the starting support it does already, & for soverign nations is seen as mission critical so as to avoid long term digital services taxations being leveraged like tarrifs against anyone who does not cooperate with whomever is leading the USA. So to me it feels as though the project is impressive, & quite likely to succeed where others have failed because so few do understand the technology enough to get in the way of progress towards openly standardizing decentralization of AI compute across soverign cloud infrastructures. Even if Aleph Alpha is not able to lead development fast enough, organizations such as OUMI (Open Universal Machine Intelligence) will be working alongside them in attempts to build out the Linuxs of AI frontier modelling. If nothing else, Open LLM guarantees a raise in the social standards of what it takes to succeed in AI long term. At worst it provides a measure bar of success of global AI innitiatives to be compared against while introducing new organizations & people to the open source ecosystem who would never of otherwise invested in it without the EU stamp of approval.