5 ms·
HuggingFace will soon release their BigScience model: https://twitter.com/BigScienceLLM/status/1539941348656168961 https://twitter.com/BigScienceLLM/status/1539
by MasterScrat 4y ago
HuggingFace will soon release their BigScience model: https://twitter.com/BigScienceLLM/status/1539941348656168961 https://twitter.com/BigScienceLLM/status/1539941348656168961
"a 176 billion parameter transformer model that will be trained on roughly 300 billion words in 46 languages"
So anything smaller than that will become worthless. May be a factor, companies have a last chance to make a PR splash before it happens.
Read more about it: https://bigscience.huggingface.co/blog/model-training-launched https://bigscience.huggingface.co/blog/model-training-launch...
- rahidz 4y agoNot necessarily, only ~30% of the database is in English, so it likely won't be as good as a smaller model trained solely or mostly on English words. https://bigscience.huggingface.co/blog/building-a-tb-scale-multilingual-dataset-for-language-modeling https://bigscience.huggingface.co/blog/building-a-tb-scale-m...
- TaylorAlexander 4y agoIt kinda seems like a model trained on multiple languages would to some extent be better at English than a model trained only on English? I mean so much of English comes from other languages, and understanding language as a concept transcends any specific language. Of course there are limits and it needs good English vocabulary and understanding, but I feel the extra languages would help rather than hinder English performance.
- lairv 4y ago"worthless" huh, not everyone can afford inference of a ~500gb models, depending on the the speed/rate you need you might definitely go for smaller model But maybe your sentence was more about "after BigScience model, open-sourcing anything smaller than that will be useless" which isn't necessarily true either, because there is still room to improve parameter efficiency, i.e. smaller models with comparabale performances