20 ms·
Mozilla Common Voice Adds 16 New Languages and 4,600 New Hours of Speech
- pkz 5y agoOpenly licensed speech data for smaller languages is great! I hope as many as possible contribute in order to get better representation across ages and pronunciation. In the end, this may be what is needed for the hyperscale companies to support speech assistants in more languages?
- satya71 5y ago> The top five languages by total hours are English (2,630 hours), Kinyarwanda (2,260), German (1,040), Catalan (920), and Esperanto (840) Some unusual suspects among the top languages, there!
- ftyers 5y agoThat's what happens when people have the opportunity and tools to support their own languages and not just rely on hand outs from big tech :)
- Anon1096 5y agoEsperanto is a hobby language for upper-middle class people in developed countries. It isn't anyone's "own language".
- bradrn 5y agoWell, it has native speakers: https://en.wikipedia.org/wiki/Native_Esperanto_speakers https://en.wikipedia.org/wiki/Native_Esperanto_speakers
- krrrh 5y agoTechnically there are a few hundred L1 speakers of Esperanto, but that doesn’t really contradict your point. https://cogsci.ucsd.edu/~bkbergen/papers/NEJCL.pdf https://cogsci.ucsd.edu/~bkbergen/papers/NEJCL.pdf
- stegrot 5y agoYou are not wrong, but besides the upper-middle-class hobby people, there is also a 130 years old culture that exists parallel to it. I've met a few native Esperanto speakers, and for them Esperanto is their identity. Traditional Esperanto clubs exists in countries like Iran, Japan, China, Burundi, Nigeria and many more. So Esperanto is both, a nerdy hobby and an old culture.
- hkt 5y agoWeirdly judgemental. Esperanto was designed to be easy to learn. It isn't an elite pursuit in the way you suggest, because its community isn't gatekept. I personally have met people of all social classes who have been interested in it. It was also never meant to be a first language, it is an auxiliary language. It is possible for an English speaker to have a conversation with a Mandarin speaker with no intermediary if both know the (comparatively easy to learn) Esperanto. Its original purpose wasn't trivial either: it was created to stop groups without a common language in the same city (Warsaw, I think?) fighting, created on the basis that they'd stop doing so if only they could speak a common language. Think of it as JVM bytecode for people.
- least 5y agoAuxiliary languages are kind of inherently doomed to fail to function as they're intended because in order for them to function as such, commitment needs to be made to adopt it multilaterally by governments with sufficient influence. If today the United States and China bilaterally decided to force Esperanto into their school curriculum it'd likely be adopted very quickly by everyone else, but that isn't the case and I doubt it ever would be under almost any circumstance, because learning English is just immediately more practical, even if it's a significantly more difficult language to be picked up. And that's how it's played out. Nearly every developed nation teaches English as a second language or is a native population of English speakers. The universal language is English. The JVM bytecode for people is English.
- hkt 5y agoSpoken like an anglophone. Tell that to Latin America and East Asia..
- least 5y agoI don't have to, you can look at pretty much any of their language curriculum and find a huge presence of English in nearly all their education systems. Certainly you will find people learning other languages for trade depending on the region, but even in East Asia, as you say, English is taught in China, Japanese, Korea. In Singapore English is the language everyone learns (and is taught in). In Vietnam the primary foreign language taught is English. In the Philippines one of its official languages is English. Argentina teaches English in elementary school. In Brazil students from grade 6 have to learn a language, which is usually English. In Venezuela English is taught from age 5. So what exactly do I have to tell them?
- samtheDamned 5y agoThey weren't exclusively talking about Esperanto. I read it as a reference to Kinyarwanda and Catalan more than anything else. In the bigger scheme of things there are a lot of languages here that are definitely a product of being able to share your own language. There's multiple native languages that are being shared here, like the thread above about Guarani.
- ndkwj 5y agoIs "upper-middle class in developed countries" meant to be an expletive?
- crvdgc 5y ago> Esperanto is a hobby language for upper-middle class people in developed countries. I wonder what gave you such an impression of Esperanto. My personal experience of Esperanto is quite different. I started to casually self-learn Esperanto about one year ago as my second foreign language apart from English. After about half a year, I was confident enough to join online Esperanto communities and it gave me a surprisingly much more diverse experience than any community I had encountered on the Internet. For example, in an online chat group, active users mainly come from US, South America, and Russia. As an person from East Asia, there is little chance for me to get in touch with the latter two groups otherwise. And there are often new users from South America who speak only Spanish and Esperanto. I myself do not identify as a upper-middle class person, and I don't know enough to assess other Esperanto speakers' class status. The impression of Esperanto speakers being upper-middle class may come from the fact people learn Esperanto as a hobby. But people not in the upper-middle class can have other hobbies, why is Esperanto different? It doesn't come with the many benefits that people may expect from learning a "practical" language, but it takes significantly less effort. I'd say it's about as hard as learning a new instrument. So it is not that exclusive to only upper-middle class people. After one year of casual learning, I am now able to contribute to the Common Voice project in Esperanto (175 recordings and 123 validations) and I actually use it as a source of learning material.
- jpetso 5y agoYou must be a fast learner. After one year of learning a new language, I personally would not feel comfortable speaking it well enough to use as examples for others.
- crvdgc 5y agoThanks to the design of the language, each letter of Esperanto has a fixed pronunciation, and the stress is always on the second-to-last syllable. So after you learn the alphabet and some diphthongs, you are able to pronounce every Esperanto text in the canonical way (even if you don't know a single thing about the meaning). No exception. This is also a great feature for self-learning. Of course, it takes time to fluently "read out" the words, and in practice, it's much easier if you just know the word and pull the pronunciation from your memory. For the Common Voice project, there are usually two or three words in a batch of five sentences that I don't know. And there are unfamiliar places and names, since most of the text come from Wikipedia. In such case, I'll take my time to use the spelling to infer the correct pronunciation and practice it several times, until I can put it into the sentence. Then I'll record. And I know it must be correct. If I am not sure about the meaning of the new word (you can usually guess from etymology or word formation), I look it up in the dictionary and learn a new word.
- 1-6 5y agoYou have a point there. I've been disappointed that Korean has been stuck in the 'In Progress' state. The Korean tech giants already have APIs to do common speech recognition. I hope more Korean grassroots efforts focus on tools that are open and accessible so it can be built scalable and better.
- yorwba 5y agoIt looks like Korean still needs a fully localized interface and a sufficiently large collection of sentences to record. You can help by translating the interface https://pontoon.mozilla.org/projects/common-voice/ https://pontoon.mozilla.org/projects/common-voice/ and collecting public-domain sentences https://commonvoice.mozilla.org/sentence-collector/ https://commonvoice.mozilla.org/sentence-collector/ and of course by getting Koreans you know excited about the project so they'll help, too.
- fleaaaa 5y agoThank you for pointing it out, I had no idea but I'd happy to contribute on this one. There is indeed a decent korean natural language process engine but it's severely tied to own ecosystem AFAIK. https://papago.naver.com/ https://papago.naver.com/
- umeshunni 5y agoAh yes, major world languages with 10s or 100s of millions of speakers (Bengali, Korean, Malayalam) are ignored or are perpetually stuck "in progress" while hobby languages like Esperanto are supported.
- stegrot 5y agoHey, I work on the Esperanto version of CV. You are right, many languages should be bigger than Esperanto, and we never planned to become this big, it just happened. We are around ten active people and a telegram group with a few hundred motivated donors. Plus, we write about the project in Esperanto magazines and talk about it on Esperanto congresses. The point is: the only reason Bengali Korean and Malayalam are stuck "in progress" is that no one is working on them. No language but English is actively supported by Mozilla, it all comes from the communities. And the success of Esperanto shows that every language can make it. I hope that people take our work as a motivation. Every language can become big if a few motivated people work on it for a year or two. Even the smallest language can make it. You just need a lot of public domain sentences, a few thousand donors and some technical knowledge then your language will grow as well :)
- umeshunni 5y agoSure, I was responding to the factitious comment above. When I can use Google or Facebook in any of these languages for 10+ years, it's silly of this project to claim some high moral ground when you can't support some of the most widely spoken languages in the world and stick to languages that hipsters in San Francisco think is cool.
- yorwba 5y agoIt can support those languages, they just need some people who actually speak them to come along and make it happen. If you can help, I'm sure it will be appreciated.
- yorwba 5y agoThe project seems to have some serious government backing in Rwanda: https://digitalumuganda.com/ https://digitalumuganda.com/
- junon 5y agoHistorically not been the biggest fan of Mozilla but I really, really love this project. I'm glad they're keeping it alive.
- ftyers 5y agoOne of the most noticeable additions in my opinion is Guarani, the first Indigenous language of the Americas to be added. Indigenous languages are extremely poorly supported and forgotten by all of the major platforms and companies, and it's great to see one getting the attention they deserve. (Disclaimer: I was involved)
- neartheplain 5y agoWhoah, 6.5 million native speakers! That's several orders of magnitude more than I was expecting. It's also significantly larger than the native-speaking populations of languages like Catalan, Basque, or Romansh, which might be more familiar to North Americans or Europeans.
- victorlf 5y agoCatalan has about 10 million speakers.
- pimterry 5y agoIn total, yes, but only about 4 million _native_ speakers.
- djoldman 5y agoOr, 20x more than Icelandic: https://en.wikipedia.org/wiki/Icelandic_language https://en.wikipedia.org/wiki/Icelandic_language
- hkt 5y agoWithout wishing to get political, is the difference that Iceland is a country but Guarani speakers don't have a nation-state of their own? Or something else?
- interactivecode 5y agoThe difference is completely and inherently political.
- danShumway 5y agoI don't really have anything of substance to add here, but I'm very happy to see Mozilla continuing to put effort into this, happy to see effort being put into broadening the support beyond just English and major languages, and I'm grateful for the work that people (inside and outside of Mozilla) have already put into getting the project this far.
- orra 5y agoIndeed, it's great to see open data corpuses expand.
- mgarciaisaia 5y agoYou arguably have something of substance to add - you can help improve the datasets by speaking or validating phrases in the project's website https://commonvoice.mozilla.org/ https://commonvoice.mozilla.org/ There are many languages available to pick from.
- S5yDyAk3XoQH5 5y agoEveryone has something to add. Go read stuff on their site in the languages you know. If people would actually do it for a few months all languages would have been done years ago.
- _gtly 5y agoA direct link to where you can donate your voice here: https://commonvoice.mozilla.org/en https://commonvoice.mozilla.org/en
- jalopy 5y agoGoing along with this: What are the latest and greatest open source speech-to-text models and/or tools out there? Would love to hear from experienced practitioners and a bit of detail on the experience. Thanks HN community!
- orra 5y agoMozilla announced Deep Speech[1] around the same time as Common Voice. Mozilla Deep Speech is an open source speech recognition engine, based upon Baidu's Deep Speech research paper[2]. Unsurprisingly, Deep Speech requires a corpus such as... Common Voice. [1] https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech [2] https://arxiv.org/abs/1412.5567 https://arxiv.org/abs/1412.5567
- zerop 5y agoVosk is my favourite. I have used deep speech too. Vosk works better.
- nshm 5y agoThank you. I deeply appreciate you mention our efforts. We spend quite some time and knowledge to build accurate speech recognition. Not that easy to get as much mentions as Mozilla, so we are thankful for every single one!
- 5y ago
- donhaker 5y agoLet's take the time to appreciate the effort of Mozilla. To add new languages with others came from the minorities, we can't deny that they are continuously putting effort into the community.
- Jnr 5y agoThe great open source community around Mozilla helps a lot. When I did not see my own language in the list a year ago, and I had no clue how to get it there, I reached out to my university contacts that I know used to translate Firefox years ago. With their help we quickly translated the whole common voice site (it was a prerequisite to start contributing a language) and provided first sets of text to start contributing. In about a week we started contributing voice for a new language. The Common Voice project is awesome and very well made.
- fisxoj 5y agoIf anyone has interest contributing, I've found this app for Android makes it very easy! https://www.saveriomorelli.com/commonvoice/ https://www.saveriomorelli.com/commonvoice/
- nmstoker 5y agoWhy on Earth would anyone use an app for this when mobile browsers work perfectly well for adding audio to Common Voice? We could possibly give the developer the benefit of the doubt that they're not doing anything inappropriate with the data but frankly why pass your data through a third party that's not part of the project. And why install an app requiring access to your shared local storage? The GitHub repo claims the website an animations are slow which sounds like BS to me. It works fine on a five year old phone I use for submitting. Just contribute here if you're so inclined, much more sensible: https://commonvoice.mozilla.org/en https://commonvoice.mozilla.org/en
- totetsu 5y agobecause mozilla fired all the cv team, and the app is under active development?
- nmstoker 5y agoYou aren't distinguishing the projects correctly. The CV project isn't the same as the DeepSpeech project (even though they were related). And your point makes little sense, because if the site was not working how could the app get voice data into the project. I've had some involvement with these projects over the years so I'm not just firing off arm-chair comments on this. They wouldn't have been able to add this new voice data if the site was under developed as you imply.
- totetsu 5y agoso what happend since this? https://discourse.mozilla.org/t/mozilla-org-wide-updates-impacts-on-common-voice/65612 https://discourse.mozilla.org/t/mozilla-org-wide-updates-imp...
- rasz 5y agoWhats the point when they killed DeepSpeech in exchange for adapting closed Nvidia thing? https://venturebeat.com/2021/04/12/mozilla-winds-down-deepspeech-development-announces-grant-program/ https://venturebeat.com/2021/04/12/mozilla-winds-down-deepsp... https://blog.mozilla.org/en/mozilla/mozilla-partners-with-nvidia-to-democratize-and-diversify-voice-technology/ https://blog.mozilla.org/en/mozilla/mozilla-partners-with-nv... $1.5mil for shutting down open source initiative, almost half of CEO salary right there.
- moralestapia 5y agoLol, these guys sell themselves for peanuts.
- jononor 5y agoWhat closed NVidia thing did they adopt? I don't see any evidence of that here.
- rasz 5y agoIm sure they signed on adopting "something", otherwise it would be receiving $1.5 million grant for closing open source initiative. $3 million a year lawyer wouldn never be this blatant.
- option 5y agohttps://github.com/NVIDIA/NeMo https://github.com/NVIDIA/NeMo which is open source, Pytorch based and regularly publishes new models and checkpoints.
- Seirdy 5y agoThe source code is under a FLOSS license, but it only works on Nvidia GPUs and uses proprietary Nvidia-specific technologies like CUDA. It's significantly closer to "nonfree" on the free-nonfree spectrum than it should be, and is another example of the difference between the guiding philosophies behind "free software" and "open source"
- russian_nukes 5y agoWhat is this voice database? Do they have russian voices?
- deleted 5y ago[deleted]
- dabinat 5y agoCommon Voice is a great project that I’m glad Mozilla kept alive. One problem is that data for speech recognition needs to be extremely accurate (i.e. the speech matches the transcript perfectly) and the human review process is infallible and there are quite a number of bad clips that made it past the review process (to be fair, Mozilla provides no official guidance to reviewers or recorders). Plus in the early days, they were recording the same small sentence pool over and over again, so the first 700 hours or so are duplicates. I hope there will be efforts in the future to clean up the existing dataset to improve its quality.
- ma2rten 5y agoWhy does data for speech recognition need to be prefect. That's certainly not the case for other machine learning applications. Can you train the less clean data and fine-tune on a clean subset?
- stegrot 5y agoHere are some draft guidelines for validation that have been translated a lot: https://discourse.mozilla.org/t/discussion-of-new-guidelines-for-recording-validation/36465 https://discourse.mozilla.org/t/discussion-of-new-guidelines... But you are right, the process has some flaws. Maybe we can review the dataset automatically on some common errors, once an STT system is ready for a language? The only other option I can think about is a validation process that includes more people per sentence. Right now, only two people validate a sentence, and if they disagree a third person decides. We could at least double check sentences with one "no" vote one more time.
- johnnyApplePRNG 5y agoJust tried rating some of the English voices and I am conflicted. Most of them were definitely speaking English, but in an Indian intonation that I was barely able to understand coming from an English as a First Language country. Some of them were reading words syllable by syllable, which is definitely English, but I would hate to have to listen to an ebook or webpage read aloud to me in that manner. By clicking yes am I training the system to speak English with an Indian intonation? Should I click no, not English? Should/does english even have a "proper" intonation?
- ma2rten 5y agoI think this dataset is mainly for speech recognition and not text to speech. Speech recognition should be able to recognize as many different accents as possible.
- jturpin 5y agoWow you're right. This is conflicting as many of the words are not pronounced properly at all. Maybe it doesn't matter to the accuracy of the speech-to-text system, but it feels like training it with bad data.
- ohgodplsno 5y agoDifferent accents isn't bad data. Your vision of the world of "english is only spoken with an american accent" is what leads to horrendous speech recognition APIs, like Google's. If your ML model can't handle multiple accents, it is worthless.
- jturpin 5y agoThere's a difference between an accent and pronouncing words wrong. I would expect an English speech recognition system to handle the various accents there are in the world (the US has several accents of course), but it shouldn't handle incorrect pronunciation of syllables if it comes at the expense of recognizing clean data. If it doesn't come at its expense then I guess it's fine.
- tsjq 5y agonice. ! news from the past about this : Initial Release of Mozilla’s Open Source Speech Recognition Model and Voice Data : https://news.ycombinator.com/item?id=15808124 https://news.ycombinator.com/item?id=15808124 Mozilla releases the largest to-date public domain transcribed voice dataset https://news.ycombinator.com/item?id=19270646 https://news.ycombinator.com/item?id=19270646
- bravura 5y agoIs anyone aware of classification (e.g. word prediction) datasets for low-resource and endangered languages? If so, we would like to use it for the HEAR NeurIPS competition: https://github.com/microsoft/DNS-Challenge/tree/master/datasets_fullband/noise_fullband https://github.com/microsoft/DNS-Challenge/tree/master/datas... The challenge is restricted only to classification tasks, and sequence modeling like full ASR is unfortunately beyond the scope of the competition.
- LoriP 5y agoTips & Tricks incoming... I find that if I can't sleep and want something that's kind of useful to do without getting too involved, contributing to common voice is a great way to spend half an hour and relax/forget whatever it is I was churning about. I would recommend it for that, plus it's a great project. Both listening and voicing...
- fareesh 5y agoIs voice transcription accessible to mere mortals yet? I have tried pretty much every API offered by big tech, and also various open source models. All of them seem to have incredibly high word error rates. This is mostly for conversations with various Indian accents.
- nshm 5y agoDid you try Vosk Indian English model? It is specifically built for Indian accent English https://alphacephei.com/vosk/models/vosk-model-en-in-0.4.zip https://alphacephei.com/vosk/models/vosk-model-en-in-0.4.zip In case you want more accuracy you can share a file with an example, we can take a look on how to make the best accuracy. For Indian ASR it is also worth to mention recently introduced Vakyansh project which builds model for major Indian languages: https://github.com/Open-Speech-EkStep/vakyansh-models https://github.com/Open-Speech-EkStep/vakyansh-models
- alpb 5y agoThis may be off-topic but: What's the relationship between Coqui (an OSS TTS startup) https://coqui.ai/about https://coqui.ai/about and Mozilla? I recall that the project at one point was called mozilla/TTS (https://github.com/mozilla/TTS/ https://github.com/mozilla/TTS/) and now I see that has a fork in the startup's own repo (https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS). Presumably Common Voice is used to train mozilla/TTS and other OSS TTS solutions?
- ftyers 5y agoCommon Voice is mostly used for STT not TTS. TTS requires single speaker, clean audio. STT requires multi speaker, noisy audio.
- say_it_as_it_is 5y ago"The top five languages by total hours are English (2,630 hours), Kinyarwanda (2,260) , German (1,040), Catalan (920), and Esperanto (840)." How did they get almost as much training for Kinyarwanda as they have English?
- stegrot 5y agoThe German Federal Ministry for Economic Cooperation and Development supported this language: https://www.bmz.de/de/aktuelles/intelligente-sprachtechnologien-fuer-afrikanische-sprachen-82916 https://www.bmz.de/de/aktuelles/intelligente-sprachtechnolog...
- say_it_as_it_is 5y agoInteresting! There's a market for this kind of audio data entry? What was the total cost for that many hours? The English data was entirely volunteer driven, correct? Maybe it's worth funding the English corpus for the additional hours needed to reach the sweet spot?
- nshm 5y agoData cost plunges these days with self-supervised and semi-supervised learning. You don't need annotated and clean data anymore, there is abundance of it. Projects like Voxpopuli or Gigaspeech with 400 thousand hours (100 times more than Mozilla's) of data easily available.
- nyx-aiur 5y agoI love the datasets but they are still way to small especially for exotic languages.
- deleted 5y ago[deleted]
- Edman274 5y agoI'm guessing that of the 4,600 new hours of speech, maybe 4,100 of those hours are of men's voices and 500 hours are of women's voices, yeah?
- LoriP 5y agoTo be fair not sure that's the best guess :) there seem to be more female voices than men to me. Anyhow, I'd wager there's at least a 50:50 mix.
- Edman274 5y agoI probably should've looked this up before I decided to comment but at least according to this: https://commonvoice.mozilla.org/en/datasets https://commonvoice.mozilla.org/en/datasets The ratio of male to female tagged voices in the English dataset is 45 percent male to 15 percent female. (The remaining 40 percent is untagged.) Odds are good that the ratio is closer to 75 25 than 50 50, at least by hours of recorded audio.
- heyhillary 5y agoThanks so much for sharing your comment. Gender equality in participation in Common Voice, is something we really want to improve and champion. As part of the Kiswahili Language community engagement, our team are implementing a gender action plan that includes both participation and use cases for the dataset. We hope to consult, adapt and replicate gender inclusion that has been done by community members and gender action plan to improve representation and involvement of all genders in open source projects such as Common Voice.
- olejorgenb 5y agoI find the recording UI a bit annoying. They make it unnecessary hard to re-record a clip. Re-recording the previous clip is likely to be a common thing to do. Instead of providing a shortcut for this, they have shortcuts for re-recording each of the individual 5 clips.. It's also impossible (?) to undo a clip. Eg.: If I've already recorded 3 clips and mistakenly begin a clip I simply can't pronounce correctly, there's no way of removing that clip without discarding the whole set. (EDIT: it is possible by re-recording that clip and pressing skip)
- Vinnl 5y agoRe-recording a clip is very rare for me. Keep in mind that it's supposed to emulate real-world conditions, with all its messiness.
- SilverRed 5y agoYeah I think minor mess ups as long as the words are correct is actually good. As well as a bit of background noise. Problem is if Moz builds up a dataset of pure recordings and someone tries to use it but they are in a noisy room and the ML was never prepared for this.
- arghwhat 5y agoPeople seem to speak extremely mechanically in these samples, which I suspect may lead to training bias against native as speech if used. I think it should be explained that one should speak naturally when reading the lines.
- stegrot 5y agoMany people are also speaking very mechanically when they use a voice assistant, though ;) I believe we need a good mix, but telling people to speak a little more naturally certainly would help.