9 ms·
Common Voice
- hartator 3y agoBut, why?
- arcastroe 3y agoHow many people here have a different "reading voice" vs their normal conversational voice? Can conversational models be trained even if much of the training data sounds "scripted"?
- woodson 3y agoI remember when they (Mozilla’s CV team) solicited feedback before they got started, I brought up that issue and proposed a different approach to gathering conversational speech data, but it wasn’t picked up. The belief that it’s better to have more but crappy data rather than less data matched to what you actually want to solve is quite pervasive.
- nojvek 3y agoAmazing. One of my hopes with OpenAI were that they were going to be truly open. Open datasets, open code, open models, open evaluation. But it is now a Microsoft puppet running on corporate profit goals. This and HuggingFace are great to see. I hope HuggingFace isn’t acquired by Microsoft like GitHub did.
- rwmj 3y agoI wish they'd concentrate on the browser.
- creata 3y agoWhat work do you want them to do that they're not already doing?
- dingnuts 3y agovoice integration in a browser for control and feedback would be great if you were blind
- culi 3y agoAnd text-to-speech. Which is already a standard: https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_... The web is in a hilarious state where it's harder to style an option in a drop down than it is to generate speech from some text
- dragonwriter 3y agoBetter in the DE than an app, even the browser, unless its like ChromeOS and the browser is the DE.
- sowbug 3y ago[flagged]
- sonicanatidae 3y agoI mean, you're right, but so are they, just in the wrong place. :)
- sowbug 3y agoThat's the problem with this kind of logical fallacy. It distracts from the discussion, often by introducing a well-settled issue that's hard to disagree with. A company can do two things at once, and it should be possible to discuss one without discussing both. This is especially true when the alternative topic is (paraphrased) that Firefox could be a better browser. Do we really want to steal from a rare chance to talk about Common Voice and overwhelm it with a thread with everyone's pet peeve about their web browser?
- dang 3y agoOmit internet tropes https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- sowbug 3y agoRoger wilco. Thanks, dang.
- joomooru 3y agoAccessibility is an important part of the browser :)
- deleted 3y ago[deleted]
- OfSanguineFire 3y agoMycroft users really wished that Mozilla had kept up efforts in this direction, because otherwise the only option for reliable speech-to-text is uploading every command you give your agent to Google or Baidu. The browser is important, and I don’t support Mozilla’s vacuous projects for social-justice cred, but there are a handful of areas where we need some non-profit to provide a privacy-respecting solution.
- rwmj 3y agoThat is indeed important, so I take it back (can't edit original post now).
- user_7832 3y agoDidn't mozilla also have a related speech to text software that got canned/moved to a different company? Or was that different?
- salynchnew 3y agoDeepSpeech? https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech
- posguy 3y agoMozilla didn't want to fund further development, most of the team ended up at Coqui.ai
- user_7832 3y agoThanks, that's the one I was thinking about. I remembered it had an odd name.
- rasz 3y agoMozilla shut that project down same day (Apr 12, 2021) as: "Mozilla is partnering with NVIDIA, which is investing $1.5 million in Mozilla Common Voice,". Aka they got paid off by Nvidia to not compete.
- yorwba 3y agoDeepSpeech is not competition for NVIDIA, quite the contrary. More people using DeepSpeech means more GPUs sold. Seems more likely that Mozilla would have shut down both projects, but NVIDIA funding saved the more important one.
- contrarian1234 3y agoDoes speech to text require that much compute? EDIT: Nvm, there seems to be a new project called Sayboard that does everything on your phone: https://github.com/ElishaAz/Sayboard https://github.com/ElishaAz/Sayboard (though switching from Swiftkey is a bit annoying)
- sxp 3y agoFF's TTS is an important project for anyone who wants a trivial to use text-to-speech system. It's built into the browser so you can just run wss = window.speechSynthesis; for (let i = 0; i < wss.getVoices().length; ++i){ str = `Voice ${i} is ${wss.getVoices()[i].name}`; s = new SpeechSynthesisUtterance(str); s.voice = wss.getVoices()[i]; wss.speak(s); console.log(str); } in the console to get various TTS examples. For some browsers, this can be done offline while others use a cloud based TTS system.
- aragonite 3y agoNice! For debugging I've been having a lot of fun supplementing stderr (for especially important messages that I don't want to miss) with the free TTS voices available on Windows (by running a Powershell command) and from Chrome (via WebSocket). It's nice to have even more voices to choose from.
- alacode 3y agoPretty cool, it printed a list of over 8000 on my machine (Ubuntu, well Kubuntu now) and then proceeded to speak the voices after printing them all.
- j45 3y agoThis is handy to know, thanks. I was just trying out Common Voice a few days ago. They have a good example of a community page for folks wanting to help with a particular language. I was just thinking today that Firefox is worthy of switching back to because it was so fast,except I hadn't had a chance to do it. If anyone else thinks it's important for there to be an independent browser dedicated to privacy and security (and independence), they could as many casual browser switchers. I'm happy to be back on a few FF extension that didn't work quite the same on any chrome based browser.
- vlod 3y agoThis also works in Chrome (My version is: 119.0.6045.199) FF has 8611 voices, chrome has 19.
- 3y ago
- imjonse 3y agoWhile this dataset is orders of magnitude smaller than what recent speech models like Whisper and Seamless got trained on, and while it is meant for supervised as opposed to self-supervised learning where data is more abundant, it can still be useful for finetuning an existing model for improving its score on a specific language.
- skrebbel 3y agoI’m sad that this is English only. I’ll love to contribute lots of voice for a Dutch TTS from an nonprofit org like Mozilla
- meepmorp 3y agoThey do collect other languages - there’s a setting for it in the annotation section, and the dataset downloads let you choose other languages. e.g.: https://commonvoice.mozilla.org/nl/listen https://commonvoice.mozilla.org/nl/listen
- skrebbel 3y agoWoops! Thanks :-)
- meepmorp 3y agoDon’t feel bad - it’s not especially obvious. I only thought about it because I’m already familiar with the project.
- dabinat 3y agoAlthough English is the most-contributed language, one of the goals of Common Voice is to support languages that wouldn’t normally receive attention from commercial providers.
- yorwba 3y agoThe most-contributed language is Catalan with 3678 hours recorded vs. 3395 hours in English https://commonvoice.mozilla.org/en/languages https://commonvoice.mozilla.org/en/languages (The language list sorts your browser's UI languages ahead of all others, which is why English may appear on top for you.)
- zerotolerance 3y agohttps://commonvoice.mozilla.org/en/about?tab=how-add-language#playbook https://commonvoice.mozilla.org/en/about?tab=how-add-languag...
- dang 3y agoRelated. Others? Mozilla Common Voice Adds 16 New Languages and 4,600 New Hours of Speech - https://news.ycombinator.com/item?id=28073016 https://news.ycombinator.com/item?id=28073016 - Aug 2021 (170 comments) Firefox Voice - https://news.ycombinator.com/item?id=24096082 https://news.ycombinator.com/item?id=24096082 - Aug 2020 (154 comments) Firefox Voice: Browse the web with your voice - https://news.ycombinator.com/item?id=23902560 https://news.ycombinator.com/item?id=23902560 - July 2020 (2 comments) Mozilla Common Voice Dataset: More data, more languages - https://news.ycombinator.com/item?id=23695377 https://news.ycombinator.com/item?id=23695377 - June 2020 (41 comments) The Common Voice Project by Mozilla reached its first goal: 1k hours in englisch - https://news.ycombinator.com/item?id=23051756 https://news.ycombinator.com/item?id=23051756 - May 2020 (1 comment) Common Voice: A Massively-Multilingual Speech Corpus - https://news.ycombinator.com/item?id=21887693 https://news.ycombinator.com/item?id=21887693 - Dec 2019 (9 comments) Common Voice – Mozilla's initiative to help teach machines how real people speak - https://news.ycombinator.com/item?id=21268579 https://news.ycombinator.com/item?id=21268579 - Oct 2019 (49 comments) Mozilla releases the largest to-date public domain transcribed voice dataset - https://news.ycombinator.com/item?id=19270646 https://news.ycombinator.com/item?id=19270646 - Feb 2019 (61 comments) Mozilla Overhauls Speech-To-Text Contribution Interface - https://news.ycombinator.com/item?id=17436958 https://news.ycombinator.com/item?id=17436958 - July 2018 (42 comments) Initial Release of Mozilla’s Open Source Speech Recognition Model and Voice Data - https://news.ycombinator.com/item?id=15808124 https://news.ycombinator.com/item?id=15808124 - Nov 2017 (88 comments) Project Common Voice - https://news.ycombinator.com/item?id=14794654 https://news.ycombinator.com/item?id=14794654 - July 2017 (57 comments) Mozilla: Project Common Voice - https://news.ycombinator.com/item?id=14786881 https://news.ycombinator.com/item?id=14786881 - July 2017 (1 comment)
- vidarh 3y agoI submitted a request for Norwegian Bokmål, and realised a complication which I'm sure must affect other languages too: Norway has two separate official languages. They are unusually close - one is relatively close to Danish, and the other started as a collection of dialects, but technically they are written languages, especially Bokmål which basically means "book language". I'm unusual in that I speak close to "pure" Bokmål. Thanks to expectations at school etc., a lot of speakers who write Bokmål will adjust or tone down their dialect if asked to read a text that is written in grammatically and orthographically correct bokmål, but will otherwise speak in a manner that can deviate fairly significantly from the written language. As such, depending on whether your goal is text to speech or speech recognition, the pronunciation you will need is very different. E.g. people I know who write Bokmål might say something like "hva erredu ser på a?" ("what are you looking at?") with hardly any gaps between words, while I would stick close to the written "hva er det du ser på?" with clear gaps. In recognition you need to handle both (and many other variations), while for generation you'd at least by default usually want the latter unless there are indications the text is written in dialect. It strikes me you'd really want people to write more detail about what it is they are speaking and/or let people tag/label data with additional info about accents. Not just for this, but for other multi-lingual speakers as well. E.g. it'd be helpful to have many foreign accents in the English (and other languages) dataset for recognition, but as much as I want speech recognition to understand me, I'm not particularly interested in teaching it to speak English with a strong Norwegian accent. That is less of an issue than the dialects in some languages that can involve much more than just speaking the same words differently. To take another example "Jeg åpnet døren og gikk ut i solen" og "Jeg åpna døra og gikk ut i sola" are both valid Bokmål. Depending on context a reader may stick strictly to the text or swap åpnet<->åpna, døren<->døra, sola<->sola, and every permutation is valid. Which exact set you use differs and some speakers will write one but use the other when speaking. E.g. I would say åpna, døra, sola, but write åpnet, døren, solen. The latter is more formal and/or old-fashioned in some parts of the country, but the perception of that also varies by region. And this totally leaves out all the dialect variations used by people who'd say their language is Bokmål, and would be recognized as such by Norwegian speakers, but who use variants of words or conjugations that aren't technically recognized as valid Bokmål. The former is more "modern" (several of the forms are only valid Bokmål as a result of successive language reforms), more common in the Eastern part of Norway outside of the posher parts of Oslo and other wealthy regions, and (weirdly) more common in 1970's radical left-wing academics (especially people involved with the Maoist Workers Communist Party/AKP-ML) as an affectation/sociolect, with each of these groups also deviating in other aspects.... If you want to maximize the utility of a dataset like this, you really would want to let each speaker at least assign a lot of tags/labels to their profile; even if you don't want to deal with the hornet nest of trying to figure out all the distinctions, even unstructured labels would be a start, and ideally allowing people to tag individual recordings as well, because there are a lot more variations than just "language" and "accent" here.
- CoBE10 3y agoI'd like to give a shout-out to Common Voice Android: https://github.com/Sav22999/common-voice-android https://github.com/Sav22999/common-voice-android It's a handy app for those interested in contributing to the project. You can record voices for the languages you speak and validate other user contributions. I used to be a frequent contributor about two years ago, and this app had a much more user-friendly design compared to the official website version. Additionally, check out the official Common Voice Matrix channel: https://chat.mozilla.org/#/room/#common-voice:mozilla.org https://chat.mozilla.org/#/room/#common-voice:mozilla.org
- jeena 3y agoWhy then is the text2speech in reader mode (which other than that is excellent) on a Linux Firefox so extremely bad? Much worse than Steven Hawkins text2speech.
- spadufed 3y agoCrowdsourced datasets like this and the ones produced by the OpenAssistant project could easily become the ONLY way to build foundational models if the courts decide that what OpenAI and co are doing is not Fair-Use. I don't think I would call this scenario unlikely, either.
- pimlottc 3y agoWith recent events in AI and deepfake technology, I would need to see some assurances before I agreed to “donate my voice” to something like this. It seems like the project is for voice recognition, not generation, but it’s not immediately clear.
- charles_f 3y agoI don't know if assurances is the right term, but everything around machine learning and generation seems to be quite liberal with respecting people's property, so indeed something called "donate your voice" made me pause. Mozilla is probably the right organization for that. Their main product however is dwindling, and I'm not sure what will happen to their data if they ceased to exist. There is a tendency for dying organizations to be pulled apart for scraps, and this would definitely become an IP of interest for a lot of companies with much lesser noble causes
- yorwba 3y agoThe recordings are available for download, so if a company wants to use them for less noble causes, they can already do that.
- thih9 3y agoWhat assurances would you like to see?
- moron4hire 3y ago> Voice datasets also underrepresent: non-English speakers, people of colour, disabled people, women and LGBTQIA+ people. How does being gay change your voice?
- pseudalopex 3y agohttps://en.wikipedia.org/wiki/LGBT_linguistics#Accents_of_English https://en.wikipedia.org/wiki/LGBT_linguistics#Accents_of_En...
- moron4hire 3y agoI'm aware of the trope. I've yet to meet anyone that adheres to it, though. Always thought it was just one of those things that Hollywood overemphasizes to "other" gay people.
- jszymborski 3y agoGP linked you to a summary of a scientif finding, not a trope. You can read more about the study here [0]. Regardless, I think the point of collecting these stats is to make sure voices of people that might normally be under-represented in a uniform sample of a relatively small size can be corrected for. These are common stats to collect in surveys for similar reasons. [0] https://www.sciencedirect.com/science/article/abs/pii/S0095447005000379?via%3Dihub https://www.sciencedirect.com/science/article/abs/pii/S00954...
- pseudalopex 3y ago> Always thought it was just one of those things that Hollywood overemphasizes to "other" gay people. They may over emphasize it. I don't know. But now you know they didn't invent it.
- gary_0 3y agoThere was someone in my high school who had the stereotypical voice. One of his friends, who'd known him since they were little, mentioned that he had talked like that his whole life.
- williamsherron 3y ago[dead]