10 ms·
Project Common Voice
- jldugger 9y agoAnd... 503'd. I didn't catch what the intended use case was before it died, but I'm guessing computer generated voice? Most of the computer generated stuff I've seen uses trained actors. Which neatly avoids the problem of trying to reconcile a myriad of accents and dialects, which was immediately apparent from the first two samples I tried. edit: back up, seems to be about voice recognition, which this could help with no problem.
- popinman322 9y agoActually, based on the site content I think they're using it to create an archive of speech data to train speech recognition systems.
- eriknstr 9y agoSite works fine from my location. I think you are correct. > Read a sentence to help our machine learn how real people speak. Check its work to help it improve. It’s that simple.
- deleted 9y ago[deleted]
- mbebenita 9y agoThat is correct. The DeepSpeech project, https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech will use this data to train and validate open source / freely available speech to text models. The training data, along with the trained models will be made available for free to all users and researchers alike.
- punchingwater 9y agoSorry about the 503s! We were adding servers to our cluster to handle the hacker news load, and a few 503s are hard to avoid. If this is consistently happening for you, please file a bug and we'll look at it. https://github.com/mozilla/voice-web/issues https://github.com/mozilla/voice-web/issues
- therealunreal 9y agoAny plans for languages other than English?
- a3_nm 9y agoThis seems planned: https://github.com/mozilla/voice-web/issues/213 https://github.com/mozilla/voice-web/issues/213
- tangue 9y agoCool project, really aligned with the mission of Mozilla, and with a pleasant UX. And if you're a non-english speaker like me validating sentences is a nice way of improving your comprehension.
- TomasSedovic 9y agoYep! Though us non-native speakers should really be recording, too. So we're not left behind in voice recognition.
- punchingwater 9y agoExactly! Part of the goals of Common Voice is to make voice recognition work better for non-north american men (which is where the vast majority of the training data comes from). If you are a non-native speaker, we need your voice!
- tangue 9y agoOkay but what should strangers submit for country and accent in the form ? The only options are anglophones countries.
- punchingwater 9y agoGood question. Sounds like we should add an "Other" to that drop down, and make it clear that we are looking for all accents?
- tangue 9y agoYes - imho "non native english speaker" would be we the field where I would look first- edit typed on phone ...
- halomru 9y agoThe accent you think you are speaking (or are emulating) https://github.com/mozilla/voice-web/issues/242 https://github.com/mozilla/voice-web/issues/242
- pebers 9y agoIs the data going to be freely available as well? It's a little unclear whether they intend to make it separately available or not.
- hardwaresofton 9y agoIt looks like the database will be open sourced later this year: https://voice.mozilla.org/faq https://voice.mozilla.org/faq I'm wondering if the format will be easily translatable to the kinds of models that software like CMUSphinx and Julius use https://cmusphinx.github.io/ https://cmusphinx.github.io/ http://julius.osdn.jp/en_index.php http://julius.osdn.jp/en_index.php
- lucb1e 9y agoIt's weird to me that they publish the project about having an open dataset of voice data, with only a promise the open it up later.
- zie 9y agoIt's Mozilla, they likely just don't have the infrastructure/code in place yet. They definitely get my trust for believing they actually will. They have been very good in the past about keeping things open. They make mistakes sometimes, but their goals are all about being open.
- punchingwater 9y agothanks for the vote of confidence! yes we will absolutely open this data up, and it's just a matter of collecting enough data to be useful, and then building the UI. we have a goal of achieving this by the end of 2017, so stay tuned!
- blackkettle 9y agoI would assume they are only going to be doing the raw data collection and maybe cleanup and annotation, and that data should be made available so you can train what you like. If you poke around github and the Kaldi lists a bit more you can see that they are experimenting with and probably planning to use Kaldi. I wonder what they plan to do for provisioning. It is one thing to collect data and train models, but quite another to make the service available over the web in an unlimited capacity. And we are not yet to the point where you can reasonably expect to run a high quality open-vocabulary STT system in your browser. The search network is typically in the GBs range.
- apeddle 9y agoThis looks great! I use voice control to program on occasion due to an rsi injury. The standard stack for this is a mess due to closed source systems that aren't designed for voice programmers. A good open solution could really save me from a lot of headaches.
- qznc 9y agoPreviously, there was VoxForge [0], but it seems dead. At least, I failed to contribute my voice there. Mozilla getting into this space is good news indeed. [0] http://www.voxforge.org/home/read http://www.voxforge.org/home/read
- oulipo 9y agoYou can take a look at what we build at https://snips.ai https://snips.ai, we will open-source the platform later this year
- Jayakumark 9y agoCool. Just curious on What is the voice engine behind snips ? and who is the provider of training data ?. Also do you have plans for supporting additional languages or can it be trained on when you open source it ?
- kgdinesh 9y agoCan't believe it's down already.
- glandium 9y agoSadly, in Demographic Data, only native english accents can be selected.
- a3_nm 9y agoI have reported this and it looks like they intend to fix this https://github.com/mozilla/voice-web/issues/242 https://github.com/mozilla/voice-web/issues/242
- breakingcups 9y agoIf I read the issue right they don't intend to fix it at all.
- ibotty 9y agoAny idea why the duplicate detection did not work for this link: https://news.ycombinator.com/item?id=14786881 https://news.ycombinator.com/item?id=14786881 Anyhow: these should be merged (even though there is no discussion on the other submission)
- pvinis 9y agoI think that's why. When the previous thread is not very active, dup's are allowed. Not sure though.
- tomhoward 9y agoYes. The system is designed to allow multiple chances for good content to get exposure. https://hn.algolia.com/?query=dang%20deliberately%20porous&sort=byPopularity&prefix&page=0&dateRange=all&type=comment https://hn.algolia.com/?query=dang%20deliberately%20porous&s...
- ZoomZoomZoom 9y agoI hope this data will be used purely for voice recognition purposes and not for voice generation, or we'll be stuck with robots talking with this horrible gurgling and clicking accent due to poor recording conditions of most participants!
- sirlantis 9y agoAccording to their FAQ they actually want those poor conditions to be present in the corpus. > We want the audio quality to reflect the audio quality a speech-to-text engine will see in the wild. Thus, we want variety. This teaches the speech-to-text engine to handle various situations—background talking, car noise, fan noise—without errors. https://voice.mozilla.org/faq https://voice.mozilla.org/faq
- cooper12 9y agoIf they're planning to make a voice recognition system, why are they using example statements that are clearly taken from novels? [0] That's not how real people talk. They use a lot more slang, a lot more stopping and starting, filler words, etc. Instead you have people saying things like "irresolute", "rumbling", and other complex words. It would be useful for training a novel dictation system, but it's not how people would speak to their browser for example. [0]: An example sentence is "a thin circle of bright metal showed between the top and the bottom of the body of the cylinder", which is from H. G. Wells' War of the Worlds.
- jpalomaki 9y agoMaybe there's not yet good open datasets available for this kind of material? This gives Amazon, Apple and Google a nice advantage since they are able to collect huge sample sets of actual voice commands used by people and to some extent also correlate them with the actual action taken by the person. How could we collect such dataset? It's a bit chicken-egg problem. I don't want to talk to some open source system unless it has fairly good chance of understanding me. Should we try to half manually (through crowd sourcing) come up with potential requests like "Check news from CNN.com", "Order me quattro stagioni" which could be then fed to platform like Common Voice? Or should we work on higher level. Come up with task descriptions ("You want to order taxi to get to airport for your morning flight at 7am") and then let people record how they would actually request this from computer with voice. This might more accurately capture the language we actually use when speaking. Through some simple automation you could generate variations of the requests and at least partly the same base material could be used for different languages (task given in English, ask person to make the request in Finnish).
- saurik 9y agoIf you want people carefully reading books, it is pretty easy to get a hold of that kind of data in the form of audio books and the work of Recording for the Blind and Dyslexic. Sure, it isn't chunked into sentences, but since you have all of the source text you could do a quite reasonable job automating the slicing, throw out places you aren't sure, and still have a near infinite amount of great data. (Note that it isn't like these sentences are perfect anyway, hence the filtering process with volunteers: while I was judging some audio files one of the issues was "person turned off microphone a little too soon".)
- albertzeyer 9y agoThe terminology is a bit confusing. They are saying that they want to build voice recognition but it seems like they actually might want to build a speech recognition engine. Speech recognition is about recognizing the speech, the spoken words. Voice recognition is about recognizing the speakers voice, i.e. identifying the speaker. Also, maybe they also want to build a text-to-speech (TTS) system but I'm not sure. No matter what, the collected data might be useful for all of that, maybe except of voice recognition actually, because I guess the data will be collected anonymously? Note that there are some other existing big open speech corpora such as LibriSpeech (http://www.openslr.org/12/ http://www.openslr.org/12/) which could already be used right now to build a quite good speech recognition system.
- amelius 9y ago> Speech recognition is about recognizing the speech, the spoken words. Voice recognition is about recognizing the speakers voice, i.e. identifying the speaker. Perhaps they want to do both eventually (?) That could explain the name.
- punchingwater 9y agoCommon Voice is only about collecting a large public database of voices. We do have a separate project around speech-to-text [1]. We haven't done much work around speaker recognition (AFAIK) or voice synthesis, but they are both very interesting both from a technical and privacy related standpoint. That said, both are out of the scope of Common Voice (which is only about the data). 1.) https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech
- john2x 9y agoCould they use the different voices to generate unique, natural-sounding voices for text-to-speech?
- ancarda 9y agoI really hope so. All the text-to-speech software I've used has a generic sounding accent for the country (your choices are typically American, Australian, British, Canadian) but there's a lot more accents out there. The software isn't bad - it sounds realistic - but I wish it sounded more how I would like it to. There's some software, e.g. Cepstral Dallas - https://www.cepstral.com/en/demos https://www.cepstral.com/en/demos but it sounds too robotic to actually use and that voice isn't available for Linux so I only have it installed on my MacBook. I guess a lot of developers at e.g. Apple live in CA so Siri is probably influenced by that.
- giancarlostoro 9y agoI wonder if implementing a new type of Recaptcha with these type of projects in mind would make sense. The data wouldn't be going to some data center in Google land, but instead to some open end that anyone should be able to get their hands on. Also a free and open source recaptcha alternative would be nice. Trick is keeping it complex enough that bots cannot just reuse the existing public data set. Maybe withhold on making some of the data public for a few years till deemed 'retired'.
- timwaagh 9y agothis is an important development. voice control has good potential. would be cool if they used it as an alternative way to control firefox and/or servo?
- mememaestro 9y agoIt's a "Mozilla tries to data mine my voice" episode
- olegkikin 9y agoMan, most people have horrible microphones.
- mbebenita 9y agoIndeed, and a big problem for STT models, which is why we're trying to collect this type of data.
- eatbitseveryday 9y agoIt would be useful to collect data from non-native speakers of a language. More and more such individuals are appearing in all countries, and devices that accept spoken words should not break because of someone's level of command of a spoken language. For example, a Swiss speaking German (Hochdeutsch), or more clearly, a Brit speaking French, etc. Some children who grow up in multi-lingual families also intermix words from multiple languages into their sentences. We can still understand them.
- punchingwater 9y agoThis is a bug with our website [1]. We actually are trying to collect non-native speakers (as well as native). We are looking into clarifying this on the site. 1.) https://github.com/mozilla/voice-web/issues/242 https://github.com/mozilla/voice-web/issues/242
- sexydefinesher 9y agoShould i have any privacy concerns with contributing? I dont want just anyone to have the data to recreate my voice digitally.