4 ms·
Does this dataset include people with voice or speech disorders (or other disabilities)? I don’t see any mention of it in this announcement or the forums, thoug
by jointpdf 6y ago
Does this dataset include people with voice or speech disorders (or other disabilities)? I don’t see any mention of it in this announcement or the forums, though I haven’t looked thoroughly (yet).
Examples: dysphonias of various kinds, dysarthria (e.g. from ALS / cerebral palsy), vocal fold atrophy, stuttering, people with laryngectomies / voice prosthesis, and many more.
Altogether, this represents millions of people for whom current speech recognition systems do not work well. This is an especially tragic situation, since people with disabilities depend more heavily on assistive technologies like ASR. Data/ML bias is rightfully a hot topic lately, so I feel that the voices of people w/ disabilities need to be amplified as well (npi).
- dabinat 6y agoIt’s not in the current dataset, but offering such a disordered speech dataset has been discussed. I imagine it’s something that will probably be offered at some point in future.
- sagz 6y agoThere's g.co/euphonia for those projects
- daanzu 6y agoGoogle is certainly doing some great work with this, both Project Euphonia and other research [0]. However, as far as I know, the Euphonia dataset is closed and only usable by Google. A Common Voice disordered speech dataset would (presumably) be open to all, allowing independent projects and research. (I would love to have access to such a dataset.) [0] https://ai.googleblog.com/2019/08/project-euphonias-personalized-speech.html https://ai.googleblog.com/2019/08/project-euphonias-personal...
- jointpdf 6y agoCan you clarify what you mean by “those” projects? Because this project is supposed to be about “everyone”: >”Common Voice is part of Mozilla’s initiative to make voice recognition technologies better and more accessible for everyone.” If people with speech/voice disorders are not represented in the dataset, then this is not inclusive of everyone.
- daanzu 6y agoGathering, collecting, and publishing such a dataset would be great, and would certainly much improve the baseline speech recognition for people with disordered speech, but it can only help so much without personalizing to a specific individual. This is true for anybody, but more so for disordered speech. This is an area where I think "generic" solutions will inevitably struggle, even if they are somewhat specialized on "generic disarthritic" speech. However, this means that the gains to be had from personalized training are greater for disordered speech than for "average" speech. I develop kaldi-active-grammar [0], which specializes the Kaldi speech recognition engine for real-time command & control with many complex grammars. I am also working on making it easier to train personalized speech models, and to fine tune generic models with training for an individual. I have posted basic numbers on some small experiments [1]. Such personalized training can be time consuming (depending on how far one wants to take it), but as my parent comment says, disabled people may need to rely more on ASR, which means they have that much more to gain by investing the time for training. Nevertheless, a Common Voice disordered speech dataset would be quite helpful, both for research, and for pre-training models that can still be personalized with further training. It is good to see (in my sibling comment) that it is being discussed. [0] https://github.com/daanzu/kaldi-active-grammar https://github.com/daanzu/kaldi-active-grammar [1] https://github.com/daanzu/kaldi-active-grammar/blob/master/docs/models.md#fine-tuning-for-individual-speakers https://github.com/daanzu/kaldi-active-grammar/blob/master/d...
- totetsu 6y agoI have heard a few people with speech disorders when validating clips. I also recall some discussion of it in the discord or issue tracker. At the moment it is entirely up to people to encourage people with voice or speech disorders to submit. So long as they meet the validation criteria they will be included. I can't see a flag in the user profile for recording a disorder, so it's not likely you can filter just these recordings from the data.
- jointpdf 6y agoAfter taking some time to dig into the project and forums, I’m more concerned. I have worked on building and validating large-scale ML datasets in tricky domains before and am best friends with an SLP, so I have some context for understanding the challenges involved with creating a dataset like this. I apologize if my tone is too harsh—my intent is to help hold the ML community to a higher standard on this issue and generate productive conversation via my criticisms. The Common Voice FAQs say the right words about the mission of the project: >”As voice technologies proliferate beyond niche applications, we believe they must serve all users equally.” >”We want the Common Voice dataset to reflect the audio quality a speech-to-text engine will hear in the wild, so we’re looking for variety. In addition to a diverse community of speakers, a dataset with varying audio quality will teach the speech-to-text engine to handle various real-world situations, from background talking to car noise. As long as your voice clip is intelligible, it should be good enough for the dataset.” However, your data validation criteria both implicitly and explicitly exclude entire classes of people from the dataset, and allow for the validators to impose an arbitrary standard of purity regarding what constitutes “correct” speech. In so doing, you are influencing who is and isn’t understood by systems built upon this data. Examples from the docs (https://discourse.mozilla.org/t/discussion-of-new-guidelines-for-recording-validation/36465/11 https://discourse.mozilla.org/t/discussion-of-new-guidelines...) >”You need to check very carefully that what has been recorded is exactly what has been written - reject if there are even minor errors.” As currently stated, this criteria leads to the categorical exclusion of people for whom speaking without “even minor errors” is not possible (ex: lalling and other phonological disorders, where certain phonemes can’t be formed), based on the validators’ subjective perception of data cleanliness. >”Most recordings are of people talking in their natural voice. You can accept the occasional non-standard recording that is shouted, whispered, or obviously delivered in a ‘dramatic’ voice. Please reject sung recordings and those using a computer-synthesized voice.” Please watch this example of a person you are defining out of your dataset: (https://m.youtube.com/watch?v=5HgD0PXq0E4 https://m.youtube.com/watch?v=5HgD0PXq0E4) Look at this kid’s face light up and tell me that’s not his new natural voice. An electrolarynx is not a computer-synthesized voice (you manipulate the muscles in your neck to generate vibrations—like an external set of vocal cords). Although it would almost definitely be mistaken for one, and summarily sent to the “clip graveyard” (https://voice.mozilla.org/en/about https://voice.mozilla.org/en/about). >”I tend to click ‘no’ and move on for extreme mispronounced words. I’m of the opinion that soon enough, another speaker from their nationality will submit a correct recording.” Again, the use of the word “correct” here is problematic. Rejecting borderline cases and waiting for “cleaner” samples is a severe trap to fall into, regardless of the domain. >”I do the same as you. Accept if it’s an elongation; reject if the reader takes two attempts to start the word.” Again, this almost categorically excludes people with a stutter and other types of speech disorders. @dabinat gets its right with this comment: >”There are uses for CV and DeepSpeech beyond someone directly dictating to their computer. In my opinion, CV’s voice archive should contain as many different ways to say something as possible.” But then... >”You may well be right. I’d be interested to hear what the programmers’ expectations are.” >”I will ping @kdavis and @josh_meyer for feedback on the ML expectations (in terms of what’s good/bad for deepspeech).” Yikes. So the data is being selected to improve performance benchmarks of the speech recognition model, and not to better reflect the nuances and variety of speech in the real world (as was the stated goal of Common Voice). It’s very easily the case that cherry picking data to improve test benchmarks will decrease generalizability of the model in other applications. Narrowing the range of human speech to make the problem easier (as in simpler to build a model that functions well for most people) is antithetical to your stated mission. We can’t keep measuring AI progress in parameters and petabytes. It has to be about the people it helps. >”I agree that we don’t want to scare off new contributors off by presenting the guidelines up-front as an off-putting wall of text that they have to read.” Limiting the amount of documentation/training available to data annotators in an effort not to scare them is a surefire way to end up with inconsistently labeled data. Although I find the above examples to be dismaying, I do not mean to ascribe any ill intent to your team or the volunteers. I understand the complexities at play here. But the outright dismissal of certain types of voices as out-of-scope or not “correct” is causing real harm to real people, because ASR systems simply do not work well for people with various disabilities. I could find no direct mention or acknowledgement of the existence of speech disorders anywhere* on the website or forum. I believe there needs to be a more deliberate effort to construct a more representative dataset in order to meet your stated mission (which I am willing to volunteer my time towards). Just some initial ideas: - Augment the dataset by folding in samples from external datasets (e.g. https://github.com/talhanai/speech-nlp-datasets https://github.com/talhanai/speech-nlp-datasets). I’m not sure on the approach, but if movie scripts can be adapted, presumably so can other voice datasets. - Retain samples with speech errors like mispronunciations and stutters (perhaps with a flag indicating the error). In fact, why not retain all samples, flagging those that are unintelligible? At least keep it available, for data provenance purposes (so it is known what was excluded and can be reversed). - Establish a relationship with speech-language pathologists to collect or validate samples (eg: universities or the VA, who have many complex/polytrauma voice patients). Sessions with SLPs often involve having patients read sentences aloud, so it’s a familiar task. This is probably the best way to collect data from people with voice disorders, so volunteer annotators aren’t responsible for analyzing a complex subset. - Use inter-annotator agreement measures to characterize uncertainty about sample accuracy, rather than binary accept/reject criteria. - Collect/solicit more samples from people >70yrs old, since they are currently underrepresented in your data. Is there anyone over the age of 80 in your dataset at all? - Improve your documentation and standards to be more explicitly transparent about the ways in which it does not currently represent everyone, and plans for bridging these gaps.