3 ms·
I have no firsthand knowledge, but I’d strongly bet that the home-assistant effort to donate training data is mostly get adult males, and nearly zero children.
by vineyardmike 7mo ago
I have no firsthand knowledge, but I’d strongly bet that the home-assistant effort to donate training data is mostly get adult males, and nearly zero children.
- ethagnawl 7mo agoOh, I'm sure you're right. I've had people in my personal life (non-technical; "AI enthusiasts") laugh at me over concerns about training bias but this is likely a real world example of it.
- stavros 7mo agoI think you can train your own wake word with microWakeWord but I've never done it.
- dghlsakjg 7mo agoThis was 2021 (so pre-llm), but I used to work for a company that gathered data for training voice commands (Alexa, Toyota, Sonos, were some clients). Basically, we paid people to read digital assistant scripts at scale. Your assumptions about training data do not match the demographics of data I collected. The majority of what our work revolved around was getting diversity into the training data. We specifically recruited kids, older folks, women, people with accented/dialected English and just about every variety of speech that we could get our hands on. The companies we worked with were insanely methodical about ensuring that different people were included.
- gmueckl 7mo agoYou are reporting on a deliberately curated effort vs. what I understand is effectively voluntary data donation without incentives. It's not surprising to me that the later dataset ends up biased due to the differences in sourcing.
- dghlsakjg 7mo agoYour understanding of the datasets I helped create seems at odds with my experience actually creating the datasets. Do you have some insider experience or knowledge with dataset curation and creation for voice assistants that contradicts my own. The guideline is that the newer your model, the more likely it is to have diverse voice recognition datasets since it solves the earlier problems caused by non representative data. The trend is moving towards better recognition for outliers. The training models are fed data that is very specific and not at all just whatever recordings they have collected in an S3 bucket. Given the amount of post recording work diarization, and QA we had to do on every single recording, I can’t imagine wanting to YOLO in bulk data.
- vineyardmike 7mo agoYou're missing the point. No one cares about the datasets you've created in a commercial context. The effort being discussed is a volunteer effort among a community of tech enthusiasts, who are disproportionately privacy-oriented vs the average person. This will undoubtable skew towards middle-aged male audiences, and will be extra-selective against children. It's a best-effort collection, they're probably not turning anyone away, and it's only what they can get, they're (AFAIK) not paying anyone to collect underrepresented demographics.
- dghlsakjg 7mo agoAh. I thought you were talking about voice assistants in general. My mustake
- bluGill 7mo agoI remember when those systems first started collecting data they were worried kids wouldn't be handled - but they didn't know how to handle the privacy issuses with recording kids so discouraged it. Women being missed is not a surprise - but not anticipated.