29 ms·
But they most likely do. Voice corpora are extremely useful to train voice recognition systems, for one.
by diego 9y ago
But they most likely do. Voice corpora are extremely useful to train voice recognition systems, for one.
- lovich 9y agoAnd people have a problem with it, and so the companies use Corp speak to try and make it seem like they aren't doing what you asked. The poster higher up asked how they could remove all doubt, and it's really easy. The problem is that they are trying to remove all doubt while continuing to do business as usual
- crucifiction 9y agoIts only useful to train if they also have the transcript by a human. Eavesdropping conversations and having armies of people transcribing them seems like a very expensive and illegal way to get that data when there are probably millions of available samples, TV shows, etc that have both voice and transcription available already off the shelf.
- artificial 9y agoThe 2017 F8 developer conference featured a lot on Machine Vision. Processing images and video for objects, such as for automatic close captioning, is where vast resources are focused. I highly recommend watching a few of the videos, they’re specially aware as well and can infer orientation of obscured things like limbs. Microsoft has real-time audio translation, doing machine transcription at scale is totally feasible.
- crucifiction 9y agoThe OP that I was replying to was insinuating that FB is collecting audio as training data to create AI models like the ones you are talking about. I was pointing out that raw audio is useless to train an AI model for recognizing words, the whole point of training data for AI is that you have an input and a known output (transcription) that you can use to train and test the model with, having just input is useless for training.
- gt_ 9y agoThat makes sense. So, given that this is data it's users have an obvious interest in keeping private, Facebook could at least inform of this. The only reason I could imagine they aren't doing this is that their voice recognition isn't used for any released products. I think that's a viable theory of what they would be doing with the voice data. EDIT: User 'crucifiction' makes a good point about needing the transcript to use it for voice recognition. So, who knows. It doesn't in any sense justify collecting data and being so elusive about their intent. We naturally have reason to get the impression that Facebook just wants the data, will try anything it can to keep collecting it until so much energy has been invested in the PR issues that the collection has an attributable effect on their market. But maybe they are using the voice data to change the world for the better and bring people together yada yada! :)