3 ms·
This is a very different kind of problem than the ones you listed. One drop of blood is going to be very similar to any other in and individual. That's not true
by Analog24 7y ago
This is a very different kind of problem than the ones you listed. One drop of blood is going to be very similar to any other in and individual. That's not true when it comes to language data (or many other types of data for that matter). The data you would record in a prepared setting (i.e. reading from some predefined set of phrases) is typically not even close to representing the full distribution of phrases/dialogues that human's use.
Furthermore, Google/Amazon/FB do use representative sampling of real user data, it's not feasible to transcribe every interaction with Google Home/Alexa/Siri. This akin to what you're suggesting but it no way addresses the privacy concerns. The only real way to do that is to use authorized data or scripted interactions, which, as described above, are not actually representative samples. It is complicated and nuanced problem.
- la_barba 7y ago>One drop of blood is going to be very similar to any other in and individual. That's not true when it comes to language data (or many other types of data for that matter). Why? Please do explain. If you claim that our biology doesn't change at all in one domain, but varies significantly in another, it would be easy to show this scientifically, or more specifically, how this variance is applicable in this context of voice recognition. Just to take a simple example of blood glucose. Using continuous glucose monitors attached at various sub-cutaneous sites over the body, it is trivial to show how the local glucose is not identical at all sites.