3 ms·
> But does the state of the art language recognition allow for broad on-device translation beyond a subset of trigger words, i.e. "Alexa", "Siri", "OK Google"?
by CodeWriter23 9y ago
> But does the state of the art language recognition allow for broad on-device translation beyond a subset of trigger words, i.e. "Alexa", "Siri", "OK Google"?
You're talking about two different processes, recognition and comprehension.
Recognition on device is quite feasible and could be expanded from the trigger phrase to several dozen or perhaps a couple hundred keywords without burning up the device. The iPhone 4 did this on-device, pre-Siri, for voice dialing. Many cheapo car stereos do this too. And it is what the intelligent speakers and phones already do without cloud interaction, scanning for that trigger word.
Comprehension on the other hand, deciding what exactly the speaker is asking for, requires compute-intensive NLP and AI processing, and yes, that's going to the cloud.
Back to the specious argument by the article's author, to create an audio ad targeter doesn't require constant streaming to the cloud. For starters, you don't need to send silence or background noise. You don't need to send every word spoken either. You can screen for keywords in the trigger list, then send what is stored on-device locally before and after the trigger and forward just that.
Again, I don't think Facebook is listening to target you. Maybe if some CIA front corporation buys them up... I do however think the author's argument, the very first argument he made (and thus my labeling it as the foundation), is a load of malarkey.