4 ms·
Well, a former PM who thinks voice analysis is not possible on-device and puts out the red herring of bandwidth and storage capacity as the foundational argumen
by CodeWriter23 9y ago
Well, a former PM who thinks voice analysis is not possible on-device and puts out the red herring of bandwidth and storage capacity as the foundational argument of why he “knows” they’re not listening.
PS I don’t think “they” are listening and use other highly effective means to target ads. But I do think arguing such targeting is only possible in the cloud, and only after recording all the sound in your environment: the air conditioner, traffic, kids playing, etc. and let’s not forget ALL the silence in between, is intellectually dishonest. Especially when we all know that any always-listening device processes the signal locally, scanning for trigger words.
- danso 9y ago"foundational argument"? He spends half the article talking about the other issues, starting with "But what if those technical realities disappeared?" I think he smudges the issue by conflating FB's theoretical ingestion of such audio with its current data storage capacity and ingestion. Presumably, raw audio could be transmitted and processed without being stored. But does the state of the art language recognition allow for broad on-device translation beyond a subset of trigger words, i.e. "Alexa", "Siri", "OK Google"?
- snowwrestler 9y ago> does the state of the art language recognition allow for broad on-device translation beyond a subset of trigger words, i.e. "Alexa", "Siri", "OK Google"? No. You think Apple, Amazon, and Google ship all those audio packets back to servers for fun? To believe that Facebook is processing ambient audio on a mobile phone, you have to believe that FB is doing things on a phone that even the phone's creators (who have no sandboxes and can access custom chips) cannot.
- CodeWriter23 9y ago> But does the state of the art language recognition allow for broad on-device translation beyond a subset of trigger words, i.e. "Alexa", "Siri", "OK Google"? You're talking about two different processes, recognition and comprehension. Recognition on device is quite feasible and could be expanded from the trigger phrase to several dozen or perhaps a couple hundred keywords without burning up the device. The iPhone 4 did this on-device, pre-Siri, for voice dialing. Many cheapo car stereos do this too. And it is what the intelligent speakers and phones already do without cloud interaction, scanning for that trigger word. Comprehension on the other hand, deciding what exactly the speaker is asking for, requires compute-intensive NLP and AI processing, and yes, that's going to the cloud. Back to the specious argument by the article's author, to create an audio ad targeter doesn't require constant streaming to the cloud. For starters, you don't need to send silence or background noise. You don't need to send every word spoken either. You can screen for keywords in the trigger list, then send what is stored on-device locally before and after the trigger and forward just that. Again, I don't think Facebook is listening to target you. Maybe if some CIA front corporation buys them up... I do however think the author's argument, the very first argument he made (and thus my labeling it as the foundation), is a load of malarkey.