4 ms·
Just out of curiosity, what's the frequency range that a typical mic can pick up the signal from? The article did not specially mention about the range instead
by devy 8y ago
Just out of curiosity, what's the frequency range that a typical mic can pick up the signal from? The article did not specially mention about the range instead it said inaudible.
And here is another article I found that mentions the normal 20-20kHz frequency response range: http://blog.shure.com/mic-basics-frequency-response/ http://blog.shure.com/mic-basics-frequency-response/
Isn't that mostly overlap with the human ear capability? I understand each person is different, etc. But just curious the specifics.
- hunter2_ 8y agoEvery model will be different, but the important thing is that the boundaries represent the frequencies within which the signal will stay above a particular threshold of amplitude. A good spec sheet will tell you that threshold, and I've seen things like -3, -6, or -10 dB. It will still pass audio outside of the range but at an undisclosed attenuation.
- jerf 8y agoYes, typical mics tend to pick up the typical human frequency range, though cheaper mics may have some really poor characteristics at the edges. Usually in the speech range they'll be pretty solid. However, there's a lot of play within the space. One difference is that microphones do a very direct recording of the sound waves, but what we hear is actually very distorted compared to the "real" sound by the nature of our ear. One of the big differences is that if there is a very loud 4000Hz sound, we can't hear a soft 4005Hz sound near it very well, but the microphone "hears" it just fine. So for instance, you could put out a loud sound for a user, but embed a very quiet command in frequencies the human couldn't hear, but if the listening model doesn't account for that (and there are reasons it wouldn't necessarily want to, because it wants to hear commands even in the presence of significant background noise), you could get commands in to a system. See https://en.wikipedia.org/wiki/Psychoacoustics https://en.wikipedia.org/wiki/Psychoacoustics for discussion about how our ears fail to pick up the "real audio" signal, and how much we've exploited that in music compression. Now, that was a very brute force example. It sounds to me like what this article is talking about are called "adversarial examples" (https://blog.acolyer.org/2017/02/28/when-dnns-go-wrong-adversarial-examples-and-what-we-can-learn-from-them/ https://blog.acolyer.org/2017/02/28/when-dnns-go-wrong-adver... ). Voice recognition doesn't listen the same way we do, it doesn't necessarily take a holistic view of the signal, but is looking for specific frequency patterns and changes and turning that into phonemes, into words, etc. (There's a lot of ways of doing this and I don't specifically know what Alexa and Siri are doing, so that's a really vague overview.) If you know what they are looking for, you can use filters to very, very selectively remove the patterns from a bit of music or something that Alexa might trigger on, and then insert just the bare minimum skeleton of the sounds that it is really recognizing. A human won't be able to hear the difference (most likely; depends on how badly the original is mangled but even if it is audible it is almost certainly not audible without an A/B test and very good ears), but the probably-neural-nets monitoring for sounds will end up superstimulated and interpret the adversarial example as words. While the adversarial examples work best with tuning to the target network, widely-shared networks like Alexa or Siri mean that such tuning is practical where attacking some custom-trained model used by one person isn't, and experiments have shown that adversarial examples travel between separately-trained nets and even non-neural-net models to a much, much greater degree than what at least my own intuition would have suggested before hand. (See previous link and look for the discussion of "Practical black-box attacks against deep learning systems using adversarial examples". It is extremely counter-intuitive to me how easy this is.)
- floatrock 8y agohmm... the big idea that MP3 figured out was you can document all these "if there is a very loud 4000Hz sound, we can't hear a soft 4005Hz sound near it very well" psychoacoustic phenomena and just throw away all that extra "can't hear it very well" information, resulting in a vastly-smaller filesize that still sounds reasonable (yeah yeah it's not FLAC and the purist needs their gold-plated Monster cables, lets not go there, that's not the point) So this attack is kinda a "reverse-MP3" that adds those lossy bits back in, but shaped with an attack payload. Or at least it adds enough pieces of the attack payload that the neural net pattern recognition triggers, while the humans say "Doesn't sound like anything to me". Is that a close-enough explain-like-im-a-freshman?
- floatrock 8y agoHey Berkeley researchers, if you're reading this and want to make a demo that will really freak people out, embed an alexa activation command into this clip: https://youtu.be/iyXtGo418TY?t=1m11s https://youtu.be/iyXtGo418TY?t=1m11s
- jerf 8y agoI primarily brought up psychoacoustics as an example of the way we don't hear the way microphones do. While you could abuse them, it would be more obvious. In this case what we're getting is the audio equivalent of adversarial examples; see the link I gave for some visual examples. What's interesting there is that they are basically invisible to us, but surprisingly robust. (As another sort of philosophical sidebar, this either proves, or provides very strong evidence, that whatever it is our brains are doing, it is not what deep learning nets are doing, nor anything else vulnerable to such trivial adversarial examples. I've seen adversarial examples against another technique that do seem to work against humans as well, but it requires such a distortion to the image that "I can't tell if that's a dog or a toaster" actually makes sense; it's not just some sort of attack against human vision or something, it's a fancy morphed thing halfway between the two that would probably confuse anything and anybody.)
- floatrock 8y ago