8 ms·
Real-Time Noise Suppression Using Deep Learning
- asargsyan 8y agoCool would it be to use this technology!
- cibaocloud 8y agoAs most of us cannot afford a $500k DGX-2, or even a $10k V100 card, renting a DGX by the day makes sense, if you want to do deep learning training, ML, or graphics rendering. Check out this new service: https://www.reddit.com/r/gpumining/comments/9x1vc2/rent_dgx_supercomputer_by_the_day_now_in_boston/ https://www.reddit.com/r/gpumining/comments/9x1vc2/rent_dgx_...
- pslam 8y agoThe story title is "AI powered Noise Cancellation" but the text never uses the term "AI" at all. It's deep (machine) learning. It doesn't need the useless marketing bonus term "AI" to make it better — it's already interesting enough without.
- adammathias 8y agoAgree, the actual title "Real-Time Noise Suppression Using Deep Learning" is much more informative.
- sctb 8y agoWe've reverted the headline from the submitted “NVIDIA on state of art in AI powered Noise Cancellation” to that of the article.
- Aic1kuir 8y agohttps:// https:// > devblogs.nvidia.com uses an invalid security certificate. Certificates issued by GeoTrust, RapidSSL, Symantec, Thawte, and VeriSign are no longer considered safe because these certificate authorities failed to follow security practices in the past.
- iamrafael8 8y agoyes, that was unfortunate as I cannot see the post, but the title seemed interesting
- orliesaurus 8y agoHere you go (as an image on imgur) - https://i.imgur.com/7DGvmcI.jpg https://i.imgur.com/7DGvmcI.jpg
- msla 8y ago> Certificates issued by GeoTrust, RapidSSL, Symantec, Thawte, and VeriSign are no longer considered safe because these certificate authorities failed to follow security practices in the past. Those are some pretty big names. Names where reasonable companies could believe that nobody would ever dare enforce the rules against them, because it would break the Web. Who says nobody ever got fired for buying IBM? Heck. If the for-pay CAs keep screwing up, Let's Encrypt could become the sane, reasonable, conservative choice, even among the most Enterprise of Enterprise Enterprises.
- gsnedders 8y agoThey were all subsidiaries of Symantec when their various faults occurred leading to the Symantec distrust; it's all just the Symantec distrust.
- aartak 8y agoIt would be nice to have a Windows version. I have no Mac and can't run the Krisp
- ccostes 8y agoReally impressive results, though I wish they had gone more into the deep learning part of it (but I guess that's probably the secret sauce). Can't help but notice how well Nvidia is positioned for what appears to be a growing wave of demand for GPUs. Surprised this hasn't reflected in their share price (feels like they could be the next Intel, but what do it know).
- jononor 8y agoDedicated chips for machine learning (inference) are being developed by many companies. The hope is that these will be used instead of (or in addition to) GPUs for ML tasks. Not that Nvidia is poorly positioned. In fact, I expect that if dedicated ML chips work out, Nvidia will also put one on the market.
- twtw 8y ago> Nvidia will also put one on the market Already done. Tegra Xavier includes DLA (deep learning accelerator).
- konschubert 8y agoI am really impressed with what Nvidia is doing here. I think there is a huge market for improving sound quality in video calls. For me, roughly every second call I make is somehow harmed by some kind of "bad audio" problems. Breathing, reverb, noise, clipping, too silent, there are so many things that can go wrong. And this really harms the productivity of video calls. I have started collecting and building tools to detect all of these sources of bad audio and am collecting them at https://www.tinydrop.io https://www.tinydrop.io Maybe these APIs can help people to improve their setup. But if software like Nvidia's comes along and just fixes the problem once and for all - that's great as well!
- davitb 8y agoDisclosure: I'm the author of the blog post and co-founder at 2Hz. This is a guest post on NVIDIA Developer Blog. The author of the technology is a startup called 2Hz (2hz.ai). Our passion is to improve voice audio quality in audio/video calls. It's a tough problem but also fun to work on. Agree, breathing, reverb, noise are all problems and should be fixed. We started with noise and already shipped a product you can try on your Mac. The app is called Krisp (krisp.ai). Reverb, breathing, voice cutting will come next.
- hathawsh 8y agoHi! As someone who seems to struggle more than most to understand people on video calls, I'd like to give you my impressions. Something struck me about the sample video. The very first sample included background noise, but it was very easy to understand regardless of the noise, probably because it was recorded by a pro microphone rather than a phone. Every other sample was far more difficult, regardless of noise removal. Noise removal doesn't really seem to help; in fact, any imperfections in the noise removal process actually make the audio more difficult to understand because I have to guess not only the speaker's voice and the noise but also the algorithm for noise removal. What does help me is low frequency pickup. I think the first sample is easy because there are plenty of low frequency components that are later lost through the phone. Low frequencies are presumably difficult to pick up due to the size of the microphone in a phone, but could there be a way to restore those frequencies through audio processing? It would be interesting to analyze the response of specific microphones to specific low frequencies and find patterns that an audio processor could use to restore the low frequency components. Anyway, kudos for doing some very interesting work. I don't know how representative my experience is.
- acd 8y agoDoes this deep learning noise cancelling also work for music with headphones? If so then we can ditch proprietary noise cancelling headphones and just use the phones?
- davitb 8y agoDisclosure: I'm the author of the blog post. Not really. What works inside noise canceling headphones is a very different technology, called Active Noise Cancellation (ANC). You don't necessarily need Machine Learning to solve the problem. The technology described in the blog post is for suppressing the noise which goes from your surrounding environment to the other participants of the call (and vice versa).
- petra 8y agoActive Noise cancelling(ANC) headphones are really sensitive to latency. Take An ANC with 0 latency, that stops cancelling noise at 8 khz, and add a 50 usec latency to it, now will stop cancelling noise at ~1.5 khz. But This article talks about 20ms latency.
- Judgmentality 8y ago> Take An ANC with 0 latency, that stops cancelling noise at 8 khz, and add a 50 usec latency to it, now will stop cancelling noise at ~1.5 khz. How did you calculate this?
- petra 8y agoIt's an simple explanation of figure 3 here: https://www.edn.com/design/analog/4458544/2/A-perspective-on-digital-ANC-solutions-in-a-low-latency-dominated-world https://www.edn.com/design/analog/4458544/2/A-perspective-on...
- wrycoder 8y agoIf I understand that figure correctly, at 8KHz with no latency one gets 12 dB cancellation. With 50 ms latency, 0 dB. And above that frequency, the cancellation actually makes the noise worse. An analog cutoff filter would be needed.
- SeanFerree 8y agoAwesome article!
- adamloving 8y agoI downloaded the mac app, configured a virtual device to send the system output to the "Krisp Speaker" and verified that it cuts most of the music out of what I'm listening to, leaving only the voice (at a some what degraded quality). I wish I could configure it _cancel_ ambient noise, not just remove it from the input signal.
- ghostly_s 8y agoIn your perception, what is the difference between "cancelling" a signal and removing it?
- fredsanford 8y agoPhase Cancellation [1] [1] https://www.sageaudio.com/blog/pre-mastering-tips/phase-cancellation.php https://www.sageaudio.com/blog/pre-mastering-tips/phase-canc...
- Doxin 8y agoIt'll not be quick enough for phase cancellation, but presumably you could diff the output with the input, phase-invert it, and get the signal you want that way.
- mehrdadn 8y agoAn interesting human problem that I imagine would come up here is that the speaker could be getting distracted with all the noise (crying baby/siren/etc.) while the listener would have no idea what's going on and think the speaker is being confused/dumb/slow/etc... very curious how this would play out in real conversations!
- samstave 8y agoThis is really evident when a speaker hears an echo of themselves, slightly delayed, back through their speakers. It's really hard to speak when what you say comes echoing back.
- deleted 8y ago[deleted]
- npunt 8y agoLove it. Don’t really love the idea of audio contents of conversations being routed to a cloud server for processing though — needs to stay on-device for privacy.
- davitb 8y agoThis technology is already integrated into Krisp app (https://krisp.ai https://krisp.ai) and it runs all locally on device.
- npunt 8y agoThanks for the heads up! I really like what you're doing - not only is it great for the general public, it's a game changer for people with difficulties hearing.
- developer2 8y agoTheir ultimate goal must be to be acquired by Apple, Google, or similar. This will never fly as a third-party app/install, even if it's "promised" to be on-device (such promises can change). Moreover, the average user isn't going to know about, care about, or seek out such an app, let alone pay for it. There's no widespread reach or profit in selling direct to consumer. The only way this works is for it to be built into each device/OS (ie: firmware shipped by manufacturer). If the tech has merit, we'll be seeing it in a couple of years, whether that involves an acquisition here or independent R&D/patenting.
- npunt 8y agoAgree, I wasn't suggesting users would download a special app to use it - I'd like to see the tech make its way into the OS for all voice input. I can see why the cloud processing makes sense for certain applications / licensers / acquirers (e.g. a VoIP provider like Xoom), but voice comms is really the domain of smartphones, and my hunch is most are plenty powerful enough to do this processing locally.
- post_break 8y agoMy friend is an airline mechanic. One thing his coworkers all had was the jawbone headset. This was back in like 2008-2009. He said he could call up a mechanic working right next to the turbine while it was running and hear him crystal clear. I wonder if any of that technology paired with software technology will make it so there is 0 noise in calls. Maybe an implanted bone mic.
- johnvanommen 8y agoIIRC, that technology originated in fighter jets, and worked it's way down to consumer goods. The company sold Bluetooth headsets for a while, but multi-mic solutions were cheaper and worked better. They tried to hang in there for a few years, diversifying into consumer electronics like the "Jawbox." It didn't work out, and they went bankrupt.
- johnvanommen 8y agoWhat's wrong with using multiple microphones? A mic element costs about thirty cents, and the processing power required for noise cancellation already exists in the CPU of the mobile device. I think it's particularly interesting that Amazon has made microphe arrays particularly cheap, due to Alexa. MiniDSP offers a microphone array for under $100, which is an unheard of price considering what these cost ten years ago. https://www.minidsp.com/products/usb-audio-interface/uma-8-microphone-array https://www.minidsp.com/products/usb-audio-interface/uma-8-m...
- sophistication 8y agoHow does multi microphone filtering work? I guess they localize different sound sources by cross-correlation (to get the timings) and triangulation (based on the timings and the speed of sound)?
- gugagore 8y agoI think the (or perhaps only one) key phrase is "beamforming". A single microphone element has a certain sensitivity pattern (e.g. it may be a very directional microphone, or be equally sensitive in all directions). With multiple pick-ups, you can emulate some different sensitivity patterns. A related idea in radar is synthetic-aperture radar (SAR).
- johnvanommen 8y agoGreat point. A lot of the interesting things in audio were inspired by radar. Dan Wiggins at Sonos used to work on radar, and Don Keele created a loudspeaker technology called "CBT" that's based on radar technology. Because microphones are basically the inverse of loudspeakers, what works in loudspeaker arrays can also work in microphone arrays.
- b_tterc_p 8y agoYou can do this with ICA (Independent component analysis, a somewhat lesser known, non-Gaussian cousin of principal component analysis). Basically you take the data with multiple components and break it down into its consistent component parts.
- Aspos 8y agoCan imagine a codec, which would suppress noise, recognize the speech, send the text along with the voice information so in case of broken signal codec on the receiving end could reconstruct the speech using the text applying speaker's voice features via style transfer. So, if my voice is distorted in a broken line, it would be reconstructed from the text and reconstruction would sound like me. I guess it will be the ultimate 1kbps codec.
- deleted 8y ago[deleted]
- 33a 8y agoThat's crazy how well it handles nonstationary noise.
- exabrial 8y agoThey have mac app!! Incredible! How could I get this into my car kit?
- TatWakie 8y agoOne of my favorite teams! Looking forward to the time when I won't hear any background noise on my calls anymore :)
- opdahl 8y agoThis is amazing! Full props to the Nvidia team that accomplished this. I downloaded the Mac app they provided [1] which I highly suggest everyone with a mac tests out. I ran it on my old MacBook Air 2013 using daily.co. It worked like a charm. Definitely using this in the next group chat, where there is always someone who forgets to turn off their microphone. One cool side effect is that it actually removes the reverberations that happen when you have two computers on the same call, where the mics keep picking up on the output of the other computers and a large high-frequency noise happens (that I'm sure we all have experienced). The system simply removed it and I didn't even know it was there until I turned off the app. Amazing work and I really hope that Skype, Apple, Google etc implement this into their voice apps, or even phone providers build this into phones. Maybe in the future, we actually can have phone conversations in windy weather and on the streets. [1]: https://krisp.ai/?utm_source=Nvidia%20blog&utm_medium=download https://krisp.ai/?utm_source=Nvidia%20blog&utm_medium=downlo...
- loa-in-backup 8y agocopied from another comment, NOT MINE: COMMENT FOLLOWS davitb 18 hours ago Disclosure: I'm the author of the blog post and co-founder at 2Hz. This is a guest post on NVIDIA Developer Blog. The author of the technology is a startup called 2Hz (2hz.ai). Our passion is to improve voice audio quality in audio/video calls. It's a tough problem but also fun to work on. Agree, breathing, reverb, noise are all problems and should be fixed. We started with noise and already shipped a product you can try on your Mac. The app is called Krisp (krisp.ai). Reverb, breathing, voice cutting will come next.
- deleted 8y ago[deleted]
- Jude2711990 8y agoHave you tried SoliCall Pro (http://solicall.com/solicall-pro/ http://solicall.com/solicall-pro/)? Once installed using virtual audio device technology it will improve the audio with multiple options like NR, PNR, RNR, and more.