4 ms·
This approach surprised me. Why are they doing feature extraction and then feeding that into a DNN? It seems much more straightforward to have the input of the
by chas 10y ago
This approach surprised me. Why are they doing feature extraction and then feeding that into a DNN? It seems much more straightforward to have the input of the network be noisy samples and the output be clean samples a la super resolution[0] in images. They probably wouldn't want to use fully-connected layers in that instance, but I don't see any fundamental barriers if they have enough computational power to run a neural network already. Am I missing something?
[0] https://arxiv.org/pdf/1603.08155.pdf https://arxiv.org/pdf/1603.08155.pdf
- eb0la 10y agoThis is the paper they submitted: http://web.cse.ohio-state.edu/~dwang/papers/CWYWH.jasa16.pdf http://web.cse.ohio-state.edu/~dwang/papers/CWYWH.jasa16.pdf I'm not an expert in the hearing field, but makes sense for me: they probably know that some filters work (and are actually used) for hearing, and feeding that preprocessed data would save them training and hyperparameter tuning time.
- Eridrus 10y agoThe only thing I've seen run on raw audio are WaveNet models, and those are way too expensive to get into an embedded chip, no public real-time implementations exist, though Baidu had a paper which claimed real-time execution speed on some server class Intel CPU last week. They do mention that their CPU implementation could be parallelized too.
- biggieshellz 10y agoThe filter bands they're talking about are Bark bands (https://en.wikipedia.org/wiki/Bark_scale https://en.wikipedia.org/wiki/Bark_scale) and are actually representative of the way the ear perceives loudness. In a traditional hearing aid, you might have a compressor for each of these Bark bands to counteract the effects of loudness recruitment (http://www.sens.com/helps/helps_d03.htm http://www.sens.com/helps/helps_d03.htm).
- adinisom 10y agoThat might work, although I think there are two limitations: 1) Hearing aids have a 10ms latency budget. So no matter how much processing they can do, they're limited by how many samples they can look ahead and that limits the design of the filters. The brain can presumably look ahead further to separate sound streams so I think it's pretty impressive that ideal binary masking works. 2) Hearing aids have a power budget. The ones I've looked at achieve low power by running a FIR filter in hardware to shape the sound while a DSP classifies the sound and adjust the filter taps. The DSP doesn't have to run at the same rate as the filter. That seems well matched to the binary filter approach. Likewise features extraction might not run at the same rate as the DNN.
- gwern 10y agoThe latency and power issues can probably be fixed, assuming a good end-to-end model, by using model distillation into a wide shallow net using low-precision or even binary operations. I don't know if that would be enough - we've seen multiple order of magnitude decreases in compute requirements (think about style transfer going from hours on top-end Titan GPUs to realtime on mobile phones) but the usual target is mobile smartphones which at least have a GPU, while it seems unlikely any hearing aids will have GPUs anytime soon... I suppose a good enough squashed low-precision model could be turned into an ASIC.
- taliesinb 10y agoNot to detract from your larger point but AFAIK the style transfer thing is different. If you're willing to hardcode the style into the net you can go realtime, but the original style transfer paper is able to do different styles without retraining. So they're different algorithms. Unless the SOTA has changed recently.
- gwern 10y agoYou shouldn't need to hardcode the style if you provide the style as an additional datapoint for it to condition on. But this doesn't really matter since for fun mobile applications it's fine to pick from 20 or 50 pretrained styles, and likewise for hearing aids.