3 ms·
This is an OK description though I think a more accurate one is that it learns to pick out things you are interested in -- it's learning to estimate masks in th
by kajecounterhack 3y ago
This is an OK description though I think a more accurate one is that it learns to pick out things you are interested in -- it's learning to estimate masks in the frequency domain that remove everything but your sounds of interest. There's no pile of stuff you're not interested in (though you can generate it by inverting the mask).
The masks are learned such that they try to preserve the phase / time difference between the two recordings, thus preserving the info you'd need to do TDOA (and thus directionality).
I wonder how much distance variation between the two mics their method can tolerate before degrading the recorded sound's bearing, since that changes the time difference. I guess if the goal is just to spatialize the audio (45-90 degree bearing), it's good enough that they could even do it on airpods, in theory.