4 ms·
Single-mic source separation is possible in an unsupervised manner today that could probably work better than beamforming both compute-wise and with regard to i
by kajecounterhack 4y ago
Single-mic source separation is possible in an unsupervised manner today that could probably work better than beamforming both compute-wise and with regard to implementation difficulty (you'd just need a lot of recordings to represent the space of sounds you want to separate).
- morcheeba 4y agoMaybe a combination? Even simple beamforming/stereo would be helpful to help display the speaker's location. For example, the "speaker 1" tag could appear on the left, center, or right of the display to give a spatial clue where they are located.
- kajecounterhack 4y agoYou could only display the speaker's location if you had some way to associate streams with the individuals. So you'd have to train an audio-visual association model. The thing you could do is train a localizer on the separated audio. Phase is estimated by the source separation process, so you can actually train an ML model provided you have some ground truth (e.g. estimated human locations from camera detections)
- rapjr9 4y agoDo you have any references for this or a link to a commercial service? I'm currently in the process or trying to extract some background voices in a video (an interview where the faint background conversation is in English and the loud overdub is in Bulgarian). I tried Melodyne but it seems to only separate on pitch, not volume, and the pitchs are too similar (mono, three voices, all female) and words are made of lots of short phonemes which each create a "note" and makes editing impossible. I looked into Izotope RX as well and it does not seem capable of doing this either. There are services that can automatically add subtitles using speech-to-text translators but they are expensive, and I'd prefer to have the background voices rather than the Bulgarian interpretation of them translated back into English. It seems possible to do, in Peter Jackson's The Beatles:Get Back they were able to separate voices from loud foreground musical sounds and other ambient talking, but that technology was custom and doesn't seem to be publicly available yet.
- kajecounterhack 4y agohttps://arxiv.org/abs/2110.10739 https://arxiv.org/abs/2110.10739 I haven't seen it provided as a commercial service or free model yet, but there is open source code for Mixit that lets you train using the open source / canned FUSS dataset.
- rapjr9 4y agoThanks! That led me to this which looks like a good place to start: https://github.com/google-research/sound-separation/blob/master/models/neurips2020_mixit/README.md https://github.com/google-research/sound-separation/blob/mas...