4 ms·
You could only display the speaker's location if you had some way to associate streams with the individuals. So you'd have to train an audio-visual association
by kajecounterhack 4y ago
You could only display the speaker's location if you had some way to associate streams with the individuals. So you'd have to train an audio-visual association model.
The thing you could do is train a localizer on the separated audio. Phase is estimated by the source separation process, so you can actually train an ML model provided you have some ground truth (e.g. estimated human locations from camera detections)