3 ms·
> Google says the visual component here is key, as the tech watches for when a person's mouth is moving to better identify which voices to focus on at a given p
by _rpd 8y ago
> Google says the visual component here is key, as the tech watches for when a person's mouth is moving to better identify which voices to focus on at a given point and to create more accurate individual speech tracks for the length of a video