3 ms·
Almost everything digital in audio seems to have been a decade or two ahead of video (non-linear editing, lossless/uncompressed quality, digital synthesis of el
by fitzroy 7y ago
Almost everything digital in audio seems to have been a decade or two ahead of video (non-linear editing, lossless/uncompressed quality, digital synthesis of elements, distribution, etc), primarily due to the bandwidth and processing power requirements being much greater for video.
So I'm curious how, with regard to deepfake technology, video seems to be well ahead of audio. Is audio deep fake technology simply less interesting to people? Are listeners far more sensitive to voice not being perfect? Is the human input just a lot less helpful in modeling the output with voice vs video (where the initial human input is only slightly more helpful than just synthesizing from text-to-speech)?