4 ms·
From my perspective, the only fool proof way of removing all audio watermarks from a conversation is to run individual speaker detection, STT detection, and voi
by mpoteat 6y ago
From my perspective, the only fool proof way of removing all audio watermarks from a conversation is to run individual speaker detection, STT detection, and voice cloning algorithms to "recreate" the conversation from scratch.
Even things like background electrical hum have been used in audio forensics.
- farnsworth 6y agoI've seen the demos of AI systems that can be trained on an individual's voice, and generate new speech in the same voice. If I was waterprinting a meeting, I would train a system like this on the fly on the people in the meeting and use it to dynamically insert filler words (um...) in a unique pattern into the audio stream for each member of the meeting. That would defeat any audio filter tricks or recreating the meeting from scratch, so you would want to summarize/rephrase the meeting too.
- bobthepanda 6y agoHow would you explain these filler words if the lips don’t match?
- loa_in_ 6y agoBut if certain frequencies are in the training data, they still will be in the output. Won't you end up with watermarked audio still?
- ardy42 6y ago> I've seen the demos of AI systems that can be trained on an individual's voice, and generate new speech in the same voice. If I was waterprinting a meeting, I would train a system like this on the fly on the people in the meeting and use it to dynamically insert filler words (um...) in a unique pattern into the audio stream for each member of the meeting. That would defeat any audio filter tricks or recreating the meeting from scratch, so you would want to summarize/rephrase the meeting too. But wouldn't doing something like that make the recording seem like a deepfake, potentially reducing it's credibility? To an outside observer, it may make a genuine recording appear to be fake. I'm reminded of the controversy over the faked Bush Texas ANG documents. They were discovered by randos on the internet realizing the text looked like the output of a modern version of Word than the 70s typewriter they should have been written on if genuine. Imagine a similar but genuine document where the content was deliberately retyped using anachronistic equipment to obscure the leaker's identity. [1] https://en.wikipedia.org/wiki/Killian_documents_controversy https://en.wikipedia.org/wiki/Killian_documents_controversy
- semi-extrinsic 6y agoI don't remember the details, but there was a lawsuit (in India IIRC) over some contractual documents, which were proven to be fake since they used Microsoft's Calibri font, but were supposed to be from before Calibri et al. was released.
- Haemm0r 6y agoyou could let your recreation algorithm do the same thing. add and remove filler words randomly. This way in the end you can't be sure whos audio it was :)
- xxpor 6y agoYeah that's a good point. Just because you've removed the watermark, doesn't mean you've removed all of the unique features.
- Triv888 6y ago> "recreate" the conversation from scratch Why not do speech to text at that point?
- rndgermandude 6y agoThat's a problem tho, because then people will claim that what you recorded is doctored or even "a deep fake".
- yitianjian 6y agoAgreed, but at this point even images and videos can be faked and doctored to a pretty high standard. Privacy of the leakers should be worth it IMO.
- eru 6y agoWell, why release 'the meeting' in the first place then? If you recreate a deep fake, you might as well start from a transcript. Or leak a transcript. Leakers need some fidelity to prove their credibility. But fidelity also identifies them.
- jessaustin 6y agoThe journalist, editor, or their lawyer might need something genuine to be comfortable publishing. If that journalist and editor have a good reputation, however, the general public shouldn't need that. "The Intercept" may not have that good reputation? They seem at least as trustworthy as the typical USA war media firm to me.
- eru 6y agoYes. But in that scenario, what's the value added compared to the journalist just publishing a transcript?
- m463 6y agoJust make parts of the session randomly unintelligible and use the absence or presence in the transcript as an identifier.
- Robin_Message 6y agoAlso need to quantize all of the pauses between speakers and the time of each new speaker starting, and the rate of speaking, since as others have pointed out, Zoom varies these anyway for delay compensation.