This is the first startup here where I think the tech should essentially be illegal.
It's cool tech, yes I'm impressed at the achievement. Nuclear weapons are impressive too.
OTOH this kind of thing is getting easier and easier to do, so what's a realistic way forward?
Our only hope is that politicians and celebrities get sick of their voices and likeness being used to scam people or sell crypto or viagra and get laws passed against this type of impersonation.
What if the voice sample is somebody saying they give specific consent to be cloned by that service?
You could of course clone a voice to generate that "consent" -- but at that point there's no additional harm done because they'd already have the clone.
It's unrealistic that this tech won't exist somewhere, even if the big actors stay away for ethical reasons. A voice auth practice strikes me as a good compromise.
I guess eventually people will go back to only meeting face to face for important communications. I don't know what the way forward is for news.
I truly do not understand people like these founders, obviously they understand the future they're creating. "If not us, someone else would do it" is not an excuse. Neither is "I like money".
As this is what every "organized crime" (people who dont want prying eyes) groups have done for centuries.
Next, normal people will adopt the Mafia's 'cover the mouth while pretending to use a tooth-pick whilst talking to prevent lip reading from remote viewers (same thing sports people do currently.
-
My granmother was deaf for the latter half of her life. She became an expert lip reader.
It was fun going to restaurants with her as she would tell me what people at the tables far away were talking about "oh that couple isnt having a happy time..."
There are some cool uses like dubbing movies in foreign languages while keeping the original "voice styles" or having your long dead relatives talking to you in some memorabilia etc. It could also cause unexpected creativity explosion e.g. in games or fan fiction movies. To avoid misuses we might perhaps find the only good use of blockchain.
>long dead relatives talking to you in some memorabilia etc.
It seems a bit weird to me though. I mean, looking back at old records can still pass as mere nostalgic behavior. Wanting new sentences pronounced in disguise of lost relative voices is not great in term of respect for these people to share my own feelings.
Also I guess that now there is not much preventing completely new songs with whatever lyrics staring voices of Elvis, Hendrix and Pavarotti. Actually a continuous flow of on the fly generated lyrics seems perfectly plausible at this level, isn't it?
But what is a realistic way forward? Do you think that scammers won't have this technology in 2 years? Can we really prevent any illegal use of neural networks at this point? With weapons that you actually have to physically buy, you can intervene on a country level (to some degree). But already with those 3D printed ones, we are basically doomed. Of course it's a tragedy of the commons type of situation. But banning all legal uses does not prevent the illegal ones.
Most scammers are incredibly lazy and honestly not all that competent. There's no need for them to change that if you can prey on the weak and vulnerable.
The difference between "the paper is out there" and "there's a button to do this" is quite obvious in cases like software exploits. A report of finding a vulnerability rarely leads to a massive automated exploitation campaign, but if that report also contains a proof of concept the amount of automated attacks radically increase. I believe the same is true for many other types of crime: even a mild bar to entry will prevent a significant amount of criminals from advancing their techniques.
I think the negative impact of these voice changers is much bigger than the advantage we gain as a society. Criminals will always exist, even crafty ones, but "we can't prevent crime so let's not bother trying to do anything about it" is not a great take in my opinion.
Of course it's a tragedy of the commons type of situation
TOTC is about resource depletion. GYI. It's not applicable here.
deleted 4y ago
[deleted]
>Introducing the National Postal Service - send a letter to anyone for a nominal fee. No need for a personal courier, armed escort, or patrician status.
>This is the kind of thing that should be illegal. Now, any Plebian could essentially write a letter to anyone, impersonating anyone. Forged letters could drag us into a war with Persia - for Jupiter's sake!
This is a hilariously bad attempt at discrediting the original argument. There's a vast difference between forging a letter and replicating the unique vocal fingerprint of any human being, on demand.
I suppose if we approach the point that we can create robotic clones of anyone, anywhere, that look, sound, and move like anyone on the planet, that will be just like the post office too, right?
Yes, good point, mail fraud used to be a major problem and we started passing laws to deal with it 150 years ago.
https://www.uspis.gov/history-spotlight/history-of-the-mail-fraud-statue https://www.uspis.gov/history-spotlight/history-of-the-mail-...
Maybe we'll need a new specialized law enforcement agency like the Postal Inspectors to deal with the inevitable wave of AI-assisted crime.
Making then the illegal would accomplish nothing since it's already out in the wild. You can generate audio with high quality on fine-tuned versions of Tortoise TTS, which was originally trained on a cluster of NVIDIA 3090's, so it's within reach for any smart person to train a from-scratch model on consumer hardware. Realistically? We have to accept that this tech exists and there will be both positive and negative outcomes from it.
You may get your wish. The FTC posted an article about this a week ago. [1]
> The FTC Act’s prohibition on deceptive or unfair conduct can apply if you make, sell, or use a tool that is effectively designed to deceive – even if that’s not its intended or sole purpose.
It seems like an awfully broad rule? But they probably could go after this startup if they noticed it.
There are some kinds of businesses where making sure the regulators like what you’re doing is pretty much a prerequisite. On the other hand, plenty of companies got where they are today by pushing the limits.
[1] https://www.ftc.gov/business-guidance/blog/2023/03/chatbots-deepfakes-voice-clones-ai-deception-sale https://www.ftc.gov/business-guidance/blog/2023/03/chatbots-...
Wow, this is a great article. Obviously writing is easier than enforcing, but I'm pretty impressed with whoever at the FTC is already thinking so clearly about this stuff.
Thanks for sharing this ^
(apologies for going off topic here)
Wow. I would have imagined an article from the FTC to be more... Bland, for want of a better term.
The FTC consistently has one of the absolute best author voices in all of government. Pick a blog post at random and see what I mean. Their index on tech is probably the area you have the most domain knowledge in and so it’s probably the best area to evaluate them: https://www.ftc.gov/business-guidance/blog/term/1428 https://www.ftc.gov/business-guidance/blog/term/1428
Clear, direct, confident, not overloaded with qualifiers, not afraid of metaphor, self-summarizing, signposting, and most importantly it always has an energy of some kind that government communication (in seeking to appear neutral) regularly lacks - having that energy is why it doesn’t feel “bland”. I wonder if they have internal documents to guide their writers, or if it’s mostly information stored in the heads of Lesley Fair and Michael Atleson (who between them seem write most - all? - of the posts).
I'm with you on this. I can't honestly think a good use case for the average user to generate audio this way. Maybe some niche use case in like movie or tv production where you can generate a missing line without flying in an actor or something. Or maybe for generating dialogues for videogames. But those are business use cases, not things for the genral public.
We're getting to the point where all voice conversations will need to be authenticated via OTP, even between family members, on the phone. Especially for banking, etc.
It is a big issue in India. We have a few Bollywood celebrities with "trademark voices" - voices so distinct you would instantly associate it with that celeb. There is a Huge mimicry culture with 100s of extremely talented mimics who can clone any voice. On top of which, there is a gigantic radio audience, so the celebs despite making millions in Bollywood films, advertise cement, coconut oil, fountain pens, tobacco, beauty creams, online casinos etc in radio clips, using their distinctive voice.
This makes for a rather explosive combination. I could, as some tobacco exec, hire some mimic to promote cigarette sales using a celebrity's distinct voice. By the time the regulators catch up, the spot has aired a few million times & made a potload of money.
A bunch of celebs[1][2] have trademarked their voice...but enforcement is spotty.
[1] https://economictimes.indiatimes.com/news/new-updates/amitabh-bachchans-voice-image-cant-be-used-without-permission-says-court/articleshow/95762504.cms?from=mdr https://economictimes.indiatimes.com/news/new-updates/amitab...
[2] https://www.financialexpress.com/archive/when-celebrities-seek-copyrights/729569/ https://www.financialexpress.com/archive/when-celebrities-se...
It's amazing how many think declaring something illegal will stop criminally-minded people having and using it.
Do you want to tell us why it should be illegal? Comparing something like this to nuclear weapons is a bit hyperbolic, at least without giving more context.
Serious question. What is the difference in implications between this and a professional voice impersonator? I don't think it's as dangerous as we think it is. All of the consequences that Play.ht bring to society are already possible today and have been for some time. The difference is that it will be easier, but I don't think that makes it any more dangerous.
there's a massive higher bar in effort and costs in getting an impersonator
whilst this is cheap and easy - increasing the potential for scams in big way - even to the point of automating the scam
> The difference is that it will be easier
Seems like you know what the difference is, you just haven't assigned it the proper weight.
Two differences I can think of, either side the debate:
1) The impersonation can be carried out in real-time by the criminal themselves. No need to employ anyone else. (No trail leading to them.)
2) Pro impersonators aren't common in society. They are limited as an asset and not duplicatable. So, using one cannot spread like wildfire and overwhelm our awareness that voice impersonation is something of a common risk.
Maybe the second could hold the first in check. I think disruptive tech like this & similar advances in visuals come with a societal impact that lessens potentials for realising the bigger fears.
But people just love fears.
Scaling + ease of use.
Compare:
1) Use lots of time to find person who can impersonate specific other person. Unless you give them a lot of money or threats to shut up you can't use them to real time decieve someone.
2) Clone 1 million voices from tiktok in 1 minute. Contact 10 million relatives with a synth voice that is intelligent enough to answer questions.
We will have billions of AI's, containers, programs, agents running around trying to deceive absolutely everyone and their grandmother 24/7 soon.
Right on the top of their page is the example: "good afternoon sir, I will just need your credit card number and security code to proceed." Wow.
I have a "large" (~40K 10s lines) corpus of captioned dialogue from a video game that I briefly investigated training a model similar to this to "clone voices" with, but pretty quickly came to the realisation that doing so would be pretty unethetical to all involved.
It became more apparent to me how icky this is as the voice actor of one of the most iconic characters in the game died suddenly 10 days ago...
> This is the first startup here where I think the tech should essentially be illegal.
Yes, I agree, only criminals should be allowed to freely run it.
You are right, the technology will become ubiquitous, therefore, at least for platforms like us, it's a responsibility to have countermeasures and safeguards to prevent abuse and harm people. There'll always be people who will find ways to abuse but making it more and more difficult and evolving on that seems like a way forward.
We have these measures in place and are working on others to make sure the technology is used towards the betterment of humanity.
1/ Auto moderation on text to block harmful/malicious speech.
2/ As someone pointed out in the comments, we had a manual review process in place where the user is required to read out a consent and a member from Play.ht would review it before approving the voice. We're working on improving and adding this back.
3/ The user facing service is paywalled so we don't allow everyone in.
4/ Users trying to create malicious content are flagged and reviewed.
5/ A classifier to detect AI generated speech
Cofounder here,
What you see in the above demo is a very rate-limited demo of our upcoming model.
We realize how dangerous this technology can be and have built a lot of mitigations on our main product (Play.ht) to reduce possible abuse:
- We strictly moderate the generated text of any sexual, offensive, racist, or threatening content. It automatically gets detected and blocked.
- We built and are offering for free a tool that can identify AI generated vs human-generated audio (https://play.ht/voice-classifier-detect-ai-voices/ https://play.ht/voice-classifier-detect-ai-voices/), we will continue to invest in this tool, and we hope it helps with deploying this technology safely.
- If we get any reports of a cloned voice without consent, we block the user and remove the voice instantly.
- The price of high-fidelity voice cloning is too high for scammers to use at scale; we have been live with it for four months and haven't had any cases of abuse so far.
Like any technology, it has the potential to be abused, and we are working hard to mitigate that and deploy it safely. We will continue to observe the use cases and user feedback and improve the safety of the service accordingly.
Since we launched voice cloning 4 months ago, we have seen enough genuine use cases which motivated us to keep moving forward and figure out safe ways to make the technology useful for all.
> haven't had any cases of abuse so far.
How do you know this?
>We strictly moderate the generated text of any sexual, offensive, racist, or threatening content.
This won't be the problem. My voice calling my parents asking for money to be sent to a random account will be the problem. And none of that will be sexual, offensive, racist, or threatening.
>we are working hard to mitigate that and deploy it safely.
How?
>we have seen enough genuine use cases
What?
4y ago
Honestly I'm starting to wonder this about AI in general. I mean realistically, there's a decent chance will be looking at general AI soon. The best-case scenario endgame of that is creating a benevolent God. It might be time to start asking ourselves if that's what we want.
Pretty sure that's something that ought to have been discussed before any of this ever started, but you know, scientists, could, should. I look forward to the chaos and destruction and all these "brilliant" software developers wringing their hands saying they couldn't possibly have imagined such horrible outcomes from their fun money-making venture that just so happened to undermine the concept of a shared reality.