9 ms·
Show HN: Supertone Shift – AI powered Real-time voice changer
Supertone's Shift offers real-time voice changing technology. It lets users immediately switch to any selected voice. Just pick a voice and begin speaking. Shift is suited for VTubers, content creators, and gamers, as well as anyone who wishes to accurately express their chosen persona's voice. Try out Supertone Shift now.
>> https://product.supertone.ai/shift https://product.supertone.ai/shift
- deleted 2y ago[deleted]
- camillomiller 2y agoExcept for purely non lucrative entertainment use cases with a very high novelty factor, I am struggling to see productive use cases for all these AI applications that don't involve some form of deception or at best disingenuous marketing.
- slipheen 2y agoWould imagine the same sort of reasons people do v tubing in general, such as safety and anonymity.
- nounaut 2y agoIf they generate good quality then I suppose voice acting could have good use of it.
- phil-martin 2y agoI can think of a few applications of this technology, although some may fall into the deception category, albeit harmless in my view: - overcoming social anxiety in voice or online calls. It doesn’t take very many bullying incidents during childhood to become convinced you have a horrible or weird voice. I can see this being used as a useful tool to make people feel more comfortable by having a different voice - amateur interactive fiction development. Having your characters have a real voice in a game in response too the players commands is a real need, and being able to record it yourself and be a different character would be a huge enabler of creating something for a solo developer. - internal HR videos/podcasts. Creating these can be very expensive, needing different persons reading out dialogue could significantly reduce the effort in recording and producing these - another instrument for music creators. Auto tune is a very common tool for music production for all skill levels, and this could be applied in a very similar way It no doubt can be used for disingenuous purposes, any technology can. But these can be real life improving tools enabling many people to do things they never thought possible. The idea of participating in Q&A session in a webinar would be far too confronting and inconceivable for many people, but to be able to do it semi-anonymously with a different voice would eliminate much of the anxiety preventing them
- gattilorenz 2y agoI also can't help thinking of the "Melanie speaks" episode of 99% invisible [1]. Of course this only works for your "online persona", but still the idea of impacting how you are perceived by working on your voice... is a thing. [1] https://99percentinvisible.org/episode/melanie-speaks/ https://99percentinvisible.org/episode/melanie-speaks/
- deleted 2y ago[deleted]
- jack_pp 2y agoI think this is huge for new content creators that are not native speakers to get rid of the accent. Also if it enables multiple people to sound the same then you can have a YouTube channel with a larger team but only one voice
- deleted 2y ago[deleted]
- kthartic 2y agoI don't believe it modifies the accent. I noticed I could hear his asian accent coming through every character, so it seems to just modify the voice but not the intonation
- hiergiltdiestfu 2y agoAgreed, at the very least it doesn't seem to change _my_ accent, just things like color, tone, pitch, etc.
- kthartic 2y agoThis is huge for indie game developers! They can voice every line of dialogue for every character themselves (or with just 1 professional voice actor). Text-to-speech AI voice generators exist, but you don't have fine control over the emotion/expressiveness/intonation of the lines like you do with this approach.
- dannyw 2y agoAs someone who makes indie games as a passion and creative outlet, tools like these drastically expand my creative possibilities.
- Tanoc 2y agoThere's a balance to the ecosystem, though. People in the creative fields have always had to rely on eachother to fill in gaps in skill because it's mutually beneficial. With things like this voice changer, one has to think what opportunities are being taken away from others compared to what opportunities the technology affords oneself. So far we've been screwing that balance up pretty egregiously with these AI tools where one implementation cuts the employment prospects and creative participation of a dozen people.
- deleted 2y ago[deleted]
- kunley 2y ago[flagged]
- hagbard_c 2y ago...needing an excuse to get access to microphones to solve the problem stated on their own front page (https://supertone.ai/ https://supertone.ai/) as We need voice source material to train the AI.
- daredoes 2y agoUnless the app is directly lying, it says "For the best performance of Shift we need to listen to your voice for 10 seconds. We do not collect your voice data from this app." I haven't put a network sniffer or anything on it yet though. Just wanted to take a peak at the UI
- jen729w 2y agoI make videos, it might be handy to be able to ‘be’ someone else. I can for sure see a use for this.
- deleted 2y ago[deleted]
- kunley 2y agoFunny to see myself being downvoted for a harmless call for a reason, so, few more comments: The fact someone did something and put a substantial effort into it is not a reason good enough to justify said effort (other than the benefits of learrning) and the product that was created. The world is full of things which actually made it worse place. Another comment is actually a meta-comment and might be shocking to some people here: downvoting is not a good method of making someone stop saying statements perceived by some as unconfortable. In fact, there is nothing wrong with earning points first and then burning them with saying comments that are feared by certain individuals to the level that they "must" be downvoted...
- 2y ago
- rcarmo 2y agoI can see this being interesting for gamers and more whimsical pursuits, but I'm more curious about neural speech synthesis for both normal speech and singing--the first because there is a pretty strong demand for automated narration of training videos, and the second because of my music hobby--other than vocaloids and a few niche DAWs, I haven't found any nice Open Source tooling for the latter (the former I can mostly do with XTTSv2).
- edwcross 2y agoFrom what I found, XTTSv2 is based on the Coqui Public Model License, which explicits disallows commercial commercial usage: "This license allows only non-commercial use of a machine learning model and its outputs." So, from what I understand, I cannot use it and then upload the training video to Youtube. Or can I?
- margorczynski 2y agoI guess if it is demonetized it should be ok? Or maybe not if your other content or activity is commercial, as even if the video in itself doesn't make money it would indirectly promote your other commercial activity. Interesting legal problem.
- watersb 2y agoVery interesting! I would like some clarity on the Terms of Service clause 4: > The content created using Supertone Shift remains your property. However, by using our Services, you grant Supertone a worldwide, non-exclusive, royalty-free license to use, reproduce, adapt, and display content solely for the purpose of operating and improving Supertone Shift. This license does not grant Supertone any rights to sell or distribute your content. Does Supertone Shift need the user content in order to further improve the product during the beta period? Or does it need the user content in normal operation (for example, running the conversion on remote servers vs local processing)? I can see some hesitation from people if you're recording everything they say, and keeping that recording for an indefinite period of time. I can appreciate that there may be a problem enforcing a "Don't use our product for evil" clause, if you can review usage. The challenge here seems overwhelming.
- htrp 2y agoLooks like facebook's ToS, we may need your data for some unspecified purpose ("AI model training") that we can't even dream of right now, so we'll just take all the rights
- deleted 2y ago[deleted]
- echelon 2y agoThere are dozens of other products in this category, including completely open source ones you can fine tune. Commercial applications like Voice.Ai and Koe are real time and have celebrity and anime voices respectively. The RVC ecosystem on GitHub has dozens of different real time open source voice changers. I haven't kept up with the SOTA, but they're incredible, fine tunable, and 100% local. https://voice.ai https://voice.ai https://koe.ai https://koe.ai https://m.youtube.com/watch?v=zkaBK5erB2c https://m.youtube.com/watch?v=zkaBK5erB2c
- andoando 2y agoIve tried making this exact product using all of these services, including using github repo koi is based on. They all use like 50% of my cpu to get real time. I was able to get actual low latency with koi, but still massive cpu usage. And theres no community of models for it either. Perhaps someone who really knows what theyre doing could optimize these open source models but its not me
- gardenhedge 2y agoThis is awesome. Very futuristic
- darkoob12 2y agomore dystopian. yet another "contribution" of AI for destroying the society via misinformation.
- Springtime 2y agoSpeaking generally, there's an undermarketed positive privacy aspect to such voice changers, in helping with both protecting a user's identity against data scraping and doxxing. Additionally, like one other comment touched on, some people have strong accents that make communication in videos challenging (and can be a turn off for audiences and prejudice initial impressions). Though by using an online service approach it means providing one's real voice to a service that may be using it for further training. Users have to make the call whether they feel they're good stewards.
- hiergiltdiestfu 2y agoMy accent seems to translate very well through the app, tho. The tool "only" changes color, pitch, tone, etc. not _how_ I say stuff, i.e. pronounciation, choice of words, ...
- qntmfred 2y agoAll technology can be used for good and evil If humanity can figure out how to make machines think, may we should also figure out how to stop doing evil to each other
- cynicalsecurity 2y agoYes, but we can't stop it.
- AlecSchueler 2y ago
- jen729w 2y ago> The installation has completed. Please restart your Mac. Seriously?
- simse 2y agoIt's worth it!
- giankam 2y agoWould have liked to know it before installing.
- earthnail 2y agoSame question. What did I just install that required a restart?
- mintplant 2y agoVirtual microphone driver, perhaps?
- giankam 2y agoNot only, it's not possible to quit the installer. Had to kill it and then look for changes done to the system. Hope I've been able to find them all but really upsetting.
- jen729w 2y agoSo, Supertone Shift creators: this is really good! The first time you hear your own voice as a K-pop star or a nymph it’s genuinely startling. Just improve the installer so I don’t feel like I’ve been scammed by malware!
- terhechte 2y agoCurious Question: Given the low latency, does it run the computation on device or over the network? If on device, are there minimum CPU requirements?
- catapart 2y agoVery interested in this answer! I'd really like to see it on the website for any AI I'm considering. It's an entirely different proposition as to whether you're getting a utility or a service.
- tiborsaas 2y agoThis looks like an amazing tool for indie game developers. Even musicians could find this an amazing help to add some unique tones.
- michaelmior 2y agoThis seems really cool and I can see some great use cases. But the marketing is very odd to me. It says it will let me express myself in a voice that is truly my own…but I can already do that with my natural voice. That seems more likely to be unique than what I would get by adjusting it in software.
- themoonisachees 2y agoI guess the wording is awkward, but as a trans person, I still resonate with it. I'm acutely aware it's not going to be "my voice", but neither is the one I have right now.
- michaelmior 2y agoThanks for the explanation. This is definitely something I hadn't considered.
- mintplant 2y agoIt's funny to me that we just had a big front-page thread full of HN users questioning the value of diversity, and then this thread where people struggle to figure out the obvious trans market for voice-changing software.
- squigz 2y agoNon-verbal people might also be interested in such things
- sdfgtr 2y agoThat particular line is definitely directed towards people with gender identity issues.
- idiotsecant 2y agoPro tip: Some people do not consider their natural voice 'their' voice.
- corytheboyd 2y ago
- itronitron 2y agoI wonder if this could be applied to educational videos to make the material seem less challenging for children.
- rtcode_io 2y agoNice to see a venture from South Korea!
- drivingmenuts 2y agoWeren't we able to do this before AI? I'm not sure I get what AI is bringing to the table/value-adding for this particular technology, except marketing hype.
- cma 2y agoWasn't that very basic pitch shifting only?
- drivingmenuts 2y agoIt was a bit more complex than that, but that's more or less what this software, which claims the benefits of AI, is doing. It's not, near as I can tell, changing inflection or tone or doing anything other than changing the pitch and maybe adding some frequencies. It's not even producing natural tones. None of the voices they're demoing sound real in any way. It's a toy, no more, no less.
- trashcluster 2y agoIf it was compatible as a VST plugin for DAWs it would be even more useful than a standalone software. From skimming through the website it seems that Supertone already make a VST plugin so it may be a matter of time before Shift becomes a VST too.
- hollowayaegis 2y agoSelf plug, but I've been developing a local AI voice changing VST [1] (bring your own RVC models, or use builtins). It works in DAWs in realtime on modern macs. [1] https://audio.sunflower.industries https://audio.sunflower.industries
- desro 2y agoThis looks cool and I've downloaded it. Clicking on the "free" tier on the subscription page brings you into Stripe checkout for the $6 tier, FYI.
- vouaobrasil 2y agoAll this technology is leading to a world where we can present second-life/alternative identities cohesively online. I wonder if this is going to cause a global decline in the ability for people to express themselves, since it is now so easy to create an identity online that is different than your real-life identity. I think it's rather sad. Yes, there are some fringe use-cases perhaps but I think this is the wrong direction for humanity. We should find more value in what we already have rather than inventing arbitrary things like this to hide away from real acceptance of ourselves.
- latexr 2y agoIt will first lead to a world where fake videos of celebrities will be used to scam you, and your own voice will be used to scam your relatives. Both of those are happening today. Ironically, this will lead to a work where we need to use these fake personas online to not have our lives messed with offline. I don’t fully agree with your first paragraph, but I do agree with the second one.
- Ukv 2y ago> and your own voice will be used to scam your relatives. Both of those are happening today. I can't really see it becoming common for cold-calls that pretend to be someone the victim knows (like the terrifying ransom calls), since the operations work at a huge scale expecting most people to not even pick up a "scam likely" call. Even given free and instant model tuning, just having to find voice clips of the person prior to each unanswered automated call seems like it would tank the quantity they're able to make. I imagine there will be plenty of unevidenced claims that this is what scammers did to them though. Victims have always said "it sounded exactly like him/her", and from there it's more comforting for someone to conclude they must've been fooled by a sophisticated attack rather than something simple. For more targetted phishing, like pretending to be a company's CEO and phoning employees to get access, I could definitely see it being used. I think we're probably going to have to move "person sounds like boss over the phone" from "plausible to fake" to "trivial to fake".
- 2y ago
- bogwog 2y agoSeems like we're getting closer and closer to Star Trek's universal translator
- jl6 2y agoWould it be possible to embed a watermark in the generated audio? Many people will use voice changing tech for honest purposes, but there will always be those acting to ruin it for the rest of us. There are just too many scenarios where faking your voice confers an illicit benefit. I know watermarks are never foolproof, but they may deter casual misuse.
- andoando 2y agoI was trying to make this myself earlier but every single AI model I found used something like 50% of my CPU or GPU. Any idea how this is possible? Voicemod does something similar and I couldn't figure it out. Is it actually AI or is this just shifting pitch/reverb/etc
- jzemeocala 2y agoFun looking product. Sad to see no Linux support (yet?). Would you be interested in any help porting/maintaining a Linux release?
- fzaninotto 2y agoWhy does the Mac installer require admin right and a restart? Giving admin rights to an installer requires trust in the vendor. Supertone Shift is just a newborn. I cancelled the installation because of that. I would love to test the technology without the risk of damaging my computer!
- moralestapia 2y agoThanks for this, I was very eager to try it out but this is a always a deal breaker.
- desro 2y agoI use the great, free, "Suspicious Package" app [0] to inspect installers like these. In fact, it was Supertone Shift's installer that prodded me to seek it out (I happened to find and install Shift a couple of weeks ago). In this case, it needs admin permissions to install to `/Library/Application Support` as well as `/Library/Audio`. It needs to restart in order for the HAL driver to be loaded (this provides the virtual audio interface for using the app with Teams, Zoom, etc.) The preinstall/postinstall scripts simply handle the app's directory in Application Support. I decided it was safe enough, and had some fun playing with it. It contacts what it claims are licensing servers (when it starts), and won't start without it. It wanted to keep contacting those servers constantly, but blocking its network access via Little Snitch didn't prevent it from functioning. The network traffic was in the single-digit kilobyte range, so I felt reasonably confident no audio data was being looted. [0] https://mothersruin.com/software/SuspiciousPackage/ https://mothersruin.com/software/SuspiciousPackage/
- WORMS_EAT_WORMS 2y agoCongrats! This is amazing work