Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
_josh_meyer_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
Prompt-to-Voice: create a new voice with Generative AI
(coqui.ai)
10 points
by
_josh_meyer_
4y ago
|
2 comments
62.
▲
by
_josh_meyer_
4y ago
Prompt-to-voice creates a new, unique voice given a text prompt. Similar to stable diffusion, but for voices. Then those voices can generate speech via TTS.
63.
▲
by
_josh_meyer_
4y ago
Newest release from https://coqui.ai Direct-able, controllable generative AI Voices
64.
▲
by
_josh_meyer_
4y ago
Voice cloning from a segment of Merkel via Coqui.ai
65.
▲
All Generative AI {video/voice/text/music} in one video
(twitter.com)
6 points
by
_josh_meyer_
4y ago
|
1 comments
66.
▲
Sequoia's Map of Generative AI
(twitter.com)
2 points
by
_josh_meyer_
4y ago
|
2 comments
67.
▲
by
_josh_meyer_
4y ago
Map of interesting companies and applications in the Generative AI Landscape, according to Sequoia (Oct. 2022)
68.
▲
Waitlist for AI Voice Studio
(twitter.com)
1 points
by
_josh_meyer_
4y ago
|
1 comments
69.
▲
by
_josh_meyer_
4y ago
Waitlist for early access to the Coqui Studio
70.
▲
by
_josh_meyer_
4y ago
imo this kind of tech is useful to supplement the artist, not replace them. You listen to a singer because you know their voice more than anything. Bob Dylan was a pretty bad singer, but I listen to him because of some emotional connection.
71.
▲
AI can now sing like Freddy Mercury
(youtube.com)
7 points
by
_josh_meyer_
4y ago
|
3 comments
72.
▲
by
_josh_meyer_
4y ago
From the Coqui TTS project < https://github.com/coqui-ai/TTS >
73.
▲
How Inflation Works: An illustrated guide for the rest of us
(finmasters.com)
3 points
by
_josh_meyer_
4y ago
|
0 comments
74.
▲
by
_josh_meyer_
4y ago
same company, yes, but major updates since our demo: 1) Quality: the underlying model is much, much better, and so is the quality of the Voice Clone. 2) Productization: previously we just had a stand-alone demo, now we've launched the
75.
▲
by
_josh_meyer_
4y ago
yours is a solid point -- would be nice to have that "wow factor" right off-the-bat
76.
▲
by
_josh_meyer_
4y ago
still getting the error? I can't replicate on MacOS Monterey
77.
▲
by
_josh_meyer_
4y ago
delicious documentation for your reading pleasure:) https://tts.readthedocs.io
78.
▲
by
_josh_meyer_
4y ago
thanks for linking -- the core project is indeed open source! the founding team all worked at Mozilla on TTS and DeepSpeech, then we spun out to create Coqui <3
79.
▲
by
_josh_meyer_
4y ago
Product Hunt --> https://www.producthunt.com/posts/coqui
80.
▲
by
_josh_meyer_
4y ago
more fine-tuned control over enunciation/emotion and the like is very much in the works... stay tuned :)
81.
▲
by
_josh_meyer_
4y ago
Thanks, Neil :D The current model definitely does an "implicit" accent conversion to American English, but in the near future we'll have something that keeps accents on separate sides of the pond, so to speak :)
82.
▲
by
_josh_meyer_
4y ago
the voices are demoed for [videogames]( https://www.youtube.com/watch?v=x8tEdwll_CY&t=4s ), [dubbing]( https://www.youtube.com/watch?v=TBjdUY3_ccQ&t=1s ), and [post-production]( https://www.yo
83.
▲
by
_josh_meyer_
4y ago
thanks for flagging!
84.
▲
by
_josh_meyer_
4y ago
why don't you try :D ? Make a clone where you're talking like Hal... see how it does ;)
85.
▲
by
_josh_meyer_
4y ago
currently it's only user-facing, but still early days:)
86.
▲
by
_josh_meyer_
4y ago
of course! we've got an open ethics discussion going here: https://github.com/coqui-ai/TTS/discussions/1036 and our chatrooms are probably the best/fastest way to get ahold of us: gitter.im/coq
87.
▲
by
_josh_meyer_
4y ago
video game demo --> https://www.youtube.com/watch?v=x8tEdwll_CY
88.
▲
by
_josh_meyer_
4y ago
walk-though video --> https://www.youtube.com/watch?v=ri1U-bJ-6vc&t=4s
89.
▲
by
_josh_meyer_
4y ago
thanks for flagging! what OS are you on? iOS seems working OK atm
90.
▲
by
_josh_meyer_
4y ago
For this zero-shot approach, there's actually no "fine-tuning", and from what we can tell ~30 seconds is optimal (the text isn't actually even used... you can say anything!) For longer outputs and legit fine-tuning, that
More ›