5 ms·
It's unfortunate that another comment that links to Rhasspy has been downvoted (I assume because it lacked any other context) so I wanted to mention the project
by follower 5y ago
It's unfortunate that another comment that links to Rhasspy has been downvoted (I assume because it lacked any other context) so I wanted to mention the project with some additional context: https://rhasspy.readthedocs.io/ https://rhasspy.readthedocs.io/
While I've not used the entire Rhasspy project myself (but trying it out is on the long list of things to do :) ) I have used the offline Text-To-Speech sub-project Larynx...
...and it is amazing!
Larynx is significantly ahead in terms of quality of output & variety of voices (fifty--across multiple languages, accents & genders) of any other FLOSS Text-To-Speech project I've tried.
I think the relative new-ness of the project is part of the reason Larynx (https://github.com/rhasspy/larynx/ https://github.com/rhasspy/larynx/) currently flies under the radar.
If the rest of Rhasspy is as good as Larynx I'd imagine it's worth trying out.
Larynx demo video: https://www.youtube.com/watch?v=hBmhDf8cl0k https://www.youtube.com/watch?v=hBmhDf8cl0k
Samples of pre-trained voices: https://rhasspy.github.io/larynx/#en-us https://rhasspy.github.io/larynx/#en-us
- lukifer 5y agoI can vouch for Rhasspy, it's an amazing and flexible piece of software, though it does require some setup and tech knowledge (albeit with a usable web GUI); and it's very DIY on defining the actual voice commands. I recommend pairing it with Node-RED [0] for routing commands to devices, it has plugins for most things. The only thing I struggled with was getting the wake-word config right: I could never find the right balance point where it responded every time, without also having annoying false positives, so I ended up turning it off. It does support multiple wake-word engines; I'm gonna have another go with Picovoice Porcupine now that they're opened up custom wake-word training for free. I'm most heavily experienced with Rhasspy's sister project, voice2json [1], which I used to build a voice-controlled car jukebox [2], and it's been working fantastically. (It triggers from a Bluetooth remote, so no wake-word issues.) The two projects share the same core engine. For hardware, Raspberry 3/4 perform quite well, and strong recommend for ReSpeaker [3] for audio (either usb or 4-mic hat). [0] https://nodered.org/ https://nodered.org/ [1] http://voice2json.org/ http://voice2json.org/ [2] https://github.com/lukifer/voicetunes https://github.com/lukifer/voicetunes [3] https://www.seeedstudio.com/category/Speech-Recognition-c-44.html https://www.seeedstudio.com/category/Speech-Recognition-c-44...
- Semaphor 5y agoSoftware has been mostly solved for a while. My issue is hardware. The way I read, is that to get proper voice recognition in many circumstances, you need a microphone array (for every DIY Alexa and Dot). Now you have a PI, the array installed on it and… it just stands around looking ugly and accumulating dust? You want a case, but then from my research, cases can easily interfere with those arrays. So you need one that’s custom-made for the array you are getting. But no array I’ve seen does actually come with such a case. Back when I asked (relevant subreddits and on tildes before I deleted my account), no one could tell me that any of my research had been wrong, but no one had a solution either. I posted the threads almost 2 years ago, so maybe things changed? I’m currently still using Alexa, but besides privacy reasons I’d also love custom software that can take the idiocy out of my assistant (mainly by using pre-configured commands that do what I want instead of sometimes guessing what I want; also for on-the-fly language switching, Alexa is atrocious when you want to request a band that’s not in your primary language) I could probably get away with the kitchen and office assistant using a normal microphone, but both bedroom and living room need to recognize voices from most directions (and in the case of the living room, also have decent recognition through music playing). If anyone has any solutions, I’d love to hear them.
- ivanhoe 5y agoWhen I researched ReSpeaker was one of the better solutions, they provide various mic-array PiHats, as well as professionally looking cases for the device. They also sell pre-build devices I think, if you're not into DIY. https://www.seeedstudio.com/category/Speech-Recognition-c-44.html https://www.seeedstudio.com/category/Speech-Recognition-c-44... (Just to male it clear, I'm not affiliated with them in any way)
- Semaphor 5y agoYeah, ReSpeaker looks like they have the best hardware for this. But the issue was, that they only sell one case, for their old array, which is barely available any more (neither case, nor array. But! Now they sell the new array as USB array in a proper case! That’s not perfect as I’ll have 2 devices instead of one to place/hide, but it’s still a lot better than last time I checked their site. Thank you, now I have plans for 2022 :) Hah, and of course there is only one left in stock :/ edit: And I found a company in Germany that not only resells them for almost the same price, they even have a lot of stock.
- 2Gkashmiri 5y agohttps://ai-service-demos.go-aws.com/polly https://ai-service-demos.go-aws.com/polly check the neural then british, Amy. the "smoothness" is uncanny. the samples are "almost" there with larynx you linked. good but take kathleen (glow_tts), there is "still" some robotic in there. is this something that can be improved by tweaking the training ? this sounds really cool to be used at home
- follower 5y agoOh, yeah, I'm definitely aware there's still a quality gap when compared to proprietary online TTS options--but I'm specifically interested in FLOSS+offline for my purposes. And Larynx is game-changingly ahead of the other FLOSS+offline options. (Probably the highest quality voice is the one that's used for the demo video narration--which when I first heard it I had to skip to the end of the video to confirm it wasn't a live human. :) ) My (mostly uninformed) impression is that there's room for training tweaking/improvement given how young the project is. And there's also multiple stages to the generation process so presumably there's opportunities at each stage.
- 2Gkashmiri 5y agoyeah, https://www.youtube.com/watch?v=hBmhDf8cl0k https://www.youtube.com/watch?v=hBmhDf8cl0k at the end says southern female english https://rhasspy.github.io/larynx/#en-us_southern_english_female-glow_tts https://rhasspy.github.io/larynx/#en-us_southern_english_fem... but the sample is NOT like the video, maybe the samples are old. the video is a great example
- avnigo 5y agoThe Rhasspy video overview is pretty great too, I'm impressed. https://www.youtube.com/watch?v=IsAlz76PXJQ https://www.youtube.com/watch?v=IsAlz76PXJQ
- deleted 5y ago[deleted]
- 3np 5y agoThis is still buried too but someone just released a promising HA integration of Rhasspy https://news.ycombinator.com/item?id=29565983 https://news.ycombinator.com/item?id=29565983 https://homeintent.io/ https://homeintent.io/