4 ms·
Demos here: https://resemble-ai.github.io/chatterbox_demopage/ https://resemble-ai.github.io/chatterbox_demopage/ (not mine) This is a good release if they're
by Mizza 1y ago
Demos here: https://resemble-ai.github.io/chatterbox_demopage/ https://resemble-ai.github.io/chatterbox_demopage/ (not mine)
This is a good release if they're not too cherry picked!
I say this every time it comes up, and it's not as sexy to work on, but in my experiments voice AI is really held back by transcription, not TTS. Unless that's changed recently.
- pinter69 1y agoRight you are. I've used speechmatics, they do a decent jon with transcription
- theyinwhy 1y ago1 error every 78 characters?
- pinter69 1y agoThe way to measure transcription accuracy is word error and not character error. I have not really checked or trusted) speechmatics' accuracy benchmarks But, from my experience and personal impression - it looks good, haven't done a quantitative benchmark
- theyinwhy 1y agoThanks for your constructive reply on my bad joke. I was referring to your original comment where you had a typo. I just couldn't resist, sorry.
- deleted 1y ago[deleted]
- ianbicking 1y agoFWIW in my recent experience I've found LLMs are very good at reading through the transcription errors (I've yet to experiment with giving the LLM alternate transcriptions or confidence levels, but I bet they could make good use of that too)
- mikepurvis 1y agoI was going to say, ideally you’d be able to funnel alternates to the LLM, because it would be vastly better equipped to judge what is a reasonable next word than a purely phonetic model.
- ianbicking 1y agoIf you just give the transcript, and tell the LLM it is a voice transcript with possible errors, then it actually does a great job in most cases. I mostly have problems with mistranscriptions saying something entirely plausible but not at all what I said. Because the STT engine is trying to make a semantically valid transcription it often produces grammatically correct, semantically plausible, and incorrect transcriptions. These really foil the LLM. Even if you can just mark the text as suspicious I think in an interactive application this would give the LLM enough information to confirm what you were saying when a really critical piece of text is low confidence. The LLM doesn't just know what are the most plausible words and phrases for the user to say, but the LLM can also evaluate if the overall gist is high or low confidence, and if the resulting action is high or low risk.
- miki123211 1y agoThis is actually something people used to do. old ASR systems (even models like Wav2vec) were usually combined with a language model. It wasn't a large language model, those didn't exist at the time, it was usually something based on n-grams.
- vunderba 1y agoPairing speech recognition with a LLM acting as a post-processor is a pretty good approach. I put together a script a while back which converts any passed audio file (wav, mp3, etc.), normalizes the audio, passes it to ggerganov whisper for transcription, and then forwards to an LLM to clean the text. I've used it with a pretty high rate of success on some of my very old and poorly recorded voice dictation recordings from over a decade ago. Public gist in case anyone finds it useful: https://gist.github.com/scpedicini/455409fe7656d3cca8959c123938f800 https://gist.github.com/scpedicini/455409fe7656d3cca8959c123...
- causal 1y agoPlay with the Huggingface demo and I'm guessing this page is a little cherry-picked? In particular I am not getting that kind of emotion in my responses.
- backnotprop 1y agoIt is hard to get consistent emotion with this. There are some parameters, and you can go a bit crazy, but it gets weird…
- lukax 1y ago[flagged]
- junon 1y agoYou should really disclaim that you're affiliated. https://news.ycombinator.com/item?id=41866830 https://news.ycombinator.com/item?id=41866830
- echelon 1y agoI absolutely ADORE that this has swearing directly in the demo. And from Pulp Fiction, too! > Any of you fucking pricks move and I'll execute every motherfucking last one of you. I'm so tired of the boring old "miss daisy" demos. People in the indie TTS community often use the Navy Seals copypasta [1, 2]. It's refreshing to see Resemble using swear words themselves. They know how this will be used. [1] https://en.wikipedia.org/wiki/Copypasta https://en.wikipedia.org/wiki/Copypasta [2] https://knowyourmeme.com/memes/navy-seal-copypasta https://knowyourmeme.com/memes/navy-seal-copypasta
- bschwindHN 1y agoHeh, I always type out the first sentence or two of the Navy Seal copypasta when trying out keyboards.
- lvl155 1y agoCan’t you get around that by synthetic data?
- gapeleon 1y agoFor English-only an non-commercial, Parakeet has been almost flawless for me. https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2 https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2 I use it for real-time chat and generating subtitles. It can do a tv show in less than a minute on a 3090. Whisper always hallucinated too much for me. It's more useful as a classifier.
- causality0 1y agoIt would be nice if there was some type of front-end integration that would present the user with a list of heteronyms found in the text and ask for clarification for each one. As well as having lists of common phrases to compare them against. There's really no excuse for an LLM to mispronounced "live feed" or "live here".
- vivzkestrel 1y agoThere seems to be a 40 second limit that nobody s talking about, once your audio crosses the 40 second length, it gets cut off