9 ms·
Chatterbox TTS
- decide1000 1y agoHow does it perform on multi-lingual tasks?
- yjftsjthsd-h 1y agoThe readme says it only supports English
- benob 1y agoWatermarking is easily disabled in the code. I a wondering when they will release model weights with embedded watermarking.
- kiririn7 1y agodefinitely worse than the new elevenlabs model(v3). that model is really good
- plangary123 1y agoI disagree
- xnx 1y agoYou can run it for free here: https://huggingface.co/spaces/ResembleAI/Chatterbox https://huggingface.co/spaces/ResembleAI/Chatterbox
- mpeg 1y agoA bit on the nose that they used a sample from a professional voice actor (Jennifer English) as the default reference audio file in that huggingface tool.
- echelon 1y agoSadly they don't publish any training or fine tuning code, so this isn't "open" in the way that Flux or Stable Diffusion are "open". If you want better "open" models, these all sound better for zero shot: Zeroshot TTS: MaskGCT, MegaTTS3 Zeroshot VC: Seed-VC, MegaTTS3 Granted, only Seed-VC has training/fine tuning code, but all of these models sound better than Chatterbox. So if you're going to deal with something you can't fine tune and you need a better zero shot fit to your voice, use one of these models instead. (Especially ByteDance's MegaTTS3. ByteDance research runs circles around most TTS research teams except for ElevenLabs. They've got way more money and PhD researchers than the smaller labs, plus a copious amount of training data.)
- Quarrel 1y agoFun to play with. It makes my Australian accent sound very English though, in a posh RP way. Very natural sounding, but not at all recreating my accent. Still, amazingly clear and perfect for most TTS uses where you aren't actually impersonating anyone.
- skatanski 1y agoHow does it work from the privacy standpoint? Can they use recorded samples for training?
- gardnr 1y agoPreviously, on Hacker News: https://news.ycombinator.com/item?id=44120204 https://news.ycombinator.com/item?id=44120204 https://news.ycombinator.com/item?id=44144155 https://news.ycombinator.com/item?id=44144155 https://news.ycombinator.com/item?id=44195105 https://news.ycombinator.com/item?id=44195105 https://news.ycombinator.com/item?id=44230867 https://news.ycombinator.com/item?id=44230867 https://news.ycombinator.com/item?id=44172134 https://news.ycombinator.com/item?id=44172134 https://news.ycombinator.com/item?id=44221910 https://news.ycombinator.com/item?id=44221910 https://news.ycombinator.com/item?id=44145564 https://news.ycombinator.com/item?id=44145564
- deleted 1y ago[deleted]
- pinter69 1y agoI did a quick google search before positing and only found a reference in a comment. But, I searched for the link to the GitHub.
- tomhow 1y agoThanks for posting this but it's conventional to only post links to past submissions if they had significant discussion, which none of these did.
- abraxas 1y agoAre these things good enough to narrate a book convincingly or does the voice lose coherence after a few paragraphs being spoken?
- pinter69 1y agoI consult a company in the space (not resemble) and I can definitely say it can narrate a book
- raincole 1y agoOnce it's good enough Audible will be flooded with AI-narrated books so we'll know soon. (The only question is whether Amazon would disclose it, ofc)
- landl0rd 1y agoFlip side is a solution where I can have a book without an audiobook auto-generated (or use an existing ebook rather than paying audible $30 for their version) and it's "good enough" is a legit improvement. AI generated isn't as good but it's better than nothing. Also, being able to interrupt and ask for more detail/context would be pretty nice. Like I'm reading some Pynchon and I have to stop sometimes and look up the name of a reference to some product nobody knows now, stuff like that.
- skygazer 1y agoIf you're willing to forgo the interactive LLM bit, kokoro-tts (just a script using Kokoro-ONNX) takes epubs and outputs a series of wavs or mp3s that need to be stitched together into chapters or audiobook m4a with some ffmpeg fu. I've listened to several generated audiobooks, and found them pretty good. Some nice generic narration-like prosody. It uses espeak-ng to generate phonemes and passes those to the model to render voice, so it generally pronounces things quite well. It comes with a handful of nice voices and several can be blended, but no easy voice cloning, like chatterbox, that I'm aware of. https://github.com/nazdridoy/kokoro-tts/blob/main/kokoro-tts https://github.com/nazdridoy/kokoro-tts/blob/main/kokoro-tts
- Mizza 1y agoDemos here: https://resemble-ai.github.io/chatterbox_demopage/ https://resemble-ai.github.io/chatterbox_demopage/ (not mine) This is a good release if they're not too cherry picked! I say this every time it comes up, and it's not as sexy to work on, but in my experiments voice AI is really held back by transcription, not TTS. Unless that's changed recently.
- pinter69 1y agoRight you are. I've used speechmatics, they do a decent jon with transcription
- theyinwhy 1y ago1 error every 78 characters?
- pinter69 1y agoThe way to measure transcription accuracy is word error and not character error. I have not really checked or trusted) speechmatics' accuracy benchmarks But, from my experience and personal impression - it looks good, haven't done a quantitative benchmark
- theyinwhy 1y agoThanks for your constructive reply on my bad joke. I was referring to your original comment where you had a typo. I just couldn't resist, sorry.
- deleted 1y ago[deleted]
- ianbicking 1y agoFWIW in my recent experience I've found LLMs are very good at reading through the transcription errors (I've yet to experiment with giving the LLM alternate transcriptions or confidence levels, but I bet they could make good use of that too)
- j2kun 1y agoThey should put the meaning of "TTS" in the readme somewhere, probably near the top. Or their website.
- byteknight 1y agoTTS is a very common initialism for Text-to-Speech going back to at least the 90s.
- j2kun 1y agoSo? Acronym soup is bad communication.
- aquariusDue 1y agoI miss glossaries.
- dylan604 1y agoGood writing rules can still be used even for repo READMEs where the first time an acronym is used it is spelled out to show what the acronym means. Too many assumptions being made that everyone is going to know it. Sometimes the author can be too inside baseball and assumes anyone reading their README will already know about the subject. Not all devs are literature majors and probably just never think about these things
- rapfaria 1y agoAn AI-powered browser extension that shows on hover the most likely acronym meaning, based on context you say?
- aquariusDue 1y agoI've used this one for a hot minute a few weeks ago: https://lumetrium.com/definer/ https://lumetrium.com/definer/ It also can be configured to use Ollama or an API key from other providers (OpenRouter included) and from what I gather the default prompt can be changed too. Sadly it's closed source.
- pryelluw 1y agoSilly question, what’s the lowest spec hardware this will run ?
- bityard 1y agoNot a silly question, I came here to ask too. Curious to know whether I need a GPU costing 4 digits or if it will run on my 12-year-old thinkpad shitbox. Or something in between.
- 01HNNWZ0MV43FF 1y agoI was going to report how it runs on an old CPU but after fussing with it for about 30 minutes, I can't even get it to run. Listing the issues in case it helps anyone: - It doesn't work with Python 3.13, luckily `uv` makes it easy to build a venv with 3.12 - It said numpy 1.26.4 doesn't exist. It definitely does, but `uv pip` was searching for it on the pytorch repo. I passed an `--index-strategy` flag so it would check other repos. This could just be a bug in uv, but when I see "numpy 1.26.4 doesn't exist" and numpy is currently on 2.x, my brain starts to cramp up. - The `pip install chatterbox-tts` version has a bug in CPU-only mode, so I cloned the Git repo - The version at the tip of main requires `protobuf-compiler` installed on Debian - I got a weird CMake error that I can't decipher. I think maybe it's complaining that the Python dev headers are not installed. Why would they be, I'm trying to do inference, not compile Python... I know anger isn't productive but this is my experience almost any time I'm running Somebody Else's Python Project. Hit an issue, back up, hit another issue, back up, after an hour it still doesn't run.
- thorum 1y agoWe’ll know AGI has arrived when it can figure out Python dependency conflicts
- kevin_thibedeau 1y agoIt'll just throw up its virtual hands and switch to something better after transpiling all the Python code in a fit.
- nmstoker 1y agoI've found it excellent with really common accents but with other accents (that are pretty common too) it can easily get stuck picking a different accent. For instance several Scottish recordings ended up Australian, likewise a fairly mild Yorkshire accent
- Quarrel 1y ago> For instance several Scottish recordings ended up Australian Funnily enough, it made my Australian accent sound very English RP. I was suddenly very posh.
- a_wild_dandan 1y agoI think this says more about Scottish than the model.
- m3sta 1y agoLike a professional actor!
- ltrg 1y agoI'm English (RP) and it gave me a Yorkshire accent and Scottish accent in turn.
- az226 1y agoHow does one train a TTS model with an LLM backbone? Practically, how does this work?
- cyanf 1y agoyou use a neural audio codec to encode audio into codebooks then you could treat the codebook entries as tokens and treat audio generation as a next token prediction task you then take the codebook entries generated and run it through the codec’s decoder and yield audio it works surprisingly well speech text models (tts model with an llm as backbone) is the current meta
- deleted 1y ago[deleted]
- teraflop 1y ago> Every audio file generated by Chatterbox includes Resemble AI's Perth (Perceptual Threshold) Watermarker - imperceptible neural watermarks that survive MP3 compression, audio editing, and common manipulations while maintaining nearly 100% detection accuracy. Am I misunderstanding, or can you trivially disable the watermark by simply commenting out the call to the apply_watermark function in tts.py? https://github.com/resemble-ai/chatterbox/blob/master/src/chatterbox/tts.py https://github.com/resemble-ai/chatterbox/blob/master/src/ch... I thought the point of this sort of watermark was that it was embedded somehow in the model weights, so that it couldn't easily be separated out. If you're going to release an open-source model that adds a watermark as a separate post-processing step, then why bother with the watermark at all?
- jchw 1y agoPossibly a sort of CYA gesture, kinda like how original Stable Diffusion had a content filter IIRC. Could also just be to prevent people from accidentally getting peanut butter in the toothpaste WRT training data, too.
- throw101010 1y agoStable Diffusion or rather Automatic1111 which was initially the UI of choice for SD models had a joke/fake "watermark" setting too which was deliberately doing nothing besides poking fun at people who were thinking that open source projects would really waste time on developing something that could easily be stripped/reverted by the virtue of being open source anyways.
- vunderba 1y agoYeah, there's even a flag to turn it off in the parser `--no-watermark`. I assumed they added it for downstream users pulling it in as a "feature" for their larger product.
- echelon 1y ago1. Any non-OpenAI, non-Google, non-ElevenLabs player is going to have to aggressively open source or they'll become 100% irrelevant. The TTS market leaders are obvious and deeply entrenched, and Resemble, Play(HT), et al. have to aggressively cater to developers by offering up their weights [1]. 2. This is CYA for that. Without watermarking, there will be cries from the media about abuse (from anti-AI outfits like 404Media [2] especially). [1] This is the right way to do it. Offer source code and weights, offer their own API/fine tuning so developers don't have to deal with the hassle. That's how they win back some market share. [2] https://www.404media.co/wikipedia-pauses-ai-generated-summaries-after-editor-backlash/ https://www.404media.co/wikipedia-pauses-ai-generated-summar...
- andy_xor_andrew 1y agoin my experience, TTS has been a "pick two" situation: - fast / cheap to run - can clone voices - sounds super realistic from what I can tell, Chatterbox is the first that apparently lets you pick 3! (have not tried it myself yet, this is just what I can deduce)
- CGamesPlay 1y agoCan you share one that is fast/cheap to run and sounds super realistic? I'm very interested in finding a good TTS and not really concerned about cloning any particular voice (but would like a "distinctive" voice that isn't just a preset one).
- pzo 1y agoIt's also about if you want multi lung support and if wanna run on edge devices. Chatterbox only support English.
- ineedasername 1y agoThe emotional exaggeration is interesting, though I don't think I've come across anything quite so versatile and easy to "sculpt" as Elevenlabs and it's ability to generate a voice on the basis of a description of how you want the voice to sound. SparkTTS allows some additional parameters, and it's project on GitHub has placeholders in its code that indicate the model might be refined for more fine grained emotional control. As it is, I've had some success with it and other models by trying to influence prosody and tonality with some heavy handed queues in the text, which can then be used with VC to get closer to desired results, but it's a much more cumbersome process than Eleven.
- causality0 1y agoAnyone know how this compares to Kokoro? I've found Kokoro very useful for generating audiobook but it almost always pronounces words with paired vowels incorrectly. Daisy becomes die-zee, leave becomes lay-ve, etc.
- BigBananaGuy 1y agoChatterbox sounds much more natural. The zero shot voice cloning and exaggeration feature is sick!
- nmstoker 1y agoIf you're running Kokoro yourself then it might be worth checking your phonemizer / espeak-ng installs in case they are messing up the phonemes for those words (which are then passed on as inputs to Kokoro itself)
- stevage 1y agoInteresting demo. A few observations, having uploaded a snippet of my own voice, and testing with some of my own text: - the output had some of the qualities of my voice, but wasn't super similar. (Then again, the fact it could even do this from such a tiny snippet was impressive) - increasing "CFG/pace" (whatever CFG is) even a little bit often just breaks down into total gibberish - it was very inconsistent whether it would come out with a kind of British accent or an American one. (My accent is Australian...) - the emotional exaggeration was interesting, but it seemed to vary a lot exactly what kind of emotion would come out
- Shopper0552 1y agoAnyone know a good free open source speech to text? Looking for something for my laptop which is running Fedora KDE plasma.
- santiagobasulto 1y agoWhisper?
- hoherd 1y agoWhisper has been great for me. I have a single-file uv powered python script that creates SRT files or timestamped text files from media stored on the filesystem. https://github.com/danielhoherd/pub-bin/blob/main/whisper-transcribe.py https://github.com/danielhoherd/pub-bin/blob/main/whisper-tr...
- tomp 1y agohttps://huggingface.co/spaces/hf-audio/open_asr_leaderboard https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
- bkitano19 1y agohttps://huggingface.co/spaces/nvidia/parakeet-tdt-0.6b-v2 https://huggingface.co/spaces/nvidia/parakeet-tdt-0.6b-v2
- pzo 1y agoWhisper large v3 turbo if need support of many languages and want fast enough for deployment even on smartphones (WhisperKit). Can also try lite whisper on HF if need even smaller weights and slightly faster speed.
- tuananh 1y agofor this, what does it take to support another language?
- iambateman 1y agoJust a regular reminder to tell your friends and family to be extra skeptical about phone conversations. It’s becoming much more likely that the friend who desperately needs a gift card to Walmart isn’t the friend at all. :(
- mattigames 1y agoMy bet is that the government at some point will have to put some pressure on Walmart and others to stop selling those gift cards completely, doing impersonations is getting too easy and too cheap for there not to be a flood of those scam calls in the near future.
- chii 1y agothe easiest way to defeat phone fraud is to ahead of time decide on a verbal password between family (and close friends, if they're close enough that you'd lend them money). In a real scenario, they'd know the verbal password and you can authenticate them. Drum it into them that this password will prevent other people from impersonating you in this brave new world of ai voices and even video.
- jimjimwii 1y agoThat is more or less what i did with my parents, but this approach is still susceptible to active mitm attacks. 2 factor authentication through a secure app or a trusted family member is probably also needed though i haven't tackled this part with them yet.
- chii 1y ago> 2 factor authentication through a secure app the problem is that the sort of emergency scenario in which family member would need the help is not often done or possible via a secured app. It's often just a telephone, with a number that you cannot recognize - imagine getting that phone call from a police station in the middle of nowhere when arrested, then you dont have access to any of your personal belongings as they're confiscated. The phone is a landline from the police station! Therefore, a verbal password is needed, as this scenario is exactly how a scammer would present as the emergency that they need help (usually, wire some dollars to this account to bail out).
- philipkiely 1y agoExample implementation with sample inference code + voice cloning example: https://github.com/basetenlabs/truss-examples/tree/main/chatterbox-tts https://github.com/basetenlabs/truss-examples/tree/main/chat... Still working on streaming
- pzo 1y agoIt's only for English sadly
- darccio 1y agoAre there any good options for non-English languages?
- jeroenhd 1y agoIt's not on the same level in terms of emotion, but I believe the research https://github.com/CorentinJ/Real-Time-Voice-Cloning https://github.com/CorentinJ/Real-Time-Voice-Cloning was based on is mostly oriented around Chinese first (and then English). It seems to work well enough if you and the voice you're cloning speak the same language though I haven't tested it much.
- hsavit1 1y agoanother TTS that is only supporting English. This really irritates me
- jeroenhd 1y agoFor what it's worth, there are also a whole bunch of models that speak Chinese. So far the US and China are spearheading AI research, so it makes sense that models optimize for languages spoken there. Spanish is an interesting omission on the US part, but that's probably because most AI researchers in the US speak English even if their native tongue is Spanish.
- nmstoker 1y agoMaybe that irritation could be channelled to contributing into one that supports not only English? Even small steps like tweaking docs, adding missing/extra examples, fielding a few issues in GH (most are usually simple misunderstandings where a quick pointer can easily help a beginner)
- andymcsherry 1y agoHere's an open-source serving implementation: https://lightning.ai/bhimrajyadav/studios/build-a-production-ready-tts-api-with-chatterbox-powered-by-litserve?section=all&view=public&query=chatterbox https://lightning.ai/bhimrajyadav/studios/build-a-production... Also, a deployable model: https://lightning.ai/bhimrajyadav/ai-hub/temp_01jwr0adpqf055yps7axvrkcr5?section=all&view=public https://lightning.ai/bhimrajyadav/ai-hub/temp_01jwr0adpqf055...
- ipsum2 1y agoYou failed to mention that this is an ad for the company you work at. Also, the links don't even work without signing up for some shitty service.
- andymcsherry 1y agoHey ipsum, sorry I could have mentioned that. We spend a ton of effort on open source and sharing our ML knowledge with the community. If you don't want to use our platform, the entire source code and a tutorial is there to run it on your own.
- andyferris 1y agoIt took me ages to understand what TTS means!
- andyferris 1y agoIn the spirit of being more constructive... https://github.com/resemble-ai/chatterbox/pull/156 https://github.com/resemble-ai/chatterbox/pull/156
- SV_BubbleTime 1y agoI don't like how for text to image/video it's T2V I2V, and reference video to video is V2V... Then when we get to text 2 it T all of a sudden.
- dragonwriter 1y agoTTS has been around as an initialism long before the current AI wave, the x2y pattern is newer. (You do see it around TTS, even though TTS itself hasn't become T2S; e.g., TTS toolchains often include a g2p—grapheme-to-phoneme—component.)
- racecar789 1y agoI’d sign up for a service that calls a pharmacy on my behalf to refill prescriptions. In certain situations, pharmacies will not list prescriptions on their websites, even though they have the prescriptions on file, which forces the customer to call by phone — a frustrating process. I do feel bad for pharmacists, their job is challenging in so many ways.
- jeroenhd 1y agoDidn't Google already demo that with Google Duplex? It's not available here so I can't test it, but I think that's exactly the kind of thing duplex was designed to do. Although, from a risk avoidance point of view, I'd understand if Google wanted to stay as far away from having AI deal with medication as possible. Who knows what it'll do when it starts concocting new information while ordering medicine.
- MrThoughtful 1y agoHow do you set the voice? On the Huggingface demo, there seems to be no option for it. It has a female voice. Any way to set it to a male voice?
- ipsum2 1y agoIt's voice cloning. Maybe not available in the demo, but you just provide a different input.
- tevon 1y agoI just tested it out locally, really excellent quality, the server was easy to set up and well documented. I'd love to get to real-time generation if that's in the pipeline? Would like to use it along with Home Assistant.
- ipsum2 1y agoThe voice cloning is okay, not as good as Eleven Labs. There's a Rick (from Rick and Morty) voice example, and the generated audio sounds muffled and low quality. I appreciate that its open source though.
- audiala 1y agoWhat is the current state of the art for open source multilingual TTS? I have found Kokoro to be great as English as well, but am still searching for a good solution for French, Japanese, German...
- barrell 1y agoI’ve also been looking for this. OpenVoice2 supports a few languages (5 IIRC), but I haven’t seen anything usable yet
- pradeepodela 1y agoWhat is the latency?
- travisvn 1y agoChatterbox is fantastic. I created an API wrapper that also makes installation easier (Dockerized as well) https://github.com/travisvn/chatterbox-tts-api/ https://github.com/travisvn/chatterbox-tts-api/ Best voice cloning option available locally by far, in my experience.
- venusenvy47 1y agoWould this be usable on a PC without a GPU?
- travisvn 1y agoIt can definitely run on CPU — but I'm not sure if it can run on a machine without a GPU entirely. To be honest, it uses a decently large amount of resources. If you had a GPU, you could expect about 4-5 gb memory usage. And given the optimizations for tensors on GPUs, I'm not sure how well things would work "CPU only". If you try it, let me know. There are some "CPU" Docker builds in the repo you could look at for guidance. If you want free TTS without using local resources, you could try edge-tts https://github.com/travisvn/openai-edge-tts https://github.com/travisvn/openai-edge-tts
- mistersquid 1y ago> Chatterbox is fantastic. > I created an API wrapper that also makes installation easier (Dockerized as well) https://github.com/travisvn/chatterbox-tts-ap https://github.com/travisvn/chatterbox-tts-ap Gave your wrapper a try and, wow, I'm blown away by both Chatterbox TTS and your API wrapper. Excuse the rudimentary level of what follows. Was looking for a quick and dirty CLI incantation to specify a local text file instead of the inline `input` object, but couldn't figure it. Pointers much appreciated.
- travisvn 1y agoThis API wrapper was initially made to support a particular use case where someone's running, say, Open WebUI or AnythingLLM or some other local LLM frontend. A lot of these frontends have an option for using OpenAI's TTS API, and some of them allow you to specify the URL for that endpoint, allowing for "drop-in replacements" like this project. So the speech generation endpoint in the API is designed to fill that niche. However, its usage is pretty basic and there are curl statements in the README for testing your setup. Anyway, to get to your actual question, let me see if I can whip something up. I'll edit this comment with the command if I can swing it. In the meantime, can I assume your local text files are actual `.txt` files?
- internet_points 1y ago> Supported Lanugage > Currenlty only English. meh
- palmfacehn 1y agoHas anyone developed a way to annotate the input to provide emotional context? In the past I've used different samples from the same speaker for this.
- dragonwriter 1y agoThere are models that are trained for some kind of (in or out of band) emotiona (or style more general) prompting, but Chatterbox isn’t one of them, so beyond building some kind of system that took in input, processed it into chunks of text to speak and the settings Chatterbox does support (mostly pace and exaggeration) for each chunk, there’s probably no real way to do that with Chatterbox.
- 3ds 1y agoThere are only english voices, even in the paid version. Using them in other languages results in an accent.
- andrewstuart 1y agoThere’s been surprisingly little advancement in TTS after a rapid leap forward three years ago or so. There’s eleven labs which is quite good but not incredible and very expensive. Everything else ……. all the big AI companies …. have TTS systems that are kinda meh. Everything else in AI has advanced in leaps and bounds, TTS remains deep in the uncanny valley.
- _andrei_ 1y agovery cherry picked
- init0 1y agoChatterbox CLI https://pypi.org/project/voice-forge/ https://pypi.org/project/voice-forge/
- bachittle 1y agoI always have issues with TTS models that do not allow you to send large chunks of text. Seems this one does not resolve this either. Always has a limit of like 2-3 sentences.
- travisvn 1y agoThat's just for their demo. If you want to run it without size limits, here's an open-source API wrapper that fixes some of the main headaches with the main repo https://github.com/travisvn/chatterbox-tts-api/ https://github.com/travisvn/chatterbox-tts-api/
- b0a04gl 1y ago[dead]
- SV_BubbleTime 1y agoFun stuff... I don't know how or why, but connecting bluetooth while on this site, made all of the audio clips play at once (Firefox, Linux). Not the best listening experience.
- lukeinator42 1y agoDoes anyone know of an open-source TTS like this that can also encode speech to do voice conversion alongside TTS? i.e. a model that would take speech as input and convert it to one of the pretrained TTS voices.
- yavorgiv 1y agoCheckout https://github.com/playht/PlayDiffusion https://github.com/playht/PlayDiffusion
- monksy 1y agoHow would I install this alongside librechat or ollama using docker?
- ojw0816 1y agoLooks good! What is the difference between the open-source version and the priced version?
- DHolzer 1y agoI love chatterbox, it's my favourite. While the generation speed is quick, i wonder what performance optimization i could try on my 3090 to improve throughput. It's not quite enough for realtime.
- ash1224 1y agowow! 200mms very good!