17 ms·
Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
Kitten TTS is an open-source series of tiny and expressive text-to-speech models for on-device applications. We are excited to launch a preview of our smallest model, which is less than 25 MB. This model has 15M parameters.
This release supports English text-to-speech applications in eight voices: four male and four female. The model is quantized to int8 + fp16, and it uses onnx for runtime. The model is designed to run literally anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required!
We're releasing this to give early users a sense of the latency and voices that will be available in our next release (hopefully next week). We'd love your feedback! Just FYI, this model is an early checkpoint trained on less than 10% of our total data.
We started working on this because existing expressive OSS models require big GPUs to run them on-device and the cloud alternatives are too expensive for high frequency use. We think there's a need for frontier open-source models that are tiny enough to run on edge devices!
- android521 1y agoit would be great if there is typescript support in the future
- divamgupta 1y agoYup it runs on the web browser. https://clowerweb.github.io/kitten-tts-web-demo/ https://clowerweb.github.io/kitten-tts-web-demo/
- Perz1val 1y agoIs the name a joke on "If the emperor had a tts device"? It's funny
- GaggiX 1y agohttps://huggingface.co/KittenML/kitten-tts-nano-0.1 https://huggingface.co/KittenML/kitten-tts-nano-0.1 https://github.com/KittenML/KittenTTS https://github.com/KittenML/KittenTTS This is the model and Github page, this blog post looks very much AI generated.
- nine_k 1y agoI hope this is the future. Offline, small ML models, running inference on ubiquitous, inexpensive hardware. Models that are easy to integrate into other things, into devices and apps, and even to drive from other models maybe.
- rohan_joshi 1y agoyeah totally. the quality of these tiny models are only going to go up.
- divamgupta 1y agoThat is our vision too!
- WhyNotHugo 1y agoDedicated single-purpose hardware with models would be even less energy-intensive. It's theoretically possible to design chips which run neural networks and alike using just resistors (rather than transistors). Such hardware is not general-purpose, and upgrading the model would not be possible, but there's plenty of use-cases where this is reasonable.
- amelius 1y agoBut resistors are, even in theory, heat dissipating devices. Unlike transistors, which can in theory be perfectly on or off (in both cases not dissipating heat).
- divamgupta 1y agoThe thing is that the new models keep coming every day. So it’s economically not feasible to make chips for a single model
- regularfry 1y agoIt's theoretically possible but physical "neurons" is a terrible idea. The number of connections between two layers of an FF net is the product of the number of weights in each, so routing makes every other problem a rounding error.
- mayli 1y agoIs this english only?
- g7r 1y agoYes. The FAQ says that multilingual capabilities are in the works.
- a2128 1y agoIf you're looking for other languages, Piper has been around in this scene for much longer and they have open-source training code and a lot of models (they're ~60MB instead of 25MB but whatever...) https://huggingface.co/rhasspy/piper-voices/tree/main https://huggingface.co/rhasspy/piper-voices/tree/main
- evgpbfhnr 1y agoI tried on some Japanese for the kicks of it, it reads... "Chinese letter chinese letter japanese letter chinese letter..." :D But yeah, if it's like any of the others we'll likely see a different "model" per language down the line based on the same techniques
- riedel 1y agoActually I found it irritating that the readme does not mention the language at all. I think it is not good practice to deduce it from the language of the readme itself. I would not like to have German language tts models with only a German readme...
- toisanji 1y agoWow, amazing and good work, I hope to see more amazing models running on CPUs!
- rohan_joshi 1y agothanks, we're going to release many more models in the future, that can run on just CPUs.
- onair4you 1y agoOkay, lots of details information and example code, great. But skimming through I didn’t see any audio samples to judge the quality?
- TheAceOfHearts 1y agoThey posted a demo on reddit[0]. It sounds amazing given the tiny size. [0] https://old.reddit.com/r/LocalLLaMA/comments/1mhyzp7/kitten_tts_sota_supertiny_tts_model_less_than_25/ https://old.reddit.com/r/LocalLLaMA/comments/1mhyzp7/kitten_...
- onair4you 1y agoThanks! Yeah. It definitely isn’t the absolute best in quality but it trounces the default TTS options on macOS (as third party developers are locked out of the Siri voices). And for less than the size of many modern web pages…
- blopker 1y agoWeb version: https://clowerweb.github.io/kitten-tts-web-demo/ https://clowerweb.github.io/kitten-tts-web-demo/ It sounds ok, but impressive for the size.
- nine_k 1y agoDoes anybody find it funny that sci-fi movies have to heavily distort "robot voices" to make them sound "convincingly robotic"? A robotic, explicitly non-natural voice would be perfectly acceptable, and even desirable, in many situations. I don't expect a smart toaster to talk like a BBC host; it'd be enough is the speech if easy to recognize.
- roywiggins 1y agoThis one is at least an interesting idea: https://genderlessvoice.com/ https://genderlessvoice.com/
- cosmojg 1y agoThe voice sounds great! I find it quite aesthetically pleasing, but it's far from genderless.
- a96 1y agoSo, what's the gender?
- degamad 1y agoInteresting concept, but why is that site filled with Top X blogspam?
- pbronez 1y agoThe YouTube video [1] was published in 2019. The Blog spam posts range from Nov 2022 to July 2023. Other than the video, the only relevant content is on the about page [2]. It says the voice is a collaboration between 5 different entities, including advocacy groups, marketing firms and a music producer. The video is the only example of the voice in use. There is no API, weights, SDK, etc. I suspect this was a one-off marketing stunt sponsored by Copenhagen pride before the pandemic. The initial reaction was strong enough that a couple years they were still getting a small but steady flow of traffic. One of the involved marketing firms decided to monetize the asset and defaced it with blog spam. [1] https://www.youtube.com/watch?v=lvv6zYOQqm0 https://www.youtube.com/watch?v=lvv6zYOQqm0 [2] https://genderlessvoice.com/about/ https://genderlessvoice.com/about/
- mlboss 1y agoReddit post with generated audio sample: https://www.reddit.com/r/LocalLLaMA/comments/1mhyzp7/kitten_tts_sota_supertiny_tts_model_less_than_25/ https://www.reddit.com/r/LocalLLaMA/comments/1mhyzp7/kitten_...
- tapper 1y agoSounds slow and like something from an anine
- ricardobeat 1y agoSpeech speed is always a tunable parameter and not something intrinsic to the model. The comparison to make is expressiveness and correct intonation for long sentences vs something like espeak. It actually sounds amazing for the size. The closest thing is probably KokoroTTS at 82M params and ~300MB.
- dvh 1y agoI think he meant overacting typical for English dubs.
- Telemakhos 1y agoThe voices sound artificial and a bit grating. The male voices especially are lacking, especially in depth: only the ultimate voice has any depth at all, while the others sound like teenagers who haven't finished puberty. None of the voices sound quite human, but they're all very annoying, and part of that is that they sound like they're acting.
- avisser 1y agoI heard a little DVa from Overwatch.
- numpad0 1y agoThe only real questions are which Chinese gacha game they ripped data from and whether they used Claude Code or Gemini CLI for Python code. I bet one can get a formant match from output this much overfit to whatever data. This isn't going to stay up for long.
- pkaye 1y agoWhere does the training data come for the models? Is there an openly available dataset the people use?
- wewewedxfgdf 1y agosay is only 193K on MacOS ls -lah /usr/bin/say -rwxr-xr-x 1 root wheel 193K 15 Nov 2024 /usr/bin/say Usage: M1-Mac-mini ~ % say "hello world this is the kitten TTS model speaking"
- wnoise 1y agoAnd what dynamic libraries s it linked to? And what other data are they pulling in?
- dented42 1y agoThat’s not a far comparison. Say just calls the speech synthesis APIs that have been around since at least Mac OS 8. That being said, the ‘classical’ (pre-AI) speech synthesisers are much smaller than kitten, so you’re not wrong per se, just for the wrong reason.
- deleted 1y ago[deleted]
- deathanatos 1y agoThe linked repository at the top-level here has several gigabytes of dependencies, too.
- satvikpendem 1y ago`say` sounds terrible compared to modern neural network based text to speech engines.
- wewewedxfgdf 1y agoSounds about the same as Kitten TTS.
- satvikpendem 1y agoTo me it sounds worse, especially on the construction of certain more complex sentences or words.
- RobKohr 1y agoWhat's a good one in reverse; speech to text?
- jasonjmcghee 1y agoWhisper and the many variants. Here's a good implementation. https://github.com/ggml-org/whisper.cpp https://github.com/ggml-org/whisper.cpp
- wenc 1y agoThis one is a whisper-based Python package https://github.com/primaprashant/hns https://github.com/primaprashant/hns
- wkat4242 1y agoHmm the quality is not so impressive. I'm looking for a really naturally sounding model. Not very happy with piper/kokoro, XTTS was a bit complex to set up. For STT whisper is really amazing. But I miss a good TTS. And I don't mind throwing GPU power at it. But anyway. this isn't it either, this sounds worse than kokoro.
- kenarsa 1y agoTry https://github.com/Picovoice/orca https://github.com/Picovoice/orca
- wkat4242 1y agoThanks!
- echelon 1y ago> Hmm the quality is not so impressive. [...] And I don't mind throwing GPU power at it. This isn't for you, then. You should evaluate quality here based on the fact you don't need a GPU. Back in the pre-Tacotron2 days, I was running slim TTS and vocoder models like GlowTTS and MelGAN on Digital Ocean droplets. No GPU to speak of. It cost next to nothing to run. Since then, the trend has been to scale up. We need more models to scale down. In the future we'll see small models living on-device. Embedded within toys and tools that don't need or want a network connection. Deployed with Raspberry Pi. Edge AI will be huge for robotics, toys and consumer products, and gaming (ie. world models).
- wkat4242 1y ago> This isn't for you, then. You should evaluate quality here based on the fact you don't need a GPU. I know but it was more of a general comment. A really good TTS just isn't around yes in the OSS sphere. I looked at some of the other suggestions here but they have too many quirks. Dia sounds great but messages must have certain lengths etc and it picks a random voice every time. I'd love to have something self hosted that's as good as openai.
- 1y ago
- andai 1y agoCan you run it in reverse for speech recognition?
- gromgull 1y agono, but whisper has a 39M model: https://github.com/openai/whisper https://github.com/openai/whisper
- divamgupta 1y agoWe will release an STT model as well.
- keyle 1y agoI don't mind so much the size in MB, the fact that it's pure CPU and the quality, what I do mind however is the latency. I hope it's fast. Aside: Are there any models for understanding voice to text, fully offline, without training? I will be very impressed when we will be able to have a conversation with an AI at a natural rate and not "probe, space, response"
- Teever 1y agoAny idea what factors play into latency in TTS models?
- divamgupta 1y agoMostly model size, and input size. Some models which use attention are O(N^2)
- blensor 1y ago"The brown fox jumps over the lazy dog.." Average duration per generation: 1.28 seconds Characters processed per second: 30.35 -- "Um" Average duration per generation: 0.22 seconds Characters processed per second: 9.23 -- "The brown fox jumps over the lazy dog.. The brown fox jumps over the lazy dog.." Average duration per generation: 2.25 seconds Characters processed per second: 35.04 -- processor : 0 vendor_id : AuthenticAMD cpu family : 25 model : 80 model name : AMD Ryzen 7 5800H with Radeon Graphics stepping : 0 microcode : 0xa50000c cpu MHz : 1397.397 cache size : 512 KB
- keyle 1y agoassuming most answers will be more than a sentence, 2.25 seconds is already long enough if you factor the token generation in between... and imagine with reasoning!... We're not there yet.
- moffkalast 1y agoHmm that actually seems extremely slow, Piper can crank out a sentence almost instantly on a Pi 4 which is a like a sloth compared to that Ryzen and the speech quality seems about the same at first glance. I suppose it would make sense if you want to include it on top of an LLM that's already occupying most of a GPU and this could run in the limited VRAM that's left.
- sandreas 1y agoCool. While I think this is indeed impressive and has a specific use case (e.g. in the embedded sector), I'm not totally convinced that the quality is good enough to replace bigger models. With fish-speech[1] and f5-tts[2] there are at least 2 open source models pushing the quality limits of offline text-to-speech. I tested F5-TTS with an old NVidia 1660 (6GB VRAM) and it worked ok-ish, so running it on a little more modern hardware will not cost you a fortune and produce MUCH higher quality with multi-language and zero-shot support. For Android there is SherpaTTS[3], which plays pretty well with most TTS Applications. 1: https://github.com/fishaudio/fish-speech https://github.com/fishaudio/fish-speech 2: https://github.com/SWivid/F5-TTS https://github.com/SWivid/F5-TTS 3: https://github.com/woheller69/ttsengine https://github.com/woheller69/ttsengine
- divamgupta 1y agoWe have released just a preview of the model. We hope to get the model much better in the future releases.
- nickpsecurity 1y agoFish Speech says its weights are for non-commercial use. Also, what are the two's VRAM requirents? This model has 15 million parameters which might run on low-power, sub-$100 computers with up-to-date software. Your hardware was an out-of-date 6GB GPU.
- maxloh 1y ago[dupe]
- girriPal 1y ago[dead]
- jainilprajapati 1y ago♥
- maxloh 1y agoHi. Will the training and fine-tuning code also be released? It would be great if the training data were released too!
- MutedEstate45 1y agoThe headline feature isn’t the 25 MB footprint alone. It’s that KittenTTS is Apache-2.0. That combo means you can embed a fully offline voice in Pi Zero-class hardware or even battery-powered toys without worrying about GPUs, cloud calls, or restrictive licenses. In one stroke it turns voice everywhere from a hardware/licensing problem into a packaging problem. Quality tweaks can come later; unlocking that deployment tier is the real game-changer.
- defanor 1y agoA Festival's English model, festvox-kallpc16k, is about 6 MB, and it is a large model; festvox-kallpc8k is about 3.5 MB. eSpeak NG's data files take about 12 MB (multi-lingual). I guess this one may generate more natural-sounding speech, but older or lower-end computers were capable of decent speech synthesis previously as well.
- Joel_Mckay 1y agoCustom voices could be added, but the speed was more important to some users. $ ls -lh /usr/bin/flite Listed as 27K last I checked. I recall some Blind users were able to decode Gordon 8-bit dialogue at speeds most people found incomprehensible. =3
- anthk 1y agoI'm not blind but spoken English it's far more difficult to grasp than written one (I'm a non-native speaker), and Flite runs on n270 netbooks at crazy speeds with really good enough voices.
- deleted 1y ago[deleted]
- rohan_joshi 1y agoyeah, we are super excited to build tiny ai models that are super high quality. local voice interfaces are inevitable and we want to power those in the future. btw, this model is just a preview, and the full release next week will be of much higher quality, along w another ~80M model ;)
- OfflineSergio 1y agoamazing! can't wait to integrate it into https://desktop.with.audio https://desktop.with.audio I'm already using KokorosTTS without a GPU. It works fairly well on Apple Silicon. Foundational tools like this open up the possiblity of one-time payment or even free tools.
- rohan_joshi 1y agowould love to see how that turns out. the full model release next week will be more expressive and higher quality than this one so we're excited to see you try that out.
- glietu 1y agoKudos guys!
- divamgupta 1y agoThanks
- wewewedxfgdf 1y agoChrome does TTS too. https://codepen.io/logicalmadboy/pen/RwpqMRV https://codepen.io/logicalmadboy/pen/RwpqMRV
- dang 1y agoMost of these comments were originally posted to a different thread (https://news.ycombinator.com/item?id=44806543 https://news.ycombinator.com/item?id=44806543). I've moved them hither because on HN we always prefer to give the project creators credit for their work. (it does however explain how many of these comments are older than the thread they are now children of)
- deleted 1y ago[deleted]
- evrennetwork 1y ago[dead]
- righthand 1y agoThe sample rate does more than change the quality.
- indigodaddy 1y agoCan coqui run in cpu only?
- palmfacehn 1y agoYes, XTTS2 has been reasonably performant for me and the cloning is acceptable.
- mg 1y agoGood TTS feels like it is something that should be natively built into every consumer device. So the user can decide if they want to read or listen to the text at hand. I'm surprised that phone manufacturers do not include good TTS models in their browser APIs for example. So that websites can build good audio interfaces. I for one would love to build a text editor that the user can use completely via audio. Text input might already be feasible via the "speak to type" feature, both Android and iOS offer. But there seems to be no good way to output spoken text without doing round-trips to a server and generate the audio there. The interface I would like would offer a way to talk to write and then commands like "Ok editor, read the last paragraph" or "Ok editor, delete the last sentence". It could be cool to do writing this way while walking. Just with a headset connected to a phone that sits in one's pocket.
- jiehong 1y agoOn Mac OS you can "speak" a text in almost every app, using built in voice (like the Siri voice or some older voices). All offline, and even from the terminal with "say".
- Fluorescence 1y agoI tried it a few months ago to narrate an epub in Apple Books and it was very broken in a weird way. It starts out decent but after a few pages, it starts slurring, skipping words, trailing off not finishing sentences and then goes silent. (I've just tried it again without seeing that issue within a few pages) > Siri voice or some older voices You can choose "Enhanced" and "Premium" versions of voices which are larger and sound nice and modern to me. The "Serena Premium" voice I was using is over 200Mb and far better that this Show HN. It's very natural but kind of ruined by diabolical pronunciation of anything slightly non-standard which sadly seems to cover everything I read e.g. people/place names, technical/scientific terms or any neologisms in scifi/fantasy. It's so wildly incomprehensible for e.g. Tibetan names in a mountaineering book, that you have to check the text. If the word being butchered is frequently repeated e.g. main character’s name, then it's just too painful to use.
- pjc50 1y ago
- deleted 1y ago[deleted]
- babycommando 1y agoSomeone please port this to ONNX so we don't need to do all this ass tooling
- victorbjorklund 1y agoIt is not the best TTS but it is freaking amazing it can be done by such a small model and it is good enough for so many use cases.
- rohan_joshi 1y agothanks, but keep in mind that this model is just a preview checkpoint that is only 10% trained. the full release next week will be of much higher quality and it will include a 15M model and an 80M model.
- khanan 1y ago"please join our DISCORD!"...
- klipklop 1y agoI tried it. Not bad for the size (of the model) and speed. Once you install all the massive number of libraries and things needed we are a far cry away from 25MB though. Cool project nonetheless.
- Dayshine 1y agoIt mentions ONNX, so I imagine an ONNX model is or will be available. ONNX runtime is a single library, with C#'s package being ~115MB compressed. Not tiny, but usually only a few lines to actually run and only a single dependency.
- divamgupta 1y agoWe will try to get rid of dependencies.
- wongarsu 1y agoThe repository already runs an ONNX model. But the onnx model doesn't get English text as input, it gets tokenized phonemes. The prepocessing for that is where most of the dependencies come from. Which is completely reasonable imho, but obviously comes with tradeoffs.
- pbronez 1y agoFor space sensitive applications like embedded systems, could you shift the preprocessing to compile time? You would need to constrain the vocabulary to see any benefits, but that could be reasonable. For example, you an enumeration of numbers, units and metric names could handle dynamic time, temperature and other dashboard items. For something more complex like offline navigation, you already need to store a map. You could store street names as tokens instead of text. Add a few turn commands, and you have offline spoken directions without on device pre-processing.
- WhyNotHugo 1y agoUsually pulling in lots of libraries helps develop/iterate faster. Then can be removed later once the whole thing starts to take shape.
- antisol 1y agoSystem Requirements Works literally everywhere Haha, on one of my machines my python version is too old, and the package/dependencies don't want to install. On another machie the python version is too new, and the package/dependencies don't want to install.
- divamgupta 1y agoWe are working to fix that. Thanks
- raybb 1y agoHave you considered offering a uvx command to run to get people going quickly?
- zelphirkalt 1y agoThough I think you would still need to have the Python build dependencies installed for that to work.
- pjc50 1y agoIf you restrict your dependencies to only those for which wheels are available, then uv should just be able to handle them for you.
- IshKebab 1y agoI think it can install Python itself too. Though I have had issues with that - especially with SSL certificate locations, which is one of Linux's other clusterfucks.
- pjc50 1y ago"Fixing python packaging" is somewhat harder than AGI.
- dlcarrier 1y ago
- countfeng 1y agoVery good model, thanks for the open source
- rohan_joshi 1y agothanks a lot, this model is just a preview checkpoint. the full release next week will be of much higher quality.
- tapper 1y agoI am blind and use NVDA with a sinth. How is this news? I don't get it! My sinth is called eloquence and is 4089KB
- mwcampbell 1y agoDoes your Eloquence installation include multiple languages? The one I have is only 1876 KB for US English only. And classic DECtalk is even smaller; I have here a version that's only 638 KB (again, US English only).
- killerstorm 1y agoI'm curious why smallish TTS models have metallic voice quality. The pronunciation sounds about right - i thought it's the hard part. And the model does it well. But voice timbre should be simpler to fix? Like, a simple FIR might improve it?
- codedokode 1y agoProbably "metallicity" is due to lack of details and cannot be fixed that easy.
- nickpsecurity 1y agoWe change our tone based on personal style, emotion, context, and other factors. An accurate generator might need to encode all that information in the model. It will be larger than a model that doesn't do all of that.
- dr_kiszonka 1y agoMicrosoft's and some of Google's TTS models make the simplest mistakes. For instance, they sometimes read "i.e." as "for example." This is a problem if you have low vision and use TTS for, say, proofreading your emails. Why does it happen? I'm genuinely curious.
- lynx97 1y agoWell, speech synthesizers are pretty much famous for speaking all sorts of things wrong. But what I find very concerning about LLM based TTS is that some of them cant really speak numbers greater then 100. They try, but fail a lot. At least tts-1-hd was pretty much doing this for almost every 3 or 4 digit number. Especially noticeable when it is supposed to read a year number.
- jpc0 1y agoNot entirely related but humans have the same problem. For scriptwriting when doing voice overs we always explicitly write out everything. So instead of 1 000 000 we would write one million or a million. This is a trivial example but if the number was 1 548 736 you will almost never be able to just read that off. However one million, five hundred and forty eight thousand, seven hundred and thirty six can just be read without parsing. Same with urls, W W W dot Google dot com.
- lynx97 1y agoRegarding humans, yes and no. If a human had constantly problems with 3 and 4 digit numbers like tts-1-hd does, I'd ask myself if they were neurodivergent in some way. And yes, I added instructions along the lines of what you describe to my prompt. Its just sad that we have to. After all, LLM TTS has solved a bunch of real problems, like switching languages in a text, or foreign words. The pronounciation is better then anything we ever had. But it fails to read short numbers. I feel like that small issue could probably have been solved by doing some fine tuning. But I actually dont really understand the tech for it, so...
- wongarsu 1y ago
- BenGosub 1y agoI wonder what would it take to extend it with a custom voice?
- junon 1y agoThis feels different. This feels like a genuinely monumental release. Holy cow. Very well done. The quality is excellent and the technical parameters are, simply, unbelievable. Makes me want to try to embed this on a board just to see if it's possible.
- ricardobeat 1y agoThe samples featured elsewhere seem to be from a larger model? After testing this locally, it still sounds quite mechanical, and fails catastrophically for simple phrases with numbers ("easy as 1-2-3"). If the 80M model can improve on this and keep the expressiveness seen in the reddit post, that looks promising.
- tecleandor 1y agoNot bad for the size (with my very limited knowledge of this field) ! In a couple tests, the "Male 2" voice sounds reasonable, but I've found it has problem with some groups of words, specially when played with little context. I think it's small sentences. For example, if you try to do just "Hey gang!", it will sound something like "Chay yang". But if you add an additional sentence after that, it will sound a bit different (but still weird).
- rishav_sharan 1y agoQuestion for the experts here; What would be a SOTA TTS that can run on an average laptop (32GB RAM, 4GB VRAM). I just want to attach a TTS to my SLM output, and get the highest possible voice quality/ human resembleness.
- kroaton 1y agoTry Unmute by Kyutai - https://unmute.sh/ https://unmute.sh/
- yahoozoo 1y agoIs there a paper describing the architecture of the model?
- zelphirkalt 1y agoWhat I am still looking for is a way to clone voice locally. I have OK hardware. For example I can use Mistral Small 3.1 or what it is called locally. Premade voices can be interesting too, but I am looking for custom voice. Perhaps by providing audio and the corresponding transcript to the model, training it, and then give it a new text and let it speak that.
- alexnewman 1y agoI'm so confused on how the model is actually made. It doesn't seem to be in the code or this stuff is way simpler than i thought. It seems to use a fancy library from japan, not sure how much it's just that
- anthk 1y agoAtom n270 running flite with a good voice -slt- vs this... would it be fast enough to play a MUD? Flite it's almost realtime fast...
- bashkiddie 1y agoTL;DR: If you are interested in TTS, you should explore alternatives I tried to use it... Its python venv has grown to 6 GBytes in size. The demo sentence > "This high quality TTS model works without a GPU" works, it takes 3s to render the audio. Audio sounds like a voice in a tin can. I tried to have a news article read aloud and failed with > [E:onnxruntime:, sequential_executor.cc:572 ExecuteKernel] Non-zero status code returned while running Expand node. Name:'/bert/Expand' > Status Message: invalid expand shape If you are interested in TTS, you should explore alternatives
- MrGilbert 1y agoA localized version of this, and I could finally build my tiny Amazon Echo replacement. I would love to see all speech synthesis performed on a local device.
- varenc 1y agoI'm doing this now with Home Assistant voice. All the TTS, STT, and LLMs involved run locally on my network. It's absurdly superior to every other voice assistant product. (Would be nice if it was just a pure multi-modal model though)
- binary132 1y agoI’m new to TTS models but is this something I can plug into my own engine like with LLMs, or does it require the Python stack it ships with?
- imprezagx2 1y agoBEAT THIS! Commodore C64 has the same feature called SAM - speaker synthesizer, speaks English and Polish. 48 kB of RAM BEAT THIS!
- a96 1y agoIt's not the same feature, but at least that's not several orders of magnitude away from "run anywhere"
- Tatiana343 1y ago[dead]
- spapas82 1y agoThis great for english, but is there something similar for other languages? Could this be trained somehow to support other languages?
- dirkc 1y agoHave you considered adding some 'rendered' examples of what the model sounds like? I'm curious, but right now I don't want to install the package and run some code.
- C-Loftus 1y agoAwesome work! Often times in the TTS space, human-similarity is given way too much emphasis at the expense of hurting user access. Frankly as long as a voice is clear and you listen to it for a while, the brain filters out most quirks you would perceive on the first pass. Hence why many blind folks still are perfectly fine using espeak-ng. The other properties like speed of generation and size make it worth it. I've been using a custom AI audiobook generation program [0] with piper for quite a while now and am very excited to look at integrating kitten. Historically piper has been the only good option for a free CPU-only local model so I am super happy to see more competition in the space. Easy installation is a big deal, since piper historically has had issues with that. (Hence why I had to add auto installation support in [0]) [0] https://github.com/C-Loftus/QuickPiperAudiobook https://github.com/C-Loftus/QuickPiperAudiobook
- thedangler 1y agoElixir folks. How would I use this with Elixir? I'm new to Elixir and could use this in about 15 days.
- bglusman 1y agoIt looks like it's Python, so it might be possible to use via https://github.com/livebook-dev/pythonx https://github.com/livebook-dev/pythonx ? But the parallel huggingface/bumblebee idea was also good, hadn't seen or thought of, that definitely works for a lot of other models, curious if you get working! Some chance I'll play with this myself in a few months, so feel free to report back here or DM me!
- bglusman 1y agoI just decided to try this quickly and hit some issues on my Mac FYI, it might work better on Linux but I hit a compilation issue with `curated-tokenizers`, possibly from a typo in setup.py or pyproject.toml in curated-tokenizers, spotted by AI: -Wno-sign-compare-Wno-strict-prototypes should be -Wno-sign-compare -Wno-strict-prototypes so could perhaps fix with a PR to curated-tokenizers or by forking it... Might well be other issues behind that, and unclear if need any other dependencies that kitten doesn't rely on directly like torch or torchaudio? but... not 5 mins easy, but looks like issues might be able to be worked through... For reference this is all I was trying basically: Mix.install([:pythonx]) Pythonx.uv_init(""" [project] name = "project" version = "0.0.0" requires-python = ">=3.8" dependencies = [ "kittentts @ https://github.com/KittenML/KittenTTS/releases/download/0.1/kittentts-0.1.0-py3-none-any.whl" ] """) to get the above error.
- dorian-graph 1y agoIt's not possible so far via Bumblebee, unfortunately[1]. [1] https://github.com/elixir-nx/bumblebee/issues/209 https://github.com/elixir-nx/bumblebee/issues/209
- akx 1y agoThis is a fun model for circuit-bending, because the voice style vectors are pretty small. For instance, try adding `np.random.shuffle(ref_s[0])` after the line `ref_s = self.voices[voice]`... EDIT: be careful with your system volume settings if you do this.
- the_arun 1y agoI like the direction we are heading. Build models that can run on CPUs & AI can become even more mainstream.
- butz 1y agoHow does one build similar model, but for different languages? I was under impression that being open source, there would be some instructions how to build everything on your own.
- peanut_merchant 1y agoI ran some quick benchmarks. Ubuntu 24, Razer Blade 16, Intel Core i9-14900HX Performance Results: Initial Latency: ~315ms for short text Audio Generation Speed (seconds of audio per second of processing): - Short text (12 chars): 3.35x realtime - Medium text (100 chars): 5.34x realtime - Long text (225 chars): 5.46x realtime - Very Long text (306 chars): 5.50x realtime Findings: - Model loads in ~710ms - Generates audio at ~5x realtime speed (excluding initial latency) - Performance is consistent across different voices (4.63x - 5.28x realtime)
- divamgupta 1y agoThanks for running the benchmarks. Currently the models are not optimized yet. We will optimize loading etc when we release an SDK meant for production :)
- don-bright 1y agoon my Intel(R) Celeron(R) N4020 CPU @ 1.10GHz it takes 6 seconds to import/load and text generation is roughly 1x realtime on various lengths of text.
- Jotalea 1y agothanks for testing on the same hardware as mine, before me.
- yunusabd 1y agoImpressive, might use this for https://hnup.date https://hnup.date
- theshrike79 1y agoLove the idea, but the text it produces is way too flowery for my taste "A new tool is stirring up excitement and debate in the programming community" Just give me the facts without American style embellishments. You're not trying to sell me anything =)
- yunusabd 1y agoSorry, just saw this. I absolutely agree, but it's really stubborn with the flowery language. I tried adding things like "DO NOT USE EMPTY PHRASES LIKE 'EVER-EVOLVING TECH LANDSCAPE'!!!!!" to the prompt, but it just can't resist. I want to give the whole system an overhaul, maybe newer models are better at this. Or maybe a second LLM pass to de-flowerize (lol) the language.
- mattfrommars 1y agoCan this work on intel npu unit?
- m00dy 1y agoI think one of the female voices belongs to Elizabeth Warren.
- gunalx 1y agoWould love to se something like this trained for multilingual purposes. It seems kinda like the same tier as piper, but a bit faster.
- 77pt77 1y agoHow does this compare to say piper-tts? I ask because their models are pretty small. Some sound awesome and there is no depdendency hell like I'm seeing here. Example: https://rhasspy.github.io/piper-samples/#en_US-ryan-high https://rhasspy.github.io/piper-samples/#en_US-ryan-high
- moomoo11 1y agoAre there any speech to text (opposite direction) that I can load on mobile app?
- system2 1y agoOne thing any GitHub project never has. A few-second demo.
- mrfakename 1y agoCool, it looks like this model is pretty similar to StyleTTS 2? Would it be possible to confirm?
- pjcodes 1y agoThis look pretty awesome. I will definitely give it a try and let you know the results
- 36OKY 1y ago[dead]
- marcobambini 1y agoIs there any way to get a .gguf version?
- alexwang123 1y agoThis is really great.
- monkfromearth 1y ago[dead]
- ghm2180 1y agoJust amazing
- OrangeMusic 1y agoIt's just so annoying and idiotic that there aren't a few samples on the home page. It didn't occur to you that it's the very first thing people would want to hear?
- Piraty 1y ago25M ? lol . the venv is 6.9G
- akrymski 1y agoNow if only we could get LLMs to this sort of size! I don't know much about how TTS works under the hood, why is it so much easier?
- csukuangfj 1y agoI have tested its speed on CPU and compared it with Piper, kokoro, and matcha. See https://github.com/KittenML/KittenTTS/issues/40 https://github.com/KittenML/KittenTTS/issues/40
- skyzouw 1y ago[flagged]
- skurtcastle 1y agoNot bad. Something I would not want to listen for long without more clarity. Could work very well for non-english speakers in various tools an such.
- jeffWrld 1y ago[dead]
- felarof 1y agoWe can integrate this into the browser directly! -- browserOS.com
- nullc 1y agoMight be useful to split it up to use the same speech features as some off the shelf vocoder, such as FARGAN used by RADE (https://freedv.org/radio-autoencoder/ https://freedv.org/radio-autoencoder/) and DRED (https://jmvalin.ca/papers/valin_dred_journal.pdf https://jmvalin.ca/papers/valin_dred_journal.pdf).
- oscar_zhou 1y agoIt looks great