5 ms·
Open source speech foundation model that runs locally on CPU in real-time
- neuphonic 1y agoWe’ve just released Neuphonic TTS Air, a lightweight open-source speech foundation model under Apache 2.0. The main idea: frontier-quality text-to-speech, but small enough to run in realtime on CPU. No GPUs, no cloud APIs, no rate limits. Why we built this: - Most speech models today live behind paid APIs → privacy tradeoffs, recurring costs, and external dependencies. - With Air, you get full control, privacy, and zero marginal cost. - It enables new use cases where running speech models on-device matters (edge compute, accessibility tools, offline apps). Repo: https://github.com/neuphonic/neutts-air https://github.com/neuphonic/neutts-air Would love feedback from HN on performance, applications, and contributions.
- makkes 1y ago"the model has a 2048 tokens limit. This includes the reference text/phones as well as the reference and generation audio tokens." https://github.com/neuphonic/neutts-air/issues/15#issuecomment-3368454690 https://github.com/neuphonic/neutts-air/issues/15#issuecomme... So "no rate limits", while true, is kind of setting different expectations.
- ranger_danger 1y agoHow does this compare to Piper? Appears to use a proprietary codec as well.
- neuphonic 1y agoPiper is a VAE model which is quite robotic. This is a speech language model, which sound quite realistic. You can listen to the model on this video => https://www.youtube.com/watch?v=YAB3hCtu5wE https://www.youtube.com/watch?v=YAB3hCtu5wE The codec is open source: https://huggingface.co/neuphonic/neucodec https://huggingface.co/neuphonic/neucodec
- ranger_danger 1y ago> The codec is open source: https://huggingface.co/neuphonic/neucodec https://huggingface.co/neuphonic/neucodec This says it was trained on proprietary data.
- yjftsjthsd-h 1y agoThen may I suggest that this should be edited? > Audio Codec: NeuCodec - our proprietary neural audio codec that achieves exceptional audio quality at low bitrates using a single codebook ( https://huggingface.co/neuphonic/neutts-air#model-details https://huggingface.co/neuphonic/neutts-air#model-details )
- snowfield 1y agoReally cool, all my experiments with local speech has been spotty. Keen to try this out. Can it run on a gpu? Would it be faster?
- leobg 1y agoSo basically kokoro, but VC backed? Both models use espeak, so it seems like it’s the same general approach. Demo sounds great.
- arkensaw 1y agoI tried it today, amazing quality but VERY slow. took nearly 2 mins to generate audio for a single sentence on an M1 mac
- arkensaw 1y agoFrom the examples it seems that every time you run it, it clones a voice from the samples folder. Is there any way to compile this so it doesn't happen every time? Sorry if I'm missing the point.