4 ms·
You can get pretty close with open source software: https://claudio.uk/posts/audiblez-v4.html https://claudio.uk/posts/audiblez-v4.html
by csantini 2y ago
You can get pretty close with open source software:
https://claudio.uk/posts/audiblez-v4.html https://claudio.uk/posts/audiblez-v4.html
- rapind 2y agoOh wow. Thanks for posting! Samples sound great (on par with eleven by my untrained ear). Will definitely use this.
- neom 2y agoHow does it hold up on long stuff? I use Elevenlabs Studio daily and once things start to get into the chapters long, the voice can really start to go off the rails. It'd say they've solved a lot of this over the past 2/3 months, but it does still happen on long stuff.
- masteruvpuppetz 2y ago>> the voice can really start to go off the rails. Do you mean the AI gets tired?
- zaptrem 2y agoIn autoregressive models error accumulates over time. He likely means the voice starts to make odd sounds/gets lower quality. It would be really interesting if OP could share a clip of this phenomenon!
- neom 2y agoVarious different things can happen, it would take me quite some time to dig up examples but at least with elevenlabs you don't get the clicks and pops you get like on notebook LM for example. 11labs instability comes in the forms of intonation, pitch, accent, garbled words or even once language. I've only seen it happen in the 3k+ words gen's I've done, usually actually around the 75% point of the narration of whatever I've converted, and on average lasting a couple of seconds top.
- wrsh07 2y agoYeah - I've experienced this with eleven reader (I don't think you can gen text this long anymore using the reader app, lol) but switching voices fixed it for me I can go back and try to repro and get a recording....
- csantini 2y agoIt holds up well, because Audiblez uses sentence splitting (via Spacy models) before audio synthesis
- ultrasounder 2y agoBravo!
- simongray 2y agoOh no, it doesn't run on Apple Silicon. That's too bad.
- _joel 2y ago> On my M2 MacBook Pro, on CPU, it takes about 1 hour, at a rate of about 60 characters per second. Umm, it does.
- simongray 2y agoMy bad. I misread the official website: > We don't currently support Apple Silicon, as there is not yet a Kokoro implementation in MLX. As soon as it will be available, we will support it. I thought that meant that it didn't support Apple Silicon in general, but they were just talking about GPU support.
- fl0id 2y agothough they wouldn't need to use MLX, could also use pytorch etc
- stoobs 2y agoI think there's an issue somewhere in Kokoro though which means it doesn't actually take advantage of MPS, I did get a modified version up and running, but it was no faster than CPU, even though it passed all the internal tests using mps. I might try using F5-TTS-MLX instead actually (https://github.com/lucasnewman/f5-tts-mlx https://github.com/lucasnewman/f5-tts-mlx) and see how that does.
- csantini 2y agoIt works on Apple Silicon, but it doesn't use the GPU. Because Kokoro has not been implemented yet in MLX
- simongray 2y agoAh my bad! I just read the "We don't currently support Apple Silicon" on the official website, but I didn't realise that only pertains to GPU support.
- tonyhart7 2y agogood, now how I can use this on mobile??
- laurentlb 2y agoInteresting! This uses the Kokoro-82M model, which has a pretty good quality, but the set of languages is still quite limited.
- anonymous344 2y agodoes this run on linux machine also?
- nkmnz 2y agothird line on the page right below the first image says: > Audiblez 4.2 running on MacOSX via wxWidgets. Linux and Windows are supported too