13 ms·
Generate audiobooks from E-books with Kokoro-82M
- jaggs 2y agoThis looks really nice. And fast too it seems.
- treetalker 2y agoFor anyone looking for an easier alternative (and one without the bugs the author describes, such as skipping some prefaces or failing to detect some chapters), Voice Dream Reader on iOS (and macOS) handles .epub and other e-books just fine and supports a variety of built-in and external voices.
- huhtenberg 2y agoAnother subscription. $80/yr. Yaaaaaay.
- treetalker 2y agoUnless something has changed, the iOS version is a one-time purchase. I bought the app many years ago (8?) and have been a happy user since. Like you, though, I had that reaction to the subscription model for macOS and therefore decided not to "buy" it when it came out.
- huhtenberg 2y agoThey got greedy and decided to milk it. That's what changed. It's $80/yr for the iOS app.
- treetalker 2y agoOof, I believe they changed ownership so I must have been grandfathered in. That's steep.
- danman2 2y agoDo you know if it's possible to train it to use my own voice?
- rhizoma 2y agoYes, I’ve used Voice Dream for years with Pocket articles & ebooks because the Pocket app took up too much space and was limited to web articles. The voice quality is ok for short pieces or stints. The choice of voices is a bit robotic, but I find it useful while making written notes in Split View.
- freefaler 2y agoKybook is 1 time payment and can use iOS TTS voices.
- jdlyga 2y agoElevenLabs Reader is the same thing, but much higher quality voices for free. I've lost my place a few times so it's not quite as reliable as VoiceDream. But you aren't paying an expensive subscription with mediocre voices.
- ivan_icin 2y agoIt is free at the moment. They clearly specify in their terms that it won't be completely free in the future. The price of their same product if accessed from the web is $100/month for the amount (barely) sufficient for book reading. It may be smarter to skip until they reveal their official pricing for the mobile app.
- ivan_icin 2y agoThere are many apps. Voice Dream isn't up to date with voice quality (which was amazing 10 years ago when it started, but now even Apple gives you voices of similar quality for free), but is up to date with prices. Here is a detailed comparison chart I have made that tracks over 100 features across most popular apps: https://speechcentral.net/speech-central-vs-voice-dream-reader-vs-speechify/ https://speechcentral.net/speech-central-vs-voice-dream-read...
- qurashee 2y agoThis looks incredible! I’ve had an idea simmering in the back of my mind for a while now: creating an audiobook from an ebook for my commute using the voice of a specific audiobook narrator I really enjoy. The concept struck me after coming across the Infinite Conversation project here on HN. Unfortunately, I just haven’t found the time to bring it to life yet. :(
- vinni2 2y agoWhat about the copyright issue? You can’t mimic the voice of a narrator without their consent. OpenAI landed in trouble after using Scarlett Johansson’s voice in a demo. https://www.theverge.com/2024/5/20/24161253/scarlett-johansson-openai-altman-legal-action https://www.theverge.com/2024/5/20/24161253/scarlett-johanss...
- notachatbot123 2y agoNo limitations on this kind of thing if you are in private use.
- benatkin 2y agoShe only won in that OpenAI decided it wasn’t worth the trouble.
- K0balt 2y agoYeah, by my ear it was pretty clearly not SJ’s voice-likeness, although there were some superficial similarities. But some people could have mistook it due to some regional accent similarities, though it would be akin to interpretation of any light southern drawl with a similar timbre as being SJ.
- gunalx 2y agoKokoro seemed pretty nice for the size. I guess it is not much mvetter than a lot of the simpler tts. But at least it sounds less machinic than a few bad ones.
- outofpaper 2y agoIt is essentially a set of voice models building on https://huggingface.co/spaces/styletts2/styletts2 https://huggingface.co/spaces/styletts2/styletts2 The odd thing is that while they are releasing these great sounding models, they are not documenting the training process. What we want to know is what magic if any allowed them to create such wonderful voices...
- mg 2y agoWould this also be the best option if you just want to convert plain text files to audio?
- bArray 2y agoMarkdown and PDF would also be cool. I think it's just a case of feeding the TTS model the right data at the right time. The special sauce is in the model, there's really not much to the code: https://github.com/santinic/audiblez/blob/main/audiblez.py https://github.com/santinic/audiblez/blob/main/audiblez.py
- cess11 2y agoI would for sure not want this for fiction, it's too obvious that the voice has no understanding whatsoever of the text, but it's probably pretty nice for converting short news texts or notifications to audio.
- vanderZwan 2y agoYour point is a valid one, but I want to add to it that it is also a matter of expectations and how one listens. Years ago, when I was dating someone who spoke Russian as one of her native languages, we had to do a funny compromise when watching films together with her parents: they didn't speak a word of English, so we'd use the Russian dub with English subtitles. I noticed that the Russian dub was just one man reading a translation in a flat voice over what was happening on the screen, no attempts at voice acting or matching the emotions. Usually the dub would have a split second delay to the actual lines, so you'd still hear the original voices for a moment (and also a little bit in the background). At first I found it very jarring, but they explained that this flatness was a feature. You'll quickly learn to "filter out" the voice while still hearing the translation, and the faint presence of the original voices was enough to bring the emotional flavor back. The lack of voice acting helped with the filtering. This turned out to apply to me as well, even though I don't speak Russian! My brain subconsciously would filter out the dub, and extract most of the original performance through the subtitles and faint presence of the original voices. Obviously the original version would have been a better experience for me, but it was still very enjoyable. Of course a generated audiobook is not a dub, as there is no "original voice" to extract an emotional performance from. But some listeners might still be able do something similar. The lack of understanding in the generated voice and its predictable monotony might allow them to filter out everything but the literal text, and then fill it in with their own emotional interpretations. Still not as great as having proper story teller who does understand the text and knows how to deliver dramatic lines, but perhaps not as bad as expected either.
- cess11 2y agoIt's not a "point", I didn't make an argument. I dislike german and russian style dubs as well, I'd rather learn a bit of the original language.
- katspaugh 2y agoSounds better than many books on Audible.
- ekianjo 2y agojapanese is not supported yet despite the claims. you can easily realize that by running the examples provided.
- laserbeam 2y agoOn the one hand, this is very convenient. Probably cool for some non-fiction. On the other, some of my favorite audio books all stood out because the narrator was interpreting the text really well, for example by changing the pacing during chaotic moments. Or those audiobooks with multiple narrators and different voices for each character. Not to mention that sometimes the only cue you get for who's speaking during dialogue is how the voice actor changes their tone. I have mixed feelings about using this and losing some of that quality. I would totally use this over amateur ebooks or public domain audiobooks like the ones on project guttenberg. As cool as it is/was for someone to contribute to free books... as a listener it was always jarring to switch to a new chapter and hear a completely different voice and microphone quality for no reason.
- ahoka 2y agoI guess this is still very useful if you are blind.
- loktarogar 2y agoYeah, for accessibility purposes on things that aren't already narrated, this is kind of thing is huge.
- em-bee 2y agothat's the thing. it's not just for accessibility. anything not already narrated is a fair target for TTS. i don't have time to sit down and read books. all reading is done on the go, while getting around or doing daily routines at home. i have a small book that i am reading now, which should take a few hours to finish, but in the time i manage to get done reading it i will probably have listened to two or three audio books. oh, and it's also a boon for those who can't afford to buy audiobooks.
- vasco 2y agoYou don't choose to spend your time reading books. You probably roll your eyes when someone tells you they don't have time for some activity you deem valuable. This is the 'no time to exercise' debate in a different shape. They are also different activities, with audio it's easier to listen to more but retention is usually lower. Not casting any elitist "you need to read" bullshit by the way, but find it odd to define it in terms of lack of time, and I really like both mediums.
- Havoc 2y agoWow that sample sounds really good
- pprotas 2y agoI would love to have an e-reader that allows me to switch between text and audio at the press of a button. Imagine reading your book on the couch and then switching into audio mode while doing the dishes seamlessly, by connecting bluetooth headphones.
- InsideOutSanta 2y agoKindles used to provide this feature, but publishers and/or the Authors Guild stopped it, because audio rights and text rights are handled differently. In other words, when Amazon sells you a text book, it does not have the right to then also do TTS on that text and let you listen to it. There's some contemporary discussion of what happened here: https://tidbits.com/2009/03/02/why-the-kindle-2-should-speak-when-permitted-to/ https://tidbits.com/2009/03/02/why-the-kindle-2-should-speak... I think there is still integration with Audible, though. If you buy a book on the Kindle and on Audible, the position will sync, and you can switch between listening and reading without losing your place in the book.
- albert_e 2y agoYes the feature is called WhisperSync -- I used it many years ago and it was pretty good. I tried it while on a treadmill so it allowed me to follow the book with more focus without sacrificing much else.
- thfuran 2y agoIsn't whisper sync the current version that relies on owning both the ebook and audiobook?
- Brybry 2y agoI used that TTS feature semi-regularly on a Kindle 2. It wasn't a good experience but it was nice to be able to keep 'reading' a book while I was exercising. It worked for me for over a decade, until I broke the device. I don't know if I never updated the firmware or if the fact I used Calibre to convert books bypassed the feature gate.
- 2y ago
- mrklol 2y agoHow can this support more languages than the model itself?
- Kye 2y agoThe model might have stumbled on the generative AI equivalent of IPA.
- msoad 2y agoTo people who are experts in AI TTS: Why elevenlabs has such a lead in this space? It sounds better than OpenAI and Google models
- swores 2y agoCan anyone recommend an open source option that would allow training on a custom voice (my own, so I'd be able to record as many snippets as it needed to train on) to allow me to use it for TTS generation without sharing it off my machine? Edit: I'll wait to see if any recommendations get made here, if not I might give this one a go: https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS
- numpad0 2y agoI think you can probably generate TTS audio by classical means, and voice2voice that audio through RVC or Beatrice V2. Haven't looked into it in a while but Beatrice is apparently super fast and CPU only.
- phrotoma 2y agohttps://github.com/DrewThomasson/ebook2audiobook https://github.com/DrewThomasson/ebook2audiobook
- esskay 2y agoIf I recall Coqui is very much a dead project, just one to be aware of.
- hm64 2y agoCoqui is great, but in practice, I found Piper easier to set up, train, and deploy as an ONNX file. Big thanks to the Sherpa development team for their helpful resources: https://k2-fsa.github.io/sherpa/onnx/tts/piper.html https://k2-fsa.github.io/sherpa/onnx/tts/piper.html and to the Rhasspy team for their training guide: https://github.com/rhasspy/piper/blob/master/TRAINING.md https://github.com/rhasspy/piper/blob/master/TRAINING.md. I also found DEMUCS + Whisper + pydub to be a super helpful combo for creating quality datasets.
- drewbitt 2y agoThere is a fork here https://github.com/idiap/coqui-ai-TTS https://github.com/idiap/coqui-ai-TTS 'coqui-tts' Though according to the TTS leaderboard, Fish Speech https://github.com/fishaudio/fish-speech https://github.com/fishaudio/fish-speech and Kokoro are higher. https://huggingface.co/hexgrad/Kokoro-82M https://huggingface.co/hexgrad/Kokoro-82M https://huggingface.co/fishaudio/fish-speech-1.5 https://huggingface.co/fishaudio/fish-speech-1.5
- lc64 2y ago"was trained on <100 hours of audio" How the hell was it trained on that little data ?
- deleted 2y ago[deleted]
- Havoc 2y agoYeah that surprised me as well - seems low vs what is used on text llms . To be fair 100 hours of speaking is a lot of speaking though
- bbminner 2y agoI suppose it means per speaker. And it is based on a simplified style tts 2 which from my small dive into the subject seems one of the smaller models achieving great quality.
- vinni2 2y agoCan it also translate? I have family who would like audiobooks in German but most are in English only.
- october8140 2y agoAll these AI text to voice models seem to ignore emotion. It always sounds like a robot.
- lyu07282 2y agoLike with almost everything, its an active area of research: https://emosphere-tts.github.io/ https://emosphere-tts.github.io/ We are getting there
- boxed 2y agoSome of those samples sound like they are emoting in Korean while speaking English.
- lyu07282 2y agoTrue, maybe an artifact of the training data, here is another one: https://www.microsoft.com/en-us/research/project/emoctrl-tts/ https://www.microsoft.com/en-us/research/project/emoctrl-tts...
- croes 2y agoEmotion is the acting part of voice acting. Hard to copy with AI
- iagooar 2y agoI wonder if AI could create a "commentary" script that instructs the TTS how to read certain words or chapters. The commentary would be like an additional meta-track to help the TTS make the best reading. That should actually be possible to do already with existing tech. I haven't seen if you can instruct Kokoro to read in a certain way, does anyone know if this is possible?
- arafalov 2y agoTry this one https://www.hume.ai/ https://www.hume.ai/ - I found the demos (voice to voice) interesting.
- nottorp 2y agoWell there was some hope with ChatGPT that people will go back to being able to process text communication. Guess it was just a matter of time till someone figured out how to use "AI" to resume encouraging illiteracy.
- stavros 2y agoThere was some hope with the rise of equestrianism that people will go back to be able to shoe horses. Guess it was just a matter of time till someone figured out how to use "cars" to resume encouraging being unable to to a basic farrier job.
- nottorp 2y agoExcept cars were faster than horses, while audio or video content is much slower than reading.
- stavros 2y agoCars also have legs while audio doesn't, a point which is equally irrelevant. If people don't need to read, they don't need to read, and no matter how much a random Internet commenter wants them to need it, it won't change anything. Skills atrophy for a reason. It's fine to let them. You may as well be lamenting the lost art of long division.
- hombre_fatal 2y ago
- floppiplopp 2y agoIt sounds okay, but it lacks emotion and is monotone for fiction, it's the voice equivalent of the uncanny valley, which is probably fine if you don't really care.
- laserbeam 2y agoAnd when I don't care... to be honest I'm even OK with the dull browser TTS implementation when reading your average substack post. Shove the phone in my pocket, go shopping, get the jist of the article.
- yoavm 2y agoWas just looking for a TTS model to run locally for reading out loud articles, and never heard about Kokoro before! This looks great. I wonder if it can run in the browser somehow - could be a nice WebExtension.
- xkriva11 2y agoWhat about the WASM running sherpa-onnx? No intallation required and can be served locally as well. https://k2-fsa-web-assembly-tts-sherpa-onnx-en.static.hf.space/index.html https://k2-fsa-web-assembly-tts-sherpa-onnx-en.static.hf.spa...
- jiehong 2y agoI think most browsers support this already. Even maybe OS wide. I know it should work for Firefox on an article in reader mode. Or in MacOS you can select text and have it read out loud.
- yoavm 2y agoI'm using Firefox and I do not see this option. Probably not working on Linux?
- sriacha 2y agoYou might need to install/setup Speech Dispatcher. I was just using this implementation with Piper: https://github.com/Elleo/pied?tab=readme-ov-file https://github.com/Elleo/pied?tab=readme-ov-file. However easier way to read articles aloud is with Read Aloud extension: https://github.com/ken107/read-aloud https://github.com/ken107/read-aloud.
- yoavm 2y agoThat worked, thanks, though I find the speech quality in both options absolutely painful to listen to. The above WASM solution sounds about 100x better...
- albert_e 2y agoI hope a plugin for Calibre ebook management software comes along that makes it easier to convert select titles from your epub library to decent audio versions -- and a decent open source app for tablets and smartphones that can let us seamlessly consume both the ebook and audiobook at will.
- deleted 2y ago[deleted]
- Reimersholme 2y ago[dead]
- cwmoore 2y agoThe word “kokoro” means “heart” in Japanese, which I learned making the (heart shaped and paperback) puzzle books at https://www.kakurokokoro.com/ https://www.kakurokokoro.com/
- terhechte 2y agoIts also the name of the AI in Terminator Zero https://villains.fandom.com/wiki/Kokoro https://villains.fandom.com/wiki/Kokoro I'm not sure if that is related here.
- tkgally 2y agoNote that kokoro (心) means “heart” in the sense of “spirit,” “soul,” “mind,” “emotions,” etc. It doesn’t mean “heart” in the sense of “internal organ that pumps blood.” That is shinzō (心臓). I once heard an American friend with so-so Japanese ability ask a Japanese woman who had recently had a heart operation how her kokoro was doing, and she looked surprised and taken aback. Side note: After I started reading HN in 2019, I was struck by how many tech products mentioned here have Japanese names. I compiled a list for a few years and eventually posted it: https://news.ycombinator.com/item?id=31310370 https://news.ycombinator.com/item?id=31310370
- TypoAtLineZero 2y agoI am having a very similar setup locally, which uses Chrome with the 'Read Aloud' plugin. I am capturing the audio stream via QJackCtl/VLC. Voices, speed, pitch can be adjusted. Efficient and quickly set up
- deleted 2y ago[deleted]
- TheChaplain 2y agoFor accessibility I think this is a great thing, but as entertainment less so. Example is Hobbit and Lord of the Rings, the narrator Rob Inglis, makes an amazing voice performance giving depth to environments and characters. And of course the songs!
- basedrum 2y agoI want to be able to seemlessly read on my ebook reader and then put in my headphones and go for a walk with the dog and resume on audio where I left off. then when I come back, my ereader is at the right place where the audio finished and I can resume reading
- llamaimperative 2y agoReadwise Reader does this. A litttttle finicky at tracking read location but it’s workable
- GaggiX 2y agoThere is also this TTS: https://github.com/rhasspy/piper https://github.com/rhasspy/piper that is pretty good (depending on the language) and extremely fast, would be cool to change the script to user Piper instead of Kokoro in case you want to use a language that is not supported by Kokoro or it's too slow, Piper supports a lot of them.
- mikkom 2y agoWhat I really want and hope that someone does is to make an audiobook service that converts books to audiobooks but so that each character has own voice. Som audiobooks have this and I think it really makes the experience much more engaging. (Also maybe some background sound effects but not sure about that, some books also have this and it's quite nice too)
- ajsnigrutin 2y agoJust tried it, and "meh"... It's one step above "normal" text-to-speech solutions, but not much above it. The epub has "Chapter 1" as the title on the page, and a lot of whitespace, and then "This was...." (actual text). The software somehow managed to ignore all the whitespace and reach "chapter 1 this was.." as a single sentance, no pauses, no nothing. Blind? A great tool. Will it replace actual audiobooks? Well.. not yet at least.
- carlosjobim 2y agoWhy isn't the audiobook market strong enough that it would make business sense to pay good narrators and actors for each book published?
- DidYaWipe 2y agoIt is. But since when is "enough" enough for monopolistic/oligopolistic corporations?
- causi 2y agoI'm not able to try it until later, but regarding the sample audio: The voice quality is quite good, but what's going on with all the random pauses between words? It's very Captain Kirk.
- cliftonpowell 2y agoThere's another project called ebook2audiobook that has produces some decent results.
- woolion 2y agoIf you look for a lot of the great classics, audiobooks results are inundated with basic TTS "audiobooks" that are impossible to filter out. These are impossible to listen to because they lack the proper intonation marking the end of sentences, making it very tiring to parse. It might be better than tuna can sounding recordings, especially if you want to ear them in traffic (a common requirement), but that's about it. The alternative, if you want real quality recordings, is to stop reading classics and instead read latest Japanime Isekai of murder mystery, these have very good options on the market. Anyway, I don't think it needs more justification that it covers a good niche usage. I'm checking what the actual quality is (not a cherry-picked example), but: Started at: 13:20:04 Total characters: 264,081 Total words: 41548 Reading chapter 1 (197,687 characters)... That's 1h30 ago, there's no kind of progress notification of any kind, so I'm hoping it will finish sometime. It's using 100% of all available CPUs so it's quite a bother. (this is "tale of a tub" by Swift, it's about half of a typical novel length)
- csantini 2y agoYeah, that's a known issue, if the book is all on a single chapter you don't get any sense of progress. I may fix that next weekend
- woolion 2y agoIt's not in one Chapter, but Chapters are called "Section" (and so ignored!). It should be simple to have a dictionary of the different units that are used (I would assume "Part" would fail too, as would the hilarious "Catpter" of some cat-themed kid book, but that's more complicated I guess?). It did finish and result is basically as good as the provided example, so I'd say quite good! I'll plan to process some book before going to bed next time! Chapter 1 read in 6033.30 seconds (33 characters per second)
- callamdelaney 2y agoIt's insufferable.
- zoidb 2y agoNot directly related to the software, but interestingly on the authors website there is a Schedule a free call with me (https://claudio.uk/templates/call.html https://claudio.uk/templates/call.html). I wonder if randos on the internet ever do that, and how it works out.
- sam_lowry_ 2y agoHis LLM will answer the call.
- rpastuszak 2y agoI've been doing it for a few years (+200 calls) and have met a ton of wonderful people this way. https://untested.sonnet.io/notes/say-hi/ https://untested.sonnet.io/notes/say-hi/ https://sonnet.io/posts/hi https://sonnet.io/posts/hi
- herculity275 2y agoVery nice! I fiddled with this idea a few months back but the models available at the time were woefully slow on a macbook. Will definitely give this a spin, there's a large category of web serials and less popular translated novels that never get audiobook releases.
- delegate 2y agoThe quality is great (amazing even), but I can't listen to AI generated voices for more than 1 minute. I don't know why, I just don't like it. I immediately skip the video on youtube if the voice is AI generated. Might be because our brains try to 'feel' the speaker, the emotion, the pauses, the invisible smile, etc. No doubt models will improve and will be harder to identify as AI generated, but for now, as with diffusion images, I still notice it and react by just moving on..
- rockemsockem 2y agoThat kinda means the quality isn't great or amazing. Good TTS should be nearly or indistinguishable from a human speaker and should include emoting, natural pauses, etc
- xdennis 2y agoAmong other things, what I don't like is the hallucinated stress. Take the classic example of: > I never said she stole my money It can have 7 different meanings based on which word you stress out. The new AI voices sound very natural at a shallow level, but overall pronounce things in odd ways. Not quite wrong, but subtly unnatural which introduces some cognitive load. Old TTS systems with their monotonic voices are less confusing, but sound very robotic.
- DidYaWipe 2y agoerroneous or inappropriate ≠ hallucinated
- CMay 2y agoHaven't really been following the latest in TTS ML, but I expected this to be better or at least as good-bad as the stuff you hear on YouTube. Somehow it sounds worse. It really is jarring to listen to any of these ML voices and can't really stand it. Nope out of every video that uses them and can't tell if YouTube never recommends them to me for that reason, or just because the recommendations around what I watch are just so rarely going to be from some low reputation channel. Take a moment here for a second though and think about it. Even if these voices got to be really good, indistinguishable almost... would I want to listen to it even then? If it was an NPC's generated voice and generated dialogue in a game to help enrich the world building, maybe in that context. On YouTube or with newscasters? Probably not. Audio books? Think I would still rather have it be a real person, because it's like they're reading a story to me and it feels better if it's coming from someone. There's also the unknown factor, where if it's ML generated it's so sterile that the unknowns are kind of gone. Think about it like this, in the movie industry we had practical effects that were charming in a way. You could think about the physical things that had to occur to make that happen. Movie magic. Now, everything is so CG it's like the magic is gone. Even though you know people put serious hard work into it, there's a kind of inauthenticity and just lack of relevance to the real world that takes something away from it. It's like a real magician has interesting tricks, while an artificial magician is most likely just a liar. Still, I grant that it makes some cool things possible and there is potential if things are done right. Some positive mixture of real humans and machine generated stuff so it isn't devoid of anything connected to real life effort.
- maxglute 2y agoSounds really nice at 3x-4x speed, which I can't say for high quality TTS options last year. I'm wondering if there's metrics out there for audio speed vs clarity.
- monkeydust 2y agoI have been looking for something credible that can voice over written emails (long form ones), documents and powerpoints locally ...this might be just the thing!
- plumbees 2y agoAs a mandarin learner, I find that the Chinese one lacks cadence, which makes it very hard as a learner to comprehend. It's like a machine gun of words without the subtle slight pause between sets of words that I would normally lean on.
- jaggs 2y agoI really like this a lot. The default provides a really good audiobook feel, especially the Isabella voice. Any chance you could add in an API hook for optional ElevenLabs use?
- therealdrag0 2y agoDo folks have a preferred toolkit for extracting text from web articles? I’d like to TTS articles friends send me.
- Dowwie 2y ago2025 may be the year where we can generate a dramatic audiobook with ambient music, sound effects, and theatrical narration using neural networks. Many of the parts already exist.
- DidYaWipe 2y agoYes, because real narrator/actors are rolling in the dough. Let's kill one more profession with trash.
- bongodongobob 2y agoIf it's trash then why would it kill the profession?
- DidYaWipe 2y agoBecause people will opt for readily-available free trash instead of paying for high quality. And then that quality isn't available to anyone at any price, so everyone loses. If you haven't observed this in many other markets, you live an unusual (or unobservant) life.
- geor9e 2y agoThis one sounds a bit robotic and takes ~4 hours per book on my M1 laptop, so I'll keep looking. For now, I'm happy my current method - EPUBReader browser extension, which opens .epub as an HTML page in Microsoft Edge browser, which has a "Read Aloud" button set to the Stephan natural voice at 1.6 speed. Best sounding voice I've ever heard, speaks fast, clear, crisp, with natural inflections to the sentences, and if I want to jump to somewhere I just left click the text at that spot. And it's instant - no conversions. Downside is I have to stay in bluetooth range of my laptop, so I'm still looking for a good phone based method. Google Play Books works okay, but gets buggy at 1.6 speed.
- nickpsecurity 2y agoThe page says it was trained on under 100 hours of audio. Then, the link says “we employ large pre-trained SLMs, such as WavLM, as discriminators with our novel differentiable duration modeling for end-to-end training.” I don’t have time to read the paper to see what that means. Depending on what that means, it might be more accurate to say it was trained on 100 hours of audio and with the aid of another, pre-trained model. The reader who thinks “only 100 hours?!” will know to look at the pretraining requirements of the other model, too.
- leecarraher 2y agoin case you are wondering how audiblez becomes an executable in the PATH from a pip install audiblez per the documentation ... audiblez book.epub -l en-gb -v af_sky. it does not, instead it installs a python package with a cli interface, to run you then have to prepend python and load the module like this: python3 -m audiblez book.epub -l en-gb -v af_sky.
- skwee357 2y agoSoon, AI will flood the market with mediocre everything: books, audio books, art, movies, websites. The saddest thing is that people will still continue to participate in consuming these AI produced “goods”.
- abroadwin 2y agoIf there's one thing our capitalist society has taught me it's that people are always willing to endure a crappier product. I'm not sure we've found the bottom yet...
- vanderZwan 2y agoI think the saddest thing is that it's highly likely that real people will start to produce aesthetics that look/sound/etc like AI slop
- skwee357 2y agoTrue, and I think with the recent news around Sporify using AI to fill their playlists, we are already getting there. Just need to condition the public that there is no better
- flypunk 2y agoI really liked it and added a variable speed argument: https://github.com/santinic/audiblez/pull/4 https://github.com/santinic/audiblez/pull/4
- sysworld 2y agoFinally! Been trying all the TTS models popping up on here for ages, and they've all been pretty average, or not work on Mac, or only work on really short text, or be reeealy slow. But this one works pretty quick, is easy to install, has some passible voices. Finally I can start listening to those books that have no audio version. I'm a slow reader, so don't read many books. If a book doesn't have an audiobook version, chances are I won't read it. PS, I have used elevenlabs in the past for some small TTS projects, but for a full book, it's price prohibitive for personal use. (elevenlabs has some amazing voices) Thank you to the dev/s who worked on this!
- crorella 2y agoNice! It would be great to have per character voices
- boznz 2y agothis would be a game changer if done right. All good voice actors can carry a dozen different 'voices' for characters
- grwthckrmstr 2y agoThis is wonderful, and so happy to see the post where the author ran it locally on their Macbook. I am curious, is there an equivalent light model for speech to text, that can run real-time on the MacBook? I'm just playing around with AI models and was looking into this (a fully locally running app that lets you talk to your computer).
- physicsguy 2y agoI’m sure they sound more natural, but honestly, the text to speech built into my Kindle more than 10 years ago was good enough. Of course, Amazon killed that off because it would cannibalise sales to Audible.
- causality0 2y agoHas anyone gotten this working on windows? No matter where I put the files Powershell insists that kokoro-v0_19.onnx and voices.json aren't in the current folder.