7 ms·
Show HN: Offline audiobook from any format with one CLI command
QuickPiperAudiobook locally generates an mp3 audiobook on Linux with one easy command. It can convert PDFs, epub, mobi, and many more by using ebook-convert. It uses any piper TTS model, and thus supports a wide variety of languages.
I've had great success using it to read more while reducing eye strain and computer usage. I think I've probably read 30 or so books this way now over the past year. Being able to listen to any content you want in audio form free and offline while going for a walk is extremely handy.
I hope it helps you as well!
Cheers
- alexkubica 2y agoA very cool project, you should build a website interface you could easily charge for it or take donations/advertise on it if you want to keep it free What would it take to add a specific language to piper? And do you know a good speech to text model?
- imurray 2y ago> And do you know a good speech to text model? OpenAI's whisper, code+model are available, and multiple projects have built on it. You could try this wrapper: https://github.com/m-bain/whisperX https://github.com/m-bain/whisperX -- or for short utterances on a smart-phone https://github.com/futo-org/whisper-acft https://github.com/futo-org/whisper-acft
- Bilal_io 2y agoDeepgram is another alternative. I use it at work, fastest service and also relatively cheap. But Whisper is better for selfhosting
- C-Loftus 2y agoThank you! Wouldn't a website interface then make it competing with and thus inferior to solutions like those from 11elevenlabs? I am not opposed to creating a SaaS offering, but I feel I do not have the economies of scale nor proprietary models a large company has. Let me know if I am wrong! Maybe I will one day do something as a separate project on the browser with WebGPU. With regards to adding languages, first check if support already exists [0]. Then there are a few tutorials that might be relevant [1] [2] [3]. Once you have the onnx model you can just put it in the QuickPiperAudiobook model directory and specify it via the cli args. [0] https://rhasspy.github.io/piper-samples/ https://rhasspy.github.io/piper-samples/ [1] https://github.com/rhasspy/piper/issues/51 https://github.com/rhasspy/piper/issues/51 [2] https://github.com/rhasspy/piper/blob/master/TRAINING.md https://github.com/rhasspy/piper/blob/master/TRAINING.md [3] https://www.youtube.com/watch?v=b_we_jma220 https://www.youtube.com/watch?v=b_we_jma220
- cbluth 2y agoVery nice, I will give it a try! I looked at piper a bit, does this support multi speaker models?
- C-Loftus 2y agoThanks! At the moment, I don't think there are any public multispeaker models I am aware of. I could be wrong though!
- tetrisgm 2y agoWould this dedrm my audible stuff?
- deleted 2y ago[deleted]
- Vuizur 2y agoNo, for this you need https://github.com/rmcrackan/Libation https://github.com/rmcrackan/Libation
- kranner 2y agoCan do it with plain old ffmpeg: https://news.ycombinator.com/item?id=23541424 https://news.ycombinator.com/item?id=23541424
- C-Loftus 2y agoIf you are looking to dedrm ebooks, that can be done via a calibre plugin. For audiobooks, I am not sure.
- dewey 2y agoThat's interesting, thanks for sharing. Does anyone know of a good solution for seamlessly switching between audiobooks and ebooks for books that are not bought from Amazon on Kindle? In this case you already have the input file, and the audio output file but I guess there would be an app that takes these two files to provide a good reading experience. As they are based on the same source it should be possible to keep the reading progress matched between them.
- noch 2y ago> Does anyone know of a good solution for seamlessly switching between audiobooks and ebooks for books that are not bought from Amazon on Kindle? Use Calibre's e-book viewer[^0] which uses Piper for text-to-speech. [^0]: https://manual.calibre-ebook.com/viewer.html#read-aloud https://manual.calibre-ebook.com/viewer.html#read-aloud
- dewey 2y agoThanks, but clarification: I meant on iOS / mobile devices as I'm not reading on my computer. On second thought, it would be an amazing feature for https://prologue.audio https://prologue.audio, which is a beautiful app and works very well for audiobooks already.
- SamBorick 2y agoReadEra Pro is an ereader app with a decent text-to-speech, I often flip between reading and listening.
- babs42 2y agoTry out Storyteller, they're working on this exact problem: https://smoores.gitlab.io/storyteller/ https://smoores.gitlab.io/storyteller/
- dewey 2y agoVery cool, thanks for sharing! I'll follow the project and hope there's some way to get this running on Kobo or other eInk readers in the future.
- tekkk 2y agoVery interesting. I have listened to an AI audiobook once and although the inflection was somewhat jarring at first you got kinda used to it. I suppose it's good enough for your own use. And audiobook prices being what they are rather affordable one as well.
- C-Loftus 2y agoYeah that is totally fair. In my experience, I feel that after a while your brain starts to tune out some of the inflection differences. Piper models are honestly pretty solid as well. I think in general, AI audiobook solutions like mine are better for non-fiction compared to fiction. Or at least that is what I read the most of
- senkora 2y agoI’ve really enjoyed moving most of my reading to TTS-generated audiobooks. I haven’t tried the newer AI voices but that certainly sounds like a step up!
- archargelod 2y agoCool app! I've had some issues with getting it to work, though: - ebook-convert is not a small dependency, it seems that it only comes bundled with calibre software. And calibre has huge number of python dependencies (>400 packages on OpenSuse) - don't know about you, but I'm not polluting my install with that for a small tool. So, I've grabbed appimage version of calibre, extracted it and added symlink to the bundled ebook-convert. It is still around ~500mb of wasted space, but atleast it's local to a single folder. Could you replace it with another tool/library, or include only necessary stuff with binary? - Then I've encountered another problem. I have no piper installed on my system, but readme says: > You don't need to have piper installed. This program manages piper and the associated models. It didn't download piper release and proceeded without errors. Then it did download some models. After that it errored out on trying to change directory to non-existent "~/.config/QuickPiperAudiobook/piper" So naturally, I looked in source code, found link to piper tarball and extracted it myself. A-ha! Now it works. Until.. - Done. Saved audiobook as /home/archargelod/Audiobooks/text.wav You could try to guess what was the problem, but I'm going tell you right away: it didn't create "Audiobooks" folder and again there were no errors. Thankfully, that was the last issue and after I created ~/Audiobooks manually, my generated wav was there.
- C-Loftus 2y agoThank you for the feedback and I'm sorry you had those issues. I cannot replicate at the moment on Ubuntu 24.04 but will check back on this. I presume it is something simple going wrong with how I am getting the home directory in golang and checking if the path exists. Your feedback on ebook-convert is very valid. I can take a look at breaking it up. (Granted I am not sure how much of a lift that would be)
- C-Loftus 2y agoIssues should be fixed now in the latest release.
- archargelod 2y agoCan confirm that everything is fixed. Thanks for the update!
- falcolas 2y agoAs a former audiobook narrator, may your cereal always be soggy and your socks too. On a more serious note, this is a cool application of the technological advancement in AI voice models, and inevitable in today's society. It just really sucks to watch this race to the bottom actively put people out of work. But hey, at least we can save a few bucks on an audiobook, right?
- fsckboy 2y ago>It just really sucks to watch this race to the bottom actively put people out of work the entire progress of civilization has depended on putting people out of work by increasing productivity and efficiency. Subsistence hunter-gatherers and subsistence farmers were put out of work by cheaper agriculture systems, and some of those unemployed realized they could support themselves by reading books to other people, a task they enjoyed much more.
- mistrial9 2y agothis broad-brush take seems so persuasive.. for about one minute of thinking.. systems of humans are built for humans first.. which work of which humans are being replaced and why? Is anyone actually driving? If the modern answer is "money answers all questions" then, who makes money simply by moving money? Anyone who is not moving money right now is fair game because money is the only decider ? this superficial thinking is full of holes from the first examination, and, actively harms others.. and is an excuse to ignore the statements of a audio book narrator here.
- falcolas 2y ago> Subsistence hunter-gatherers and subsistence farmers were put out of work by cheaper agriculture systems, and some of those unemployed realized they could support themselves by reading books to other people, The replacement of hunter-gatherers by farming is a change that took centuries to take hold. Nobody lost their ability to feed their family because their ability to hunt and gather was automated away. Ironically, the move away from hunter/gatherer subsistence took free time away (for things like storytelling) instead of adding to it, in exchange for greater reliability in their sustenance. The loss of entire swaths of employment is a fairly new development. As is the lack of safety nets (US Centric for obvious reasons) for those who become injured or otherwise unable to sustain themselves.
- OptionOfT 2y agoTangentially related: I like to leave my phone at home when I go exercise, and just listen to books via my watch (21st century problems...) But to this date I cannot use Apple's Books app on the watch to listen to audiobooks I have on mp3/mp4a/... It only works with audiobooks you have purchased in their walled garden.
- C-Loftus 2y agoYeah unfortunately I am not sure about the Apple Watch. On iOS I personally use BookPlayer [0] and find it easy to transfer mp3 files via USB. I think there are cloud sync options as well. Been very happy with that if you are looking for other mobile options. [0] https://apps.apple.com/us/app/bookplayer/id1138219998 https://apps.apple.com/us/app/bookplayer/id1138219998
- biomcgary 2y agoIs Piper currently the best open source TTS model? I occasionally review open models to see if they match elevenlabs and have been disappointed. However, Piper sounds better than the last time I listened around.
- TiredGuy 2y agoListening to the piper demos [1] and comparing to coqui [2], I'd say coqui sounds better to me, but I'd love to hear others' opinions. Looks like Piper's latest commits were 3 months ago [3] while Coqui's were 8 months ago [4], so they both seem similar in recency. In terms of ease of use though, especially with this project, personally Piper seems way less overwhelming. [1] https://rhasspy.github.io/piper-samples/ https://rhasspy.github.io/piper-samples/ [2] https://huggingface.co/spaces/coqui/xtts https://huggingface.co/spaces/coqui/xtts [3] https://github.com/rhasspy/piper https://github.com/rhasspy/piper [4] https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS
- lucius_verus 2y agoFor anyone who is interested, CoquiTTS (formerly, MozillaTTS) was great, but the project isn't maintained anymore (athough there's been some confusion about whether or not it's active. See: https://github.com/coqui-ai/TTS/issues/4022 https://github.com/coqui-ai/TTS/issues/4022). Looks like there's an effort to keep an actively maintained fork here, though: https://github.com/idiap/coqui-ai-TTS https://github.com/idiap/coqui-ai-TTS
- richerram 2y agoThis is awesome, it was pretty easy to set up and start using it. I have just one question/note to make: I tried a book in the Mexican Spanish language and noticed that it fails to catch the accents on the words (emphasis on words with tildes and strong accents on that syllable) and I am thinking it is because of the .pdf parsing since the Piper Voice Sample on their webpage example does it properly (on both avbailable voices). Do you have an idea of what could exactly be happening and how I can try to solve it? Thank you very much for the tool again!!! Update: Ohh ok I just checked the repo Issues and found the one about polish accents, I tried "--speak-diacritics" but got the same "Error: failed to read file passed as input to piper: read /tmp/ebook-convert-xxxxxxx.txt file already closed". If I skip the diacritics option it converts fine.
- richerram 2y agoUpdate 2: I went to look at the code and although I have never done anything with Go I was pleased with how easy it is to read plus your code was pretty well structured. I realized the removal of diacritics was happening at the function RemoveDiacritics inside lib/textProcessing.go on line 26 and modified the definition(?) to not modify special characters, compiled again and voila! It worked great. After that I used Calibre to convert a couple .pdfs to .txt and with a pretty simple python script got rid of page footnotes/headers/page_numbers and I just ended up with pretty decent Audiobooks. Thanks again for the great tool!