10 ms·
Nerd-dictation, hackable speech to text on Linux
- abetusk 5y agoI've never even heard of VOSK-API [0], the underlying offline speech to text engine that this project uses. Does anyone have experience using it? Is it any good? [0] https://github.com/alphacep/vosk-api https://github.com/alphacep/vosk-api
- Arnavion 5y agoI use it to transcribe English robocalls. Vosk gets all the words right as long as I use the "Accurate generic US English" model. PocketSphinx (with the default en-us.lm.bin model in the distro package, no idea what it is) didn't get a single word right IIRC. I didn't try anything else.
- commoner 5y agoVosk powers Dicio, a free and open source voice assistant for Android. If you have an Android device, this app is another way to try out Vosk: - F-Droid: https://f-droid.org/packages/org.dicio.dicio_android/ https://f-droid.org/packages/org.dicio.dicio_android/ - Source: https://github.com/Stypox/dicio-android https://github.com/Stypox/dicio-android - HN: https://news.ycombinator.com/item?id=29762526 https://news.ycombinator.com/item?id=29762526 The accuracy of the English language recognition is not bad. I'm glad to see an implementation of Vosk for desktop Linux.
- flas9sd 5y agocan second Dicio to give Vosk a try. For a local model it worked surprisingly well. But you can't yet mix languages mid sentence - difficult when searching for restaurants that have english names but are not located in a english speaking country.
- Nimitz14 5y agoIt's very well known among ppl who know the field. It's quite good, the lead has a nice blog too.
- woodson 5y agoVosk-api isn't an SST engine itself, it is built using the Kaldi speech recognition toolkit (https://github.com/kaldi-asr/kaldi https://github.com/kaldi-asr/kaldi) and nicely implements and packages an API for Kaldi chain/LF-MMI models.
- foothebar 5y agoResults depend heavily on which speech files you use. You can even guess which it was, looking at the errors it makes.
- follower 5y agoYeah, I was really impressed with the project when I encountered it last year when trying out a bunch of FLOSS Speech-To-Text options. It was significantly better than the other FLOSS options I looked at--both in terms of getting it going initially & the quality of the speech to text results. I tested it with a lightly modified version of this example script: https://github.com/alphacep/vosk-api/blob/master/python/example/test_microphone.py https://github.com/alphacep/vosk-api/blob/master/python/exam... What I found particularly interesting was when you have the "partial" recognition output shown in real-time you get to see how--at the end of a sentence--it may change a word earlier in the sentence in the final recognition output based on (I guess) the additional context of the full sentence. (I just did a quick test again (with the installs from my testing last year) using an internal laptop microphone & the test script recognized a significant chunk of my speech (using a headset definitely improves things though) whereas with the same environment a test with `mic_vad_streaming` (from `DeepSpeech-examples-r0.9` with `deepspeech-0.9.0-models.pbmm`) failed to recognize any words at all.)
- zoomablemind 5y agoVosk, is it "wax" in Russian ("воск")? I think of wax recording rolls - old days CDs, aka Phonograph cylinder: https://en.m.wikipedia.org/wiki/Phonograph_cylinder https://en.m.wikipedia.org/wiki/Phonograph_cylinder
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- yjftsjthsd-h 5y agoThis is better than any other speech-to-text setup I've ever encountered, for one simple reason: I followed the dead-simple install steps in the readme, started the program, and it worked. Bonus points for the install being a git clone and pip install away. I don't know why this is a hard bar to clear, but bravo. (I suspect that it's because a lot of FOSS speech recognition is from academia where "follow the following 13 steps, including hand-crafting recognition parameters" is more normal and acceptable because everyone involved is already a domain expert, whereas I, as a user, just want "plug in a mic, run this thing, and get text on stdout".)
- melony 5y agoMost TTS and speech synthesis can be easy to install if you get rid of the GPU requirement. Both AMD and Nvidia have horrible workflows for installing their drivers and neural network/linear algebra kernels. Real time speech recognition/synthesis on generic consumer grade Intels/AMD cores is very, very difficult to do well which is why most providers are cloud based. (The alternative is targeting Mac only as they have standardized hardware everywhere)
- yjftsjthsd-h 5y agoYeah, text-to-speech is and has been easy for ages; I'm pretty sure I used espeak like a decade ago. On the other hand, I have tried... pretty much all the big names in speech-to-text, without success, or at best "I kind of got the demo to work but couldn't figure out how to do anything useful with it". Kaldi, sphinx, julius, a handful of tiny PoC things I found online... maybe I'm just bad at following instructions, or I'm trying to do something that they're not trying to optimize for, but I have not had a good time.
- johnisgood 5y ago> I followed the dead-simple install steps in the readme, started the program, and it worked. Bonus points for the install being a git clone and pip install away. I don't know why this is a hard bar to clear, but bravo. Woah, we really do write 2022.
- 2Gkashmiri 5y agoYou know... I have an idea. How about we use vosk and this tech to integrate with ffmpeg somehow so that peertube videos can get subtitles while being transcoded. Once we get English SRT, we could use libretranslate to translate that English SRT to multiple languages. This could be similar to what YouTube does with it's automatic subtitles. What do you guys say?
- harryvederci 5y agoSounds great! Do it!
- saYu 5y ago
- follower 5y agoThe package that this project is built on (vosk-api) mentions & includes some examples to demonstrate exactly that type of use case with ffmpeg: "...continuous large vocabulary transcription, zero-latency response with streaming API, reconfigurable vocabulary and speaker identification ... can also create subtitles for movies, transcription for lectures and interviews." * https://github.com/alphacep/vosk-api/blob/master/python/example/test_ffmpeg.py https://github.com/alphacep/vosk-api/blob/master/python/exam... * https://github.com/alphacep/vosk-api/blob/master/python/example/test_srt.py https://github.com/alphacep/vosk-api/blob/master/python/exam... * https://github.com/alphacep/vosk-api/blob/master/python/example/test_webvtt.py https://github.com/alphacep/vosk-api/blob/master/python/exam... (Edit: Also, thanks for introducing me to "libretranslate", looks like an interesting project.)
- 2Gkashmiri 5y agogreat. someone should link peertube github with this. i am sure the great people will do it much faster and more elegantly :-)
- follower 5y agoI just checked and apparently they are already aware: https://github.com/Chocobozzz/PeerTube/issues/3325#issuecomment-756237193 https://github.com/Chocobozzz/PeerTube/issues/3325#issuecomm... :) (Tho I'll admit I have no idea what "bluffing" means in that context. :D )
- suifbwish 5y agoVery cool. Does it have an erotic voice? Asking for a friend.
- commoner 5y agoNerd Dictation does speech-to-text (voice recognition), not text-to-speech. If you want to speak to your computer in an erotic voice, nobody's stopping you.
- jancsika 5y agoThere is actually an API for this: evStart: start speaking in an erotic voice evStop: stop speaking in an erotic voice evQuery: query whether you are speaking in an erotic voice evLinuxClassic: enable the inability to speak until Firefox is closed (experimental)
- dvh 5y agoFestival and install Scottish or French voice (whatever floats your boat)
- allanrbo 5y agoNice. Another notable mention in this space is Talon. Useful for automating all OS tasks with voice commands, as well as just dictation: https://talonvoice.com/ https://talonvoice.com/
- dotancohen 5y agoRunning it as a user prompts: > $ ./run.sh > [+] Prompting for admin to set up Tobii udev rule > [sudo] password for dotancohen: That does not build trust. I would prefer an instruction on how to set up a udev rule, or better yet, I would prefer that requirement to be relaxed. What does it need more than standard microphone access that e.g. nerd-dictation or even Telegram need?
- yencabulator 5y agoTalon has an EULA that is enough to send me scrambling for the hills: https://talonvoice.com/EULA.txt https://talonvoice.com/EULA.txt Meanwhile, the meat of this speech recognizition is Vosk, which is just Apache-2: https://github.com/alphacep/vosk-api/blob/master/COPYING https://github.com/alphacep/vosk-api/blob/master/COPYING
- kristopolous 5y agoI'm throwing another hat in the ring as this technology totally working most of the time. I used it to write this comment. This should make my life a lot easier because I find myself going to my phone and using the dictation feature a lot recently. It's not as good as the one on my android, but it's 95% of the way there.
- ideasman42 5y agoInteresting are you using the full model? I found with a good microphone and the full 1 gigabyte language model that the quality is quite good compared to other people's phones I have used from time to time.
- dotancohen 5y agoHow did you add the punctuation and capitalization?
- kristopolous 5y agoThat was by hand but it's a small task.
- dotancohen 5y agoI'm imagining remapping some VIM shortcuts for even easier capitalizing words and adding common punctuation. This remap makes the "." key add a period after the previous word, capitalize the current word, then move on to the next word: :noremap . bea.<esc>w~w
- dotancohen 5y agoReading your comment with nerd-dictation returns this for me: > i'm throwing another hat in the ring as this technology totally working most > of the time i used it to write this comment this should make my life ah lot > easier because i find myself going to my phone and using third dictation > feature a lot recently it's not as good as the one on my android for it's > ninety five percent of the way fair For use with no training that looks great. I'm sure that as I learn to speak more clearly, your 95% estimate is achievable.
- sundarurfriend 5y agoI was wondering how well it dealt with accents, them I saw that the Vosk API page specifically mentions "English, Indian English, German, French, ..." :D I don't know the story behind "Indian English" specifically being listed as a separate language, but I'm glad to see it's supported.
- bruce343434 5y agoWell, I'm not Indian but I can see it being a separate dialect, much like American English. For instance, an Indian might say "I have a doubt" instead of "I have a question". And as you mentioned, there is an accent, just like with American English.
- dotancohen 5y agoAdding Indian English was definitely doing the needful.
- sundarurfriend 5y agoThat's true, but that's why I mentioned "specifically listed" in my comment - the same things apply to Australian English for eg., and possibly to Filipino English and South African English and many others. Indian English is the only _dialect_ mentioned in the list, the rest are languages, which made me curious. The reason is probably something pragmatic, like perhaps a large enough corpus was available for that specifically.
- zelphirkalt 5y agoHas anyone used this somehow inside Emacs or knows how to make Emacs take its output and put it into a buffer?
- dotancohen 5y agoJust open emacs. This program outputs like a keyboard. And, in English at least, it works really well. I cannot believe it.
- dotancohen 5y agoReading my comment back, right here in the browser: > just open a max this program output like a keyboard and in > english at least it works really well i cannot believe it The only problems that I see are: 1. Capitalization and punctuation. 2. Doesn't know what emacs is, so it got that wrong. A user-installed dictionary might help here. 3. "outputs" came out as "output". I just tried a few more times, and I got the same results. I suspect that like "emacs", the word "outputs" is not in the dictionary.
- ideasman42 5y agoRegarding 1) I have 2 key bindings, 1 that starts a sentence and another binding but doesn't. while punctuation remains an issue (comers brackets and question marks for example) - since I'm mostly using this to save typing longer passages of text having to manually deal with punctuation isn't all that much of a hassle. But I can understand anyone attempting to go completely hands free would need something to support entering literal characters and punctuation.
- dotancohen 5y agoHow do the keybound scripts know when to stop listening? After a certain delay? Does the sentence-script just capitalize the first word and add a period to the end? I'd love to see the scripts if you don't mind.
- phantom_oracle 5y agoThis is such an amazing technology for the many tech people who are having to deal with hand/finger/elbow issues after extensive usage for years on their keyboards. I was looking for this type of tech for at least 2 years and I am glad it now exists. FOSS is amazing!
- deknos 5y agois there an offline good program for text to speech for german,french,spanish,english? and no, festival and espeak are not what i would consider good. the at&t website with text to speech as audio file which were used in these anonymous publications are good, but not espeak. if i had sth like this for european (and russian and arab languages) as open source standalone, i would be happy :(
- follower 5y agoYes! The project is called Larynx, and it is amazing: https://github.com/rhasspy/larynx/ https://github.com/rhasspy/larynx/ I waxed lyrical about it recently in this thread about private alternatives to Alexa: https://news.ycombinator.com/item?id=29562526 https://news.ycombinator.com/item?id=29562526 I can only vouch for the quality/variety in English but it does note support for 50 voices over 9 languages, including all the first group of languages you mentioned, and also Russian. (I've "played" with all those languages to test them but can't really vouch for how a native speaker/listener might find it. :D ) It is miles ahead of any of the other Free/Open Source TTS solutions I've tried, including the ones you mentioned. (It's still synthesized speech but the output quality is so good and the project is still extremely early days.) And there's a range of options in accent & gender--which are in general sorely lacking in other FLOSS TTS options. (In terms of licensing, some voices are licensed more freely than others but the majority are without significant restriction.) I like Larynx so much that I've been working on an editor for it to assist in "auditioning" & recording speech in a narrative context, e.g. game/film pre-viz.
- deknos 5y agoThanks, i will look it up! Thank you! :)