Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
synesthesiam
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
synesthesiam
4y ago
The "new" CEO (Michael Lewis) has been there for a few years now, so this doesn't really affect Mycroft going forward. Source: I work for Mycroft AI.
32.
▲
by
synesthesiam
5y ago
I usually announce things on the Rhasspy voice assistant forums: https://community.rhasspy.org/ I also have a Twitter account (@rhasspy) used almost exclusively for announcements.
33.
▲
by
synesthesiam
5y ago
Due to licensing constraints, I wrote my own text frontend (MIT): https://github.com/rhasspy/gruut I plan to release a version of Larynx that uses eSpeak for phonemization, since it covers many of those corner cases.
34.
▲
by
synesthesiam
5y ago
Larynx author here, glad you're enjoying it! I'm implementing a new TTS model (VITS) which sounds better and is about 2x faster in practice. Should be ready in the next month or so :)
35.
▲
by
synesthesiam
5y ago
You're welcome! Glad to get one more person off the unnecessary part of the cloud.
36.
▲
by
synesthesiam
5y ago
Rhasspy author here. Thanks for the shout-out :)
37.
▲
by
synesthesiam
5y ago
Thanks! Try the --raw-stream option for listening to long texts: https://github.com/rhasspy/larynx#long-texts For speech-dispatcher, I'd start a Larynx HTTP server and use curl to get audio. I have an undocumented
38.
▲
by
synesthesiam
5y ago
The shape warnings don't seem to matter (something to do with the onnx runtime). Interactive mode needs sox installed or for you to specify a --play-command
39.
▲
by
synesthesiam
5y ago
Thanks! They seemed to work fine with en_us phonemes, so I haven't created a separate en_gb set yet. Maybe someday :)
40.
▲
by
synesthesiam
5y ago
Yes! There's a Docker image and Debian package for both 32-bit and 64-bit ARM. The 64-bit version is significantly faster (especially with low quality set).
41.
▲
by
synesthesiam
5y ago
I created Larynx ( https://github.com/rhasspy/larynx ) to address shortcomings I saw in Linux speech synthesis: * Licensing (MIT) * Quality (judge for yourself: https://rhasspy.github.io/larynx/ ) *
42.
▲
by
synesthesiam
5y ago
You might give Larynx a try: https://github.com/rhasspy/larynx Demo: https://youtu.be/hBmhDf8cl0k (I'm the author)
43.
▲
by
synesthesiam
5y ago
I think many of the available Kaldi/DeepSpeech models would pass, at least with the "Type-F Reproducibility". The pocketsphinx models would not, however, since they were trained on private datasets. My aim has been to train &
44.
▲
by
synesthesiam
5y ago
I'd recommend using Vosk directly for that: https://alphacephei.com/vosk/ voice2json is better suited for limited domain speech, where each sentence is a specific voice command (think home automation).
45.
▲
by
synesthesiam
5y ago
My templating language was inspired by JSGF, which seems to have informed the ABNF version of the W3C Speech Grammars. I don't support probabilities, though, since those are derived during the n-gram model generation. I would have pref
46.
▲
by
synesthesiam
5y ago
I haven't seen this yet, but I imagine it would involve running at least "voice2json record-command | voice2json transcribe-wav | jq .text". This will record a single command (until silence), and output the text transcription
47.
▲
by
synesthesiam
5y ago
I plan to add Vosk support soon. The goal of voice2json is to provide a common layer on top of existing open source engines. This common layer lets you train custom speech/intent models with having to know the details of each engine.
48.
▲
by
synesthesiam
5y ago
Agreed. I've at least added a "Recommended" option in the web UI that's language-specific. Part of the problem is that language support varies dramatically between components. There's usually a pretty obvious "
49.
▲
by
synesthesiam
5y ago
I need to cycle back and update voice2json. Rhasspy (the full voice assistant) supports DeepSpeech 0.9.3.
50.
▲
by
synesthesiam
5y ago
Author here. Thanks to everyone for checking out voice2json! The TLDR of this project is: a unified command-line interface to different offline speech recognition projects, with the ability to train your own grammar/intent recognizer i
51.
▲
by
synesthesiam
5y ago
If you're willing to record a public domain dataset, I'll help train a voice :)
52.
▲
by
synesthesiam
5y ago
Shameless plug for Rhasspy: https://rhasspy.readthedocs.io/en/latest/
53.
▲
by
synesthesiam
5y ago
Larynx TTS has a similar goal: https://rhasspy.github.io/larynx/ It was originally based on Mozilla TTS, but I've since moved to exporting models to Onnx for speed.
54.
▲
by
synesthesiam
6y ago
It supports open-ended transcription too: https://voice2json.org/commands.html#open-transcription Users have reported good accuracy with the English Deepspeech profile: https://github.com/synesthesiam/v
55.
▲
by
synesthesiam
6y ago
You may be interested in voice2json for offline batch processing: https://voice2json.org Here's an example using GNU parallel: http://voice2json.org/recipes.html#parallel-wav-recognition
56.
▲
by
synesthesiam
6y ago
> I found mycroft, Jarvis and a few others, but either got bogged down in dependencies or configuration. More recently, there is also Rhasspy ( https://rhasspy.readthedocs.io ) and voice2json ( https://voice2json.org
57.
▲
by
synesthesiam
7y ago
Rhasspy author here in case you have any questions :) If you're looking for something for the command-line, check out https://voice2json.org
58.
▲
by
synesthesiam
7y ago
https://rhasspy.readthedocs.io/en/latest/ Rhasspy is fully offline and uses grammars to increase accuracy just like you said.
59.
▲
by
synesthesiam
7y ago
We have, though my tests prior to 0.6 were not as promising as I'd hoped. With 0.6, though, we're planning to add support for English and French (a German model is apparently in the works: https://github.com/AASHIS
60.
▲
by
synesthesiam
7y ago
I haven't ever tried running this on Windows. You may get lucky with Docker, but audio input might be difficult. A workaround might be to stream audio in: https://rhasspy.readthedocs.io/en/latest/audio-input&#
More ›