Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ftreml
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
by
ftreml
7y ago
we used a context-free data set, augmented with additional domain-specific sound samples and it worked out fine, although the additiinal samples made nearly no difference
62.
▲
by
ftreml
7y ago
we used existing recipies for training (tuda german), slightly adapted and with additional training data. i am sure with more knowledge we would have made some extra percent ... performance is not really good, but good enough for our purpos
63.
▲
by
ftreml
7y ago
from my experience, when trained with the same data, kaldi is slightly better and with custom recipes adaptable to changing conditions. deepspeech has way better documentation and is more developer friendly. wav2letter seems to be the quick
64.
▲
by
ftreml
7y ago
with "performance" i meant the error rate, not the speed. speed was not a criteria for us, so we didnt evaluate it. i read that deepspeech is way quicker in training phase as it is smarter in using gpu computing power.
65.
▲
by
ftreml
7y ago
https://speech.botiumbox.com just a small server, hope that it wont crash when posting the link here
66.
▲
by
ftreml
7y ago
in our tests the performance was comparable, so we had no reason to switch from kaldi to something else (german data only). after all, what matters the most is the available training data. are there ready trained models available for wav2le
67.
▲
by
ftreml
7y ago
good point, will add this information. in short, german and english, because those are the languages i am comfortable with. i hope to find native speakers contributing more languages. deepspeech was evaluated, but right now kaldi provides b
68.
▲
by
ftreml
7y ago
yes absolutly! it should be a good mix of freely available packages, with meaningful default configuration.
69.
▲
by
ftreml
7y ago
it means: easy to install, easy to use, medium performance. no further know-how needed. compared to the total effort for selecting, training, deploying speech recogniction and speech synthesis it provides an extremly quick boilerplate to a
70.
▲
by
ftreml
7y ago
no, it was a special call center software called vacapo.
71.
▲
by
ftreml
7y ago
This project is the result of a one year long learning process in speech recognition and speech synthesis. The original task was to automate the testing of a voice-enabled IVR system. While we started with real audio recordings, very soon i
72.
▲
Show HN: Text-to-speech and speech-to-text open-source software stack
(github.com)
438 points
by
ftreml
7y ago
|
85 comments
73.
▲
by
ftreml
7y ago
i try to keep a healthy distance between private space and work space. not so easy. and may not apply to everyones lifestyle.
74.
▲
by
ftreml
7y ago
as i spend all of my energy to build a great product i dont have any energy left to convince co-founders or co-workers every few weeks. i wouldnt ever found a business or project with someone not as committed to it as myself. so: if you hav
75.
▲
by
ftreml
7y ago
Functional programming with functions as first-class-citizens. And related: asynchronous programming. Both with Node.js, a couple of years ago.
76.
▲
Show HN: Botium, the Selenium for Chatbots (Open Source)
(botium.at)
1 points
by
ftreml
7y ago
|
0 comments