Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ftreml
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
ftreml
7y ago
the german model is from the kaldi tuda recipe with WER of 15%. the english is from the tedlium recipe with WER of 7%. room for improvement, but for our original purpose it was sufficient.
32.
▲
by
ftreml
7y ago
Few years ago, when building a chatbot for an Austrian telecommunication provider, we noticed that none of the available test automation frameworks was really helping us in testing and training. So we started to build something from scratch
33.
▲
Show HN: Open-Source Stack for Testing and Analytics of Conversational AI
(github.com)
1 points
by
ftreml
7y ago
|
1 comments
34.
▲
by
ftreml
7y ago
Here is the link to the Github repository: https://github.com/codeforequity-at/botium-core
35.
▲
by
ftreml
7y ago
Few years ago, when building a chatbot for an Austrian telecommunication provider, we noticed that none of the available test automation frameworks was really helping us in testing and training. So we started to build something from scratch
36.
▲
Show HN: Open-Source Stack for Testing and Analytics of Conversational AI
(chatbotsmagazine.com)
4 points
by
ftreml
7y ago
|
2 comments
37.
▲
by
ftreml
7y ago
so this is a real cool project. as soon as you finished your training I will be happy to add it as option (or maybe default setup) in the botium speech processing setup (if you want that). Do you have any experience with online decoding in
38.
▲
by
ftreml
7y ago
are there any wav2letter models ready for download ?
39.
▲
by
ftreml
7y ago
picotts is a command line tool. marytts is a client/server tool. (both included in botium speech processing and callable with curl). high quality with google cloud speech and amazon polly.
40.
▲
by
ftreml
7y ago
this seems like a joke ...
41.
▲
by
ftreml
7y ago
wonder if this works in real life without additional training phrases for an FAQ ?
42.
▲
by
ftreml
7y ago
I point you to this article: https://medium.com/@klintcho/creating-an-open-speech-recogni... It basically describes the thing you mentioned - matching freely available audio books with the source text and using some to
43.
▲
by
ftreml
7y ago
are you ready for a pull request ?
44.
▲
by
ftreml
7y ago
It's an API, of course it is hard to use without any real user interface ... but as an STT/TTS API, it won't get more easy than that ... Of course it is not a competitor to Google in any sense.
45.
▲
by
ftreml
7y ago
see here: https://speech.botiumbox.com
46.
▲
by
ftreml
7y ago
yes there are plenty of them. just google for something like "kaldi vs google". in short: not surprising the blockbuster cloud services provide better results as they have way more training data. tradeoff between price, privacy, q
47.
▲
by
ftreml
7y ago
40GB is maybe too much, but when building the docker images there is some space wasted. The image size after building is around 20GB (12GB marytts, 3GB kaldi de, 6GB kaldi en)
48.
▲
by
ftreml
7y ago
marytts supports a high number of voices in several languages. you could try to use another voice.
49.
▲
by
ftreml
7y ago
forgot to mention: when doing realtime parallel processing, the default configuration of this project is not a feasible setup. you have to run way more decoder workers, maybe distributed on various machines.
50.
▲
by
ftreml
7y ago
the project includes a websocket endpoint for realtime decoding. will add it to the docs. we are already using it for a callcenter with around 50 parallel audio streams.
51.
▲
by
ftreml
7y ago
just a guess: marytts is rather heavy weight
52.
▲
by
ftreml
7y ago
zamia-speech: asr training scripts for research purposes, several ready trained asr models for download, based on voxforge data. zamia-speech is the (very hard in terms of know-how, hardware and software requirements) training part to be do
53.
▲
by
ftreml
7y ago
with ssml formatting it is good enough for a customer facing ivr, though clearly recognizable as robotic voice - if this should be an issue i would not recommend it
54.
▲
by
ftreml
7y ago
currently included german and english. contributions for other languages welcome, native speakers will have better insights into quality of speech output and recognition model
55.
▲
by
ftreml
7y ago
for me it was exactly that, yes. we had powerful hardware and budget available so there was no reason to stick to statistical models. the available benchmarks showed that kaldi easily outperforms cmusphinx - when starting from scratch you t
56.
▲
by
ftreml
7y ago
very very interesting, didnt know this one. as soon as i am having some hours left will try to run some evaluation on this. after all, for my project only performance in german language counts
57.
▲
by
ftreml
7y ago
its always an issue with domain specific utterances. for the freely available data they are oov and have to be handled somehow (fe generate pronounciation with seqitur)
58.
▲
by
ftreml
7y ago
it requires hundreds of hours of speech data to make any difference, at least for a context-free asr. and gpu-powered high-end hardware, and several days for training. training an asr model is totally different requirement than using it as
59.
▲
by
ftreml
7y ago
speech recogniztion model filed are quite big
60.
▲
by
ftreml
7y ago
when using right ssml formatting, output quality is not so bad. but it cannot compete commercial engines, thats right. i guess thats for the same reason that google is dominating the speech recogniction world: they have tons of training dat
More ›