3 ms·
I would assume they are only going to be doing the raw data collection and maybe cleanup and annotation, and that data should be made available so you can train
by blackkettle 9y ago
I would assume they are only going to be doing the raw data collection and maybe cleanup and annotation, and that data should be made available so you can train what you like.
If you poke around github and the Kaldi lists a bit more you can see that they are experimenting with and probably planning to use Kaldi.
I wonder what they plan to do for provisioning. It is one thing to collect data and train models, but quite another to make the service available over the web in an unlimited capacity. And we are not yet to the point where you can reasonably expect to run a high quality open-vocabulary STT system in your browser. The search network is typically in the GBs range.