2 ms·
You seem to be using a slightly tweaked CTC-based architecture built in tensorflow (possibly with Baidu's warp-ctc) but marketing it as some super-secret techno
by asrbash 9y ago
You seem to be using a slightly tweaked CTC-based architecture built in tensorflow (possibly with Baidu's warp-ctc) but marketing it as some super-secret technology you invented in-house. I don't see any performance benchmarks or WER results we can compare with other APIs, but the pricing is the same. Surely character-based approach lets you add new words without pronunciations, but that process is not as flawless as you make it seem, especially when you lack language model data for new words. Now I'm still a bit confused why somebody would use AssemblyAI over other APIs given the same price. And FYI you are not using Kaldi / Sphinx because the guys behind them did not endorse CTC and are purposefully avoiding putting it in there, though for example Kaldi's chain models are also sequence based. There was also Eesen that tried to implement CTC on top of Kaldi. Sorry if this came off too harsh, but I am a little suspicious about the novelty of the approach here.