3 ms·
There's actually two main reasons, related to each other: In order to do the kind of open domain recognition google (and others) do here, you need data. Lots
by ezy 15y ago
There's actually two main reasons, related to each other:
In order to do the kind of open domain recognition google (and others) do here, you need data. Lots of data. That implies that the models themselves are very large, and (generally) require more CPU to recognize with. You can't do that on a mobile device, mainly because of the space issue, but fancy new algorithms probably consume more CPU than even the latest mobile hw can manage[1]
The second, related issue is that, again, in order to do this kind of open domain recognition, you need to constantly improve the models[2]. Even a moderately sized set of models would be a pain in the rear to send back up to every android device using the system.
[1] You can typically do the front-end feature processing (and adaptation) on the device, and most vendor's solutions end up doing a bit of that.
[2] Even for speaker dependent modelling (dictation), you'd probably want to associate a key with the model and adapt that model in the cloud rather than sending model updates back to the device because of the space and transmit times.