5 ms·
You just said the models pretty much all work the same way, then you said doing what I described won't help. I'm confused. Apple and Google both offer real time
by coder543 4y ago
You just said the models pretty much all work the same way, then you said doing what I described won't help. I'm confused. Apple and Google both offer real time, on device transcription these days, so something clearly works. And if you say the models already all do this, then running it 30x as often isn't a problem anyways, since again... people are used to that.
I doubt people run online transcription for long periods of time on their phone very often, so the battery impact is irrelevant, and the model is ideally running (mostly) on a low power, high performance inference accelerator anyways, which is common to many SoCs these days.
- fxtentacle 4y agoI meant that most research that has been released in papers or code recently uses the same architecture. But all of those research papers use something different than Apple and Google. As for running the AI 30x, on current hardware that'll make it slower than realtime. Plus all of those 1GB+ models won't fit into a phone anyway.
- coder543 4y ago> Plus all of those 1GB+ models won't fit into a phone anyway. I don't think that's a requirement here. I've been playing with Whisper tonight, and even the tiny model drastically outperformed Siri dictation for me in my testing. YMMV, of course.