4 ms·
I wonder if fast enough for wakeword detection in WASM. Picovoice worked extremely well for this but it's proprietary.
by cjdell 3y ago
I wonder if fast enough for wakeword detection in WASM. Picovoice worked extremely well for this but it's proprietary.
- kkielhofner 3y agoThere's also OpenWakeWord[0]. The models are readily available in tflite and ONNX formats and are impressively "light" in terms of compute requirements and performance. It should be possible. [0] - https://github.com/dscripka/openWakeWord https://github.com/dscripka/openWakeWord
- regularfry 3y agoIt's probably still too big to be helpful with these model sizes, but if someone helpful runs the same training on `small.en` (and smaller) we might have something. Yes, this is me praying to the benevolent HN gods that someone will pick this up and run with it. I don't have a GPU anywhere close to capable...
- FL33TW00D 3y agoYou'd be surprised how capable old GPUs are! I've had great success with people running Whisper-Turbo in the browser on really old hardware: https://whisper-turbo.com/ https://whisper-turbo.com/
- kkielhofner 3y agoWe have benchmarks[0] for Willow Inference Server using Whisper + ctranslate2 + some of our own optimizations. TLD a six year old ~$100 used GTX 1070 is roughly 5x faster than a Threadripper PRO 5955WX at a fraction of the cost and power. [0] - https://heywillow.io/components/willow-inference-server/#benchmarks https://heywillow.io/components/willow-inference-server/#ben...
- RecycledEle 3y ago> TLD a six year old ~$100 used GTX 1070 is roughly 5x faster Did you mean TIL?
- regularfry 3y agoIt's not the inference, it's the training. They say in the paper: "We train with a batch size of 256 for a total of 80,000 optimisation steps, which amounts to eight epochs of training." That's a fair chunk of time. Mind you, `small.en` has smaller decoder layers than `medium.en`...
- srush 3y agoThe model targets the decoder part of the system which is the speed bottleneck. So for tasks like classification it is not likely to be helpful. However a similar method could be used for that use case. (Coauthor)
- bane 3y agoI would think that using any version of whisper for this use-case would be like digging a posthole in your front yard with an orbital directed energy cannon powered by a fusion reactor.
- kkielhofner 3y agoThis is a common viewpoint. Have you used Echo/Alexa and seen what people do with it? "Alexa make an entry on my calendar for lunch with Guillermo, Brian, and Kyle next week Wednesday at noon at Giordano's on Ohio street in Chicago". From 10-15 feet away, often with all kinds of noise, echo, who knows what. A child mumbling french can get within range of an Echo device and do this (with varying degrees of success). Yes a lot of that is handled on device in the audio frontend and elsewhere but it often still bleeds through and makes the fundamental speech recognition challenging. Not to mention bring your accent/voice/speech pattern. That's firmly Whisper territory and doesn't even get into the flexible grammar, integrations, etc with entire other stacks. Plus, many hundreds of millions of dollars and nearly a decade later Alexa still struggles with this.
- bane 3y agoGood response. However, wouldn't your described use-case be an activity that occurs after wakeword activation? Then handoff the rest of the audiostream to Whisper for transcription?
- kkielhofner 3y agoThanks! Yes, that's exactly what we do[0] (just like the commercial stuff). Wake word and VAD are low-resource and even an ESP chip can handle that + stream. The ESP-BOX-3 is actually our main target device for voice hardware interface. It's the nearly infinite audio, speech, grammar, language, etc variability and complexity where you need the "big guns". Another thing that seems to be getting lost on people - user expectations for voice interfaces are pretty high. If wake fails, a transcript is wrong, speech rec is slow, etc it's easier, faster, and far less frustrating to just take your phone out of your pocket. At that point why even have something poorly attempting to do voice? [0] - https://heywillow.io/how-willow-works/#willow-inference-server-mode-flow https://heywillow.io/how-willow-works/#willow-inference-serv...