Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nshm
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
nshm
4y ago
Picovoice technology is really nice! I happily recommend it to clients looking for lightweight ASR. You might not understand but there is a huge amount of work behind this simple demo.
32.
▲
by
nshm
4y ago
There are Whisper TFLite ports with model 40Mb size and tflite itself is about 3Mb. So nowhere near a gig. https://github.com/usefulsensors/openai-whisper
33.
▲
by
nshm
4y ago
Other people reporting accuracy issues https://github.com/openai/whisper/discussions/657
34.
▲
by
nshm
4y ago
Hard to guess. We test on variety of datasets. New version has more hallucinations in the end of audio.
35.
▲
by
nshm
4y ago
Easily, you can ask questions about speech recognition system design. For example you can ask ChatGPT about upsampling vs downsampling, it gives correct answer. It also can correctly recommend you architecture for streaming speech recogniti
36.
▲
by
nshm
4y ago
I've just tested it quickly, large-v2 is actually slightly less accurate than v1. So pick the model carefully.
37.
▲
by
nshm
4y ago
Spanish and Italian are very easy to recognize due to very simple phonetic structure.
38.
▲
by
nshm
4y ago
Looks like they plugged GPT-4 AI into speech recognition research and now they are going to release huge updates every month.
39.
▲
by
nshm
4y ago
Seems related to energy prices in Europe. They try to diversify. Feels like a lot of German companies will move to US soon.
40.
▲
by
nshm
4y ago
Coral and Jetson are not very powerful actually, you can't run even medium model in realtime. https://github.com/openai/whisper/discussions/417 And there is latency issue too, you won't get response
41.
▲
by
nshm
4y ago
I never read the book but I certainly agree with the title. I hope you find it useful https://www.amazon.com/Better-Good-Machine-Than-Person/dp/19...
42.
▲
by
nshm
4y ago
Try Sepia https://sepia-framework.github.io , it is an open source assistant, works very good.
43.
▲
by
nshm
4y ago
DeepSpeeech is very old software. Vosk works just fine https://github.com/alphacep/vosk-api . People even run tiny Whisper on Pi, though they have to wait ages.
44.
▲
by
nshm
4y ago
There is also a strange story of speech developer leaving them a week ago https://community.rhasspy.org/t/rhasspy-is-joining-nabu-casa...
45.
▲
by
nshm
4y ago
And you didn't even include crazy people using Alpine Linux with musl.
46.
▲
by
nshm
4y ago
After many years of Python ABI pain we moved to python-cffi. Life become way easier.
47.
▲
by
nshm
4y ago
We did comparison of recent Vosk and Whisper models here: https://alphacephei.com/nsh/2022/10/22/whisper.html In general, Whisper is more accurate but much more resource heavy. Vosk runs on single core w
48.
▲
by
nshm
4y ago
Well, you can give them the link to The Song of Hiawatha, largely inspired by Kalevala https://www.youtube.com/watch?v=r9b_VTW4FSE
49.
▲
by
nshm
4y ago
You need a Raspberry Pi with respeaker microphone. Many assistants to install around - for example https://github.com/iamsrp/dexter and many more. Whisper is a bit slow on RPI and not working well for short phrases. It
50.
▲
by
nshm
4y ago
Browser models are too small, unlikely they recognize accurately. They are more for simple predefined phrase. You can probably try vosk-api on the desktop-grade machine. You need big models from https://alphacephei.com/vosk&
51.
▲
by
nshm
4y ago
You can use open source assistant instead like Dicio https://github.com/Stypox/dicio-android and configure it the way you like.
52.
▲
by
nshm
4y ago
If you interested in unix-like software design and not yet familiar with kaldi toolkit, you definitely need to check it https://kaldi-asr.org It extended Unix design with archives, control lists and matrices and enabled really f
53.
▲
by
nshm
4y ago
Coming weeks we will see voice assistants in every IDE and photo editor! Coolness!
54.
▲
by
nshm
4y ago
They could easily use offline transcription like Vosk
55.
▲
by
nshm
4y ago
Cool. Consider embedding speech recognition features too!
56.
▲
by
nshm
4y ago
The whole value of this model is in 680 000 hours of training data and to reuse this value you need large model, not smaller ones. Smaller versions just don't have enough capacity to represent training data properly.
57.
▲
by
nshm
4y ago
Try https://github.com/alphacep/vosk-api/blob/master/csharp/demo...
58.
▲
by
nshm
4y ago
You can apply public punctation model from Vosk on top of Kaldi output, you can also get speaker labels with existing open source software. On quick video transcription test this model is more accurate than AssemblyAI and Rev AI. It will be
59.
▲
by
nshm
4y ago
It is interesting how they compare with wav2vec2 instead of nemo conformer (which is more accurate) in Table 2.
60.
▲
by
nshm
4y ago
You properly mentioned timestamps. There are many other important properties of good ASR system like vocabulary adaptability (if you can introduce new words) or streaming. Or confidences. Or latency of the output. Compared to Vosk models t
More ›