Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
java_beyb
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
java_beyb
2y ago
i haven't them recently but in the past I found smaller companies are much better at customization. try assemblyai, deepgram, picovoice or speechmatics. picovoice is on-device, you gotta fine-tune the model, but it's pretty easy a
2.
▲
by
java_beyb
3y ago
they have been around since pre chatgpt era, and not relevant to chatgpt. chatgpt understands text, coqui reads text. they're a text-to-speech / voice cloning/voice generation company. their founders were working at Mozilla w
3.
▲
by
java_beyb
3y ago
what's advanced speech recognition?
4.
▲
by
java_beyb
3y ago
have you tried it? i mean for fun, it wouldn't hurt for sure and ggerganov is doing amazing stuff. kudos to him. but whisper is designed to process audio files in 30-second batches if I'm not mistaken. it's been a while since
5.
▲
by
java_beyb
3y ago
well, deepgram might be the fastest among cloud-dependent APIs, like Speechmatics and Assembly AI mentioned above. -but- it cannot be faster than local or smaller models as you mentioned. Among local solutions, Whisper SDK doesn't sup
6.
▲
by
java_beyb
3y ago
how does it 4x cheaper than amazon or google? your basic plan cost per 1M char is $16, so are Google and Amazon.
7.
▲
by
java_beyb
3y ago
what you're looking for is called diarization. almost all enterprise STTs do that, you can find individual libraries on GitHub too. fine-tuning whisper is a nightmare, I don't know what the interviews are for, but again most enter
8.
▲
by
java_beyb
3y ago
edge brings compute close to where data is generated, cloud brings data to compute. even processing something in a web browser is called edge. i guess due to this impression the industry is moving towards "on-device"
9.
▲
by
java_beyb
3y ago
first, good initiative! thanks for sharing. i think you gotta be more diligent and careful with the problem statement. checking the weather in Sofia, Bulgaria requires cloud, current information. it's not "random speech". ESP
10.
▲
by
java_beyb
3y ago
if your decision is cost-oriented, then Whisper API is the cheapest - at least based on what other API companies promote on their websites. however, depending on what you're building, you may consider local speech-to-text by running sp
11.
▲
Ask HN: Has anybody used picovoice instead of krisp?
1 points
by
java_beyb
3y ago
|
0 comments