Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jilijeanlouis
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Show HN: Gladia CLI: transcribe audio from your terminal in one command
(github.com)
3 points
by
jilijeanlouis
3mo ago
|
0 comments
2.
▲
by
jilijeanlouis
3mo ago
Pretty cool !
3.
▲
by
jilijeanlouis
6mo ago
did you try gladia: ranking #1 on STT blind test https://compare-stt.com/
4.
▲
by
jilijeanlouis
6mo ago
same for gladia it's ranked top 1 in the STT blind tests: https://compare-stt.com/
5.
▲
by
jilijeanlouis
7mo ago
Author here. We built this because we kept seeing different word error rates (WER) for the same models depending on who was testing and how. Normalization rules ended up being a big reason why this was happening, so we decided to release a
6.
▲
Show HN: Reproducible open-source STT API benchmarks with full methodology
(github.com)
1 points
by
jilijeanlouis
7mo ago
|
1 comments
7.
▲
Real-time processing is the next frontier of audio transcription APIs
(techcrunch.com)
2 points
by
jilijeanlouis
2y ago
|
0 comments
8.
▲
by
jilijeanlouis
2y ago
Having worked with Whisper for quite some time now, it's true that hallucinations can be a real pain point. Long pauses between sentences / silence and background noise make it worse, among other factors. Good news is, there are w
9.
▲
by
jilijeanlouis
2y ago
This is really easy to do: it's just an embedding of your voice. So typically like 10/30 sec max of your voice to configure this. You already do a similar setup for faceId. I agree with you, I don't understand why they don&#x
10.
▲
by
jilijeanlouis
2y ago
Our API, Gladia, supports speaker diarization. We use a hybrid enterprise-grade ASR system for speech-to-text, with our own version of Whisper at its core, and state-of-the-art open source models for diarization. We process large audio file
11.
▲
by
jilijeanlouis
2y ago
Language detection in the presence of strong accents is, in my opinion, one of the most under-discussed biases in AI. Traditional ASR systems struggle when English (or any language) is spoken with a heavy accent, often confusing it with ano
12.
▲
by
jilijeanlouis
2y ago
Definitely twitter. This is where everything is announced and commented
13.
▲
by
jilijeanlouis
2y ago
There are actually a big opportunity for companies like loccus: https://www.loccus.ai/
14.
▲
by
jilijeanlouis
2y ago
Did you try other providers such as 11labs or open source like voice craft or openvoice
15.
▲
by
jilijeanlouis
2y ago
It fun that this question comes today, last night I was saying to myself that all audiobooks from audible sounded the same. But AWS TTS quality is bad so the closest would be play.ht, 11labs and lately open source voice craft and open voice
16.
▲
by
jilijeanlouis
2y ago
Despite regulation I don’t Believe it will actually get implemented even governments have regulations around audio accessibility for education in particular, for years now, and still nothing moves.
17.
▲
by
jilijeanlouis
2y ago
+1000
18.
▲
by
jilijeanlouis
2y ago
BTW I’ve heard from researchers from a big GAFAM (can’t name) that they actually use gpt to do this on the ground truth and spot that correction and do a second pass of human labelling on the ground truth to have better dataset.
19.
▲
by
jilijeanlouis
2y ago
Why would you use a dedicated app? Does it have to be natively embedded in android ?
20.
▲
by
jilijeanlouis
2y ago
That’s a bit problem with Arabic especially because of dialects and the lack of good datasets (when I say I mean robust - including noisy dirty audios).Would you mind testing gladia.io and give feedbacks ?
21.
▲
by
jilijeanlouis
2y ago
I think if they open source the model people will find ways to fork it to clone anyways.
22.
▲
by
jilijeanlouis
2y ago
Congrats that's pretty impressive !
23.
▲
by
jilijeanlouis
3y ago
I need it yes. that would be amazing tbh.
24.
▲
by
jilijeanlouis
3y ago
this is really cool probably the most wanted feature in LLM open source. It would be nice to have regex support too? do you plan to support regex?
25.
▲
Whisper-Zero, new ASR model designed to mitigate Whisper hallucinations
(gladia.io)
2 points
by
jilijeanlouis
3y ago
|
1 comments
26.
▲
by
jilijeanlouis
3y ago
French startup Gladia releaases an API,born out of complete rework of Whisper ASR that eliminates hallucinations and drastically improves accuracy. Built using over 1.5 million hours of audio, including phone and noisy data
27.
▲
by
jilijeanlouis
3y ago
Thanks for the feedbacks sebastien, to preserve original language you can select "automatic multiple languages" it will perform code switching.
28.
▲
by
jilijeanlouis
3y ago
Thanks for mentioning Gladia, this is not exactly how it works however, our version of Whisper is modified from the original one to avoid hallucinations, we are releasing a new model in a few days that is even better regarding this matter.
29.
▲
Don't Use Whisper Timestamps
(twitter.com)
3 points
by
jilijeanlouis
3y ago
|
0 comments
30.
▲
by
jilijeanlouis
3y ago
Depends on your partner - 3 cases: - Some are former operators / joined successful firms at very early stage and lived the growth and startup problems. Startups all have the same problems at different stages - Some are former founders
More ›