Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
MbBrainz
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
MbBrainz
7mo ago
Love it! Solving the latency problem is essential to making voice ai usable and comfortable. Your point on VAD is interesting - hadn't thought about that.
2.
▲
Show HN: TTSLab – Text-to-speech that runs in the browser via WebGPU
(ttslab.dev)
3 points
by
MbBrainz
7mo ago
|
0 comments
3.
▲
by
MbBrainz
7mo ago
Amazing, Im happy that you like it! The voice agent indeed uses Silero VAD(v5) - There is an onnx wasm file available for the underlying VAD model, so we can perfectly run that with the onnxruntime-web, just like we run the TTS and STT mode
4.
▲
by
MbBrainz
7mo ago
Maker here. A few technical notes that might be interesting to this crowd: The Voice Agent chains three models in the browser: Whisper for STT → a local LLM → Kokoro/SpeechT5 for TTS. All inference runs on-device via WebGPU. The latenc
5.
▲
Show HN: TTSLab – A voice AI agent and TTS lab running in the browser via WebGPU
(ttslab.dev)
5 points
by
MbBrainz
7mo ago
|
3 comments