Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kwindla
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
kwindla
2y ago
We've helped a number of Pipecat users hook into a variety of content moderation systems or use LLMs as judges. The most common approach is to use a `ParallelPipeline` to evaluate the output of the LLM at the same time as the TTS infer
32.
▲
by
kwindla
2y ago
There's a really nice implementation of phrase endpointing here: https://github.com/pipecat-ai/pipecat/blob/d378e699d23029e8ca7cea7fb675577becd5ebfb/src/pipecat/vad/vad_analyzer.py
33.
▲
by
kwindla
2y ago
Hi!
34.
▲
by
kwindla
2y ago
What? No. That’s crazy. (I believe you. I’ve just … never heard of giving up IP rights because you participated in a hackathon.) This is about community and building fun things. I can’t speak for all the sponsors, but what I want is to show
35.
▲
by
kwindla
2y ago
If you're interested in low-latency, multi-modal AI, Tavus is sponsoring a hackathon Oct 19th-20th in SF. (I'm helping to organize it.) There will also be a remote track for people who aren't in SF, so feel free to sign up wh
36.
▲
GStreamer and WebRTC HTTP Signalling
(arunraghavan.net)
2 points
by
kwindla
2y ago
|
0 comments
37.
▲
Show HN: A Voice AI Homage to Metal Gear Solid
(mgsx-daily-bots.vercel.app)
2 points
by
kwindla
2y ago
|
0 comments
38.
▲
by
kwindla
2y ago
"Generative" AI/ML is moving so fast in so many directions that keeping up is a challenge even if you're trying really hard to stay current! I'm part of a team building developer tools for real-time AI use cases (vo
39.
▲
by
kwindla
2y ago
You are right! Thank you. I went back and looked at actual benchmark numbers from a couple of years ago and the numbers I got were ~26ms one-way. I rounded up to 30 to be conservative, but then double-counted in the table above. Will fix in
40.
▲
by
kwindla
2y ago
Full technical write-up here: https://www.daily.co/blog/the-worlds-fastest-voice-bot/
41.
▲
by
kwindla
2y ago
Yes, Opus is the fastest and best option for real-time audio. It was designed to be flexible and to encode/decode at fairly low latencies. It sounds good for narrow-band (speech) at low bitrates but also works well at higher bitrates f
42.
▲
by
kwindla
2y ago
Voice models are getting both faster and more natural at a, well, a fast clip.
43.
▲
Show HN: Voice bots with 500ms response times
(fastvoiceagent.cerebrium.ai)
315 points
by
kwindla
2y ago
|
99 comments
44.
▲
by
kwindla
2y ago
I've been having this "you don't want to use TCP" conversation a lot lately with people who are building real-time voice + LLM applications. Almost everybody who hacks together a voice + LLM prototype starts with WebSock
45.
▲
by
kwindla
2y ago
RTSP is the control protocol. Some other protocol is needed for the actual audio/video streaming. That's usually RTP, these days. RTP is a core part of WebRTC, for example. When you're doing a video call in a web browser, you
46.
▲
by
kwindla
2y ago
I read arvix papers and all the other tabs I have open that I've been meaning to read). These days, that means I need to purchase WiFi unless I take time to specifically download PDFs. (Because <gesturing vaguely around at the compl
47.
▲
Swarm Parallelism: Training Large Models on Poorly Connected Devices
(arxiv.org)
2 points
by
kwindla
2y ago
|
0 comments
48.
▲
by
kwindla
2y ago
Adding to your list: https://vapi.ai -- really nice tools. (I try to keep up with all the different layers/players in this space.)
49.
▲
by
kwindla
2y ago
Yeah, seems to be a drop-in replacement for the existing inference APIs. But I haven't found any docs yet for streaming audio/video input.
50.
▲
by
kwindla
2y ago
Here's a translation demo in Pipecat using the now ancient and arthritic GPT-4 Turbo model. :-) https://github.com/pipecat-ai/pipecat/tree/main/examples/tra... As soon as GPT-4o audio input is
51.
▲
by
kwindla
2y ago
An audio-to-audio model is definitely a step forward. And I do think that's where things are going to go, generally speaking. For context relating to real-time voice AI: once you're down below ~800ms things are fast enough to feel
52.
▲
Show HN: An open source framework for voice assistants
(github.com)
346 points
by
kwindla
2y ago
|
39 comments
53.
▲
Co-Maintaining Needle in a Haystack
(truesparrow.com)
1 points
by
kwindla
3y ago
|
0 comments
54.
▲
How to Estimate Speed with Computer Vision
(blog.roboflow.com)
3 points
by
kwindla
3y ago
|
0 comments
55.
▲
Palisades Ski area closed Avalanche KT22 opening day
(old.reddit.com)
1 points
by
kwindla
3y ago
|
0 comments
56.
▲
by
kwindla
3y ago
They have not adequately funded the video product for a very long time. Other companies built out better infrastructure and a much broader range of features. The market also turned out to be smaller and differently structured than all of us
57.
▲
by
kwindla
3y ago
I have a dog in this fight, but strongly recommend that you do not move to [the] Zoom [Video SDK]. tldr: architecture that doesn't work for web apps, performance issues, missing features https://www.daily.co/blog/z
58.
▲
My role as a founder CTO: Year Six
(miguelcarranza.es)
6 points
by
kwindla
3y ago
|
0 comments
59.
▲
Doom, Dark Compute, and AI
(petewarden.com)
3 points
by
kwindla
3y ago
|
0 comments
60.
▲
by
kwindla
3y ago
The first Willa Cather I read was O Pioneers. Long after college. I thought I was reasonably well read. I knew the name, Willa Cather, of course, but assumed that since I hadn't read any Cather (hadn't been induced to read any Cat
More ›