Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nicktikhonov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
nicktikhonov
7mo ago
From what I've seen, it's really easy to get PersonaPlex stuck in a death spiral - talking to itself, stuttering and descending deeper and deeper into total nonsense. Useless for any production use case. But I think this kind of e
2.
▲
by
nicktikhonov
7mo ago
Yep. Seems like caching more broadly is something worth exploring next if I were to do a pt2.
3.
▲
by
nicktikhonov
7mo ago
Yep. I've been learning Chinese for the past 3 months, so the name was a fold-in inspiration from my other hobby :)
4.
▲
by
nicktikhonov
7mo ago
Glad to hear! I built my blog on top of NextJS - it basically just renders .mdx files with contentlayer. One of the things I discovered is that you can easily vibe-code these explainer widgets. Seems like a perfect use case for vibe coding
5.
▲
by
nicktikhonov
7mo ago
You're probably right, at least at scale this could help
6.
▲
by
nicktikhonov
7mo ago
One thing you can get the LLM to do is to call a "skip turn" tool, which will basically trigger the system to wait without saying anything. Then all it will take is clever prompting to get the desired result.
7.
▲
by
nicktikhonov
7mo ago
I feel like you could get pretty far with a raspberry pi and microphone/speaker. I think the hard part is running a model that can detect a "Hey agent" on-device, so that it can run 24/7 and hand off to the orchestrator
8.
▲
by
nicktikhonov
7mo ago
This is fascinating, thanks for sharing! I wonder why amazon/google/apple didn't hop on the voice assistant/agent train in the last few years. All 3 have existing products with existing users and can pretty much define a
9.
▲
by
nicktikhonov
7mo ago
I'd say it was a collaboration. I had to hand-hold Claude quite a bit in the early stages, especially with architecture, and find the right services to get the outcome I wanted. But if you care most about where the code came from - it
10.
▲
by
nicktikhonov
7mo ago
100% - I thought about that shortly after writing this up. One way to make this work is to have a tiny, lower latency model generate that first reply out of a set of options, then aggressively cache TTS responses to get the latency super lo
11.
▲
by
nicktikhonov
7mo ago
Gross
12.
▲
by
nicktikhonov
7mo ago
A friend built this, everything working in-browser: https://ttslab.dev/voice-agent
13.
▲
by
nicktikhonov
7mo ago
Very cool! starred and on my reading list. Would love to chat and share notes, if you'd like
14.
▲
by
nicktikhonov
7mo ago
If you're of that opinion, you'll enjoy the new stuff coming out from nvidia: https://research.nvidia.com/labs/adlr/personaplex/
15.
▲
by
nicktikhonov
7mo ago
I'm sure LiveKit or similar would be best to use in production. I'm sure these libraries handle a lot of edge cases, or at least let you configure things quite well out of the box. Though maybe that argument will become less and l
16.
▲
by
nicktikhonov
7mo ago
I was using Twilio, and as far as I'm aware they handle any echos that may arise. I'm actually not sure where in the telephony stack this is handled, but I didn't see any issues or have to solve this problem myself luckily.
17.
▲
by
nicktikhonov
7mo ago
I didn't try Soniox, but I made a note to check it out! I chose Flux because I was already using Deepgram for STT and just happened to discover it when I was doing research. It would definitely be a good follow-up to try out all the di
18.
▲
by
nicktikhonov
7mo ago
If you read the post, you'll see that I used Deepgram's Flux. It also does endpointing and is a higher-level abstraction than VAD.
19.
▲
Show HN: I built a sub-500ms latency voice agent from scratch
(ntik.me)
570 points
by
nicktikhonov
7mo ago
|
153 comments
20.
▲
Show HN: BitClaw – A self-upgrading AI agent in 1,500 lines of code
(github.com)
4 points
by
nicktikhonov
7mo ago
|
0 comments
21.
▲
by
nicktikhonov
8mo ago
I'd do this for friends, but at scale this is unfortunately customs fraud
22.
▲
Building an AI voice agent from scratch
(ntik.me)
4 points
by
nicktikhonov
8mo ago
|
2 comments
23.
▲
by
nicktikhonov
8mo ago
I spent a day (~$100 in API credits) rebuilding the core orchestration loop of a real-time AI voice agent from scratch instead of using an all-in-one SDK. The hard part isn’t STT, LLMs, or TTS in isolation, but turn-taking: detecting when t
24.
▲
Pleasure of Learning
(supermemo.guru)
1 points
by
nicktikhonov
1y ago
|
0 comments
25.
▲
by
nicktikhonov
1y ago
might be possible to solve this with prompt configuration. e.g. you'd be able to explain to the llm all the weird naming conventions and unintuitive mappings
26.
▲
by
nicktikhonov
1y ago
and yet this was on the front page of hacker news for an entire day :D it's all about friction. why spend minutes writing a query when you can spend 5 seconds speaking the result you want and get 90-100% of the way there.
27.
▲
by
nicktikhonov
1y ago
Interesting lead. What else would they be looking for in a tool like this? My bad re the video, I'll make sure not to toggle dark mode in the next one.
28.
▲
by
nicktikhonov
1y ago
OpenAI LLM is used to generate SQL based on a combination of a user prompt and the database schema.
29.
▲
by
nicktikhonov
1y ago
Merged! Thanks Stephan
30.
▲
by
nicktikhonov
1y ago
I'm using my own production databases at the moment. But it might be quite nice to be able to generate complex databases with dummy data in order to test the prompts at the higher levels of complexity! And thank you for offering to con
More ›