Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lcolucci
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
lcolucci
9mo ago
Good question! Yes and to do this you'd want to use our "Self-Managed Pipeline": https://lemonslice.com/docs/self-managed/overview . You can combine any TTS, LLM and STT combination with LemonSlice a
32.
▲
Show HN: LemonSlice – Upgrade your voice agents to real-time video
133 points
by
lcolucci
9mo ago
|
133 comments
33.
▲
by
lcolucci
1y ago
Thank you! Very much agree that we need to improve speed to make the conversation more comfortable. Our target is <2sec latency (as measured by time to first byte). The other building blocks of the stack (like interruption handling, etc)
34.
▲
by
lcolucci
1y ago
thank you! No concrete paper plan yet as we're focused on shipping product features. anything specific you'd want to read about?
35.
▲
by
lcolucci
1y ago
haha one of the reasons launching on HN is great!
36.
▲
by
lcolucci
1y ago
thanks so much for the kind words! we agree that the leap to real-time feels huge. so excited to share this with you all
37.
▲
by
lcolucci
1y ago
thank you! We have an architecture diagram and some more details in the tech report here: https://lemonslice.com/live/technical-report And yes, exactly. In between each character interaction we need to do speech-to-tex
38.
▲
by
lcolucci
1y ago
haha this is amazing! Just made him a featured character. Folks can chat with him by searching for "Devil"
39.
▲
by
lcolucci
1y ago
We wouldn't build it ourselves, but there are several companies like Etched, Groq, and Cerebras working on purpose-built hardware for transformer models. Here's more: https://www.etched.com/announcing-etched
40.
▲
Show HN: Lemon Slice Live – Have a video call with a transformer model
195 points
by
lcolucci
1y ago
|
84 comments
41.
▲
by
lcolucci
2y ago
I use both. Sonauto sounds more "real" and varied than what I can get with suno
42.
▲
by
lcolucci
2y ago
The transition btw two songs demo is super cool! I often need to do this when editing videos but used to have no way to do it. Not to mention that now you can have playlists that transition seamlessly btw two songs. Low-cost party DJ?
43.
▲
by
lcolucci
2y ago
Yep for sure! EMO is a good one. VASA-1 (Microsoft) and Loopy Avatar (ByteDance) are two others from this year. And thanks!
44.
▲
by
lcolucci
2y ago
this made me laugh out loud
45.
▲
by
lcolucci
2y ago
This is the best compliment :) and also a good idea… could a trained lip reader understand what the videos are saying? Good benchmark!
46.
▲
by
lcolucci
2y ago
This is a good observation. Can you share the videos you’re seeing this with? For me, normal talking tends to work well even on long generations. But singing or expressive audio starts to devolve with more recursions (1 forward pass = 8 sec
47.
▲
by
lcolucci
2y ago
thank you! it's for sure an interesting time to be alive... can't complain about it being boring
48.
▲
by
lcolucci
2y ago
We are big fans of Hedra. Do you know if they've publicly commented on their model architecture? As far as we know, our particular choice of an end-to-end diffusion + transformer is novel. We don't know what Hedra is doing. It cou
49.
▲
by
lcolucci
2y ago
I didn't know about Persona Collective - very cool! I think the issues in your video are more related to the style of the image and the fact that she's looking sideways than the race. In our testing so far, it's done a pretty
50.
▲
by
lcolucci
2y ago
Very cool! If we release an API, you could use it across the different Ragdoll experiences you're creating. I agree personalized character experiences are going to be a huge thing. FYI we plan to allow users to save their own character
51.
▲
by
lcolucci
2y ago
We use more than one but ElevenLabs is a major one. The voice names in the dropdown menu ("Amelia", "George", etc) come from ElevenLabs
52.
▲
by
lcolucci
2y ago
Thank you! It's interesting you've noticed the last frame breakdown happening more with low-res images. This is a good hypothesis that we should look into. We've been trying to debug that issue
53.
▲
by
lcolucci
2y ago
I think you've made the 1st ever talking dog with our model! I didn't know it could do that
54.
▲
by
lcolucci
2y ago
Wow this worked so well! Sometimes with long hair and paintings, it separates part of the hair from the head but not here
55.
▲
by
lcolucci
2y ago
Woah that's a good find Andrew! That low-res video looks pretty good
56.
▲
by
lcolucci
2y ago
I'd say the 5 year ballpark is about right, but it'll involve combining a bunch of different models and tools together. I follow a lot of great AI filmmakers on Twitter. They typically make ~1min long videos using 3-8 different to
57.
▲
by
lcolucci
2y ago
Nice! Earlier checkpoints of our model would "gender swap" when you had a female face and male voice (or vice versa). It's more robust to that now, which is good, but we still need to improve the identity preservation
58.
▲
by
lcolucci
2y ago
that's a great one!
59.
▲
by
lcolucci
2y ago
Our model can recursively extend video clips, so theoretically we could generate your 5-7min talking head videos today. In practice, however, error accumulates with each recursion and the video quality gets worse and worse over time. This i
60.
▲
by
lcolucci
2y ago
This is a bug in the model we're aware of but haven't been able to fix yet. It happens at the end of some videos but not all. Our hypothesis is that the "breakdown" happens when there's a sudden change in audio leve
More ›