4 ms·
This is a good observation. Can you share the videos you’re seeing this with? For me, normal talking tends to work well even on long generations. But singing or
by lcolucci 2y ago
This is a good observation. Can you share the videos you’re seeing this with? For me, normal talking tends to work well even on long generations. But singing or expressive audio starts to devolve with more recursions (1 forward pass = 8 sec). We’re working on this.