Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jmvalin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
jmvalin
3y ago
Actually, what we're doing from DRED isn't that far from what you're suggesting. The difference is that we keep more information about the voice/intonation and we don't need the latency that would otherwise be added
2.
▲
by
jmvalin
3y ago
What the PLC does is (vaguely) equivalent to momentarily freezing the image rather than showing a blank screen when packets are lost. If you're in the middle of a vowel, it'll continue the vowel (trying to follow the right energy)
3.
▲
by
jmvalin
3y ago
Well, there's different ways to make things up. We decided against using a pure generative model to avoid making up phoneme or words. Instead, we predict the expected acoustic features (using a regression loss), which means that model
4.
▲
by
jmvalin
3y ago
Quoting from our paper, training was done using "205 hours of 16-kHz speech from a combination of TTS datasets including more than 900 speakers in 34 languages and dialects". Mostly tested with English, but part of the idea of rel
5.
▲
by
jmvalin
3y ago
As part of the packet loss challenge, there was an ASR word accuracy evaluation to see how PLC impacted intelligibility. See https://www.microsoft.com/en-us/research/academic-program/au... The good news is th
6.
▲
by
jmvalin
3y ago
(Opus author here) I'm curious what kind of "glitch" this is referring to.
7.
▲
by
jmvalin
5y ago
> This is entirely my fault, and I take all the blame for that. You shouldn't be blaming yourself, it was the best thing to do. Some people may have been confused over who the "good guys" were in this mess. By taking over
8.
▲
by
jmvalin
6y ago
No, exactly none of that data was used for training. The training was done before the demo that was asking for noise contributions. The contributions are CC0, but were never used (i.e. totally unknown dataset quality).
9.
▲
by
jmvalin
6y ago
All major browsers now implement WebRTC, including Opus support. Also, most browsers now support Opus playback in HTML5, though AFAIK Safari only supports it in the CAF container. See https://caniuse.com/#search=opus
10.
▲
by
jmvalin
8y ago
I didn't say "impossible", merely "not simple". The minute you bring in a GAN, things are already not simple. Also, I'm not aware of any work on a GAN that works with a network that does conditional sampling (l
11.
▲
by
jmvalin
8y ago
In theory, it wouldn't be too hard to implement with an neural network. In theory. In practice, the problem is figuring out how to do the training because I don't have 2 hours of your voice saying the same thing as the target voic
12.
▲
by
jmvalin
8y ago
The ceptrum that takes up most of the bits (or the LSPs in other codecs) is actually a model of the larynx -- another reason why it doesn't do well on music. Because of the accuracy needed to exactly represent the filter that the laryn
13.
▲
by
jmvalin
8y ago
Actually, what's in the demo already includes pruning (through sparse matrices) and indeed, it does keep just 1/10 of the weights as non-zero. In practice it's not quite a 10x speedup because the network has to be a bit bigge
14.
▲
by
jmvalin
8y ago
Iridium appears to be using a vocoder called AMBE. Its quality is similar to the one of the MELP codec from the demo and it also runs at 2.4 kb/s. LPCNet at 1.6 is a significant improvement over that -- if you can afford the complexity
15.
▲
by
jmvalin
8y ago
Well, in the case of music, what happens is that due to the low bit-rate there are many different signals that can produce the same features. The LPCNet model is trained to reproduce whatever is the most likely to be a single person speakin
16.
▲
by
jmvalin
8y ago
Keep in mind that the very first CELP speech codec (in 1984) used to take 90 seconds to encode just 1 second of speech... on a Cray supercomputer. Ten years later, people had that running in their cell phones. It's not just that hardwa
17.
▲
by
jmvalin
8y ago
Actually, this won't work at all for music because it makes fundamental assumptions that the signal is speech. For normal conversations, it should work, though for now the models are not yet as robust as I'd like (in case of noi
18.
▲
LPCNet: DSP-Boosted Neural Speech Synthesis
(people.xiph.org)
4 points
by
jmvalin
8y ago
|
0 comments
19.
▲
by
jmvalin
8y ago
Like many other audio codecs, Opus lets the encoder decide how to spend the bits is has -- on what frame and on what frequency bands. On top of that is has a few special features that also require decisions from the encoder. So while the de
20.
▲
by
jmvalin
8y ago
The reason we are not calling it Opus 2 is that it could confuse some people into thinking we broke compatibility. Opus 1.3 is perfectly compatible with Opus 1.0, and all future releases will keep that compatibility.
21.
▲
by
jmvalin
8y ago
Thanks for reporting that. It's fixed now.
22.
▲
by
jmvalin
8y ago
Getting something like AOM would have been easy back in 1993 because the costs would have been much lower. Back then complexity had to be really low, which means most of the complicated modern tools were off the table. Coming up with someth
23.
▲
by
jmvalin
9y ago
The C part is mostly for low-level functions and was brought in to help bootstrap development (it's easier to work on improving a working encoder than one that doesn't work yet). The amount of Rust code is expected to increase a l
24.
▲
by
jmvalin
9y ago
There's already a Rust AV1 encoder project you might want to look at: https://github.com/xiph/rav1e
25.
▲
by
jmvalin
9y ago
There's a good reason all the listening tests have stopped at 96 kb/s. Above that, the quality of Opus, Vorbis and AAC is so close to transparency that it's pretty much impossible to get statistically significant results. Eve
26.
▲
by
jmvalin
9y ago
A lot of people get the impression it's only cancelling where there's no speech, but it's also cancelling during speech -- just not as much. If you look at the spectrogram at the top of the demo, you can see HF noise being at
27.
▲
by
jmvalin
9y ago
The main problem here is that you're depending on implementation-specific behaviour. If you train on a device, you have to run on a device with exactly the same behaviour. On top of that, some FPUs have very slow (trapping) denormal ha
28.
▲
by
jmvalin
9y ago
My comment about intelligibility refers to a human (with normal audition) directly listening to the output. When the output is used in a hearing aid, a cochlear implant, or a low bitrate vocoder, then noise suppression may be able to help i
29.
▲
by
jmvalin
9y ago
What you're describing is more or less why noise suppression algorithms in general cannot really improve intelligibility of the speech. Unless they're given extra cues (like with a microphone array), there's nothing they can
30.
▲
by
jmvalin
9y ago
For training I've had to use some non-free data, but there's also some free stuff around. The speech from the examples is from SQAM ( https://tech.ebu.ch/publications/sqamcd ) and I've also used a free spe
More ›