4 ms·
There is really no such thing as 'instantaneous frequency response'. For any frequency to meaningfully exist, you need data for the corresponding period. e.g. I
by mbell 7y ago
There is really no such thing as 'instantaneous frequency response'. For any frequency to meaningfully exist, you need data for the corresponding period. e.g. If the audio contains content to 20hz, you need at least 1/40th to 1/20th of a second of data for that to materialize.
Put another way - what you are proposing it looping the buffer which is what some devices do, portable CD players were kinda notorious for it and it doesn't sound much better than cracks or pops. Computers also have a tendency to fall into buffer looping when the system hangs (which is likely the result of the failure mode of realtek codecs).
- amelius 7y ago> There is really no such thing as 'instantaneous frequency response' Yes that's true, I'm proposing something that uses an approximation of it. Consider it from a different angle: the inner ear essentially performs a Fourier transform. At every moment the "instantaneous" spectrum determines which hair cells are triggered. Now what I propose is to keep triggering those same hair cells (and not any others) when the buffer runs dry. The exact way of accomplishing this is left as an exercise (though using short windows where you take a FFT could be a good approximation).
- human20190310 7y ago> The exact way of accomplishing this is left as an exercise Perhaps you should undertake this exercise and let us know how it sounds :) EDIT: In my experience with audio, when I have a bug that introduces even the slightest discontinuity (or even just a cusp) in the audio, well short of a pop to silence, I can still hear a "weirdness". Ears are pretty attuned to things that sound unnatural. I'm not confident that essentially "forging" the audio is going to sound natural.
- StavrosK 7y agoWhat if you train a deep neural network on the song so far, so it can generate plausible-sounding music whenever the buffer drops? You can even hang intentionally to generate original music! (/s, please don't)
- squeaky-clean 7y agoI like the idea but one problem is that you usually encounter a buffer underrun when the cpu can't keep up, and so adding an extra step would require something like leaving enough processing headroom each buffer to halt it early and run the approximater. edited typo
- nitrogen 7y agoAs another comment mentioned, this is done by conferencing software to deal with packet loss. It sounds like they either loop the previous DCT frame and gradually fade out, or feed the time domain output into a reverb, then cross-fade from 100% dry to 100% wet if the buffer is about to run out (someone on HN mentioned a while back that this approach was patented). You could maybe make an argument that it would be useful in live music settings to prevent a bad situation from sounding even worse, and maybe you'd put it on some audio software so you can sort of still enjoy playing music on a crappy system, but really, it's best to have hardware and software that can 100% guarantee keeping up with audio processing.