7 ms·
It's unfortunate that the article doesn't get into the fundamental limits of spectrogram resolution which are based on the famous uncertainty principle(https://
by LeegleechN 5y ago
It's unfortunate that the article doesn't get into the fundamental limits of spectrogram resolution which are based on the famous uncertainty principle(https://en.wikipedia.org/wiki/Fourier_transform#Uncertainty_principle https://en.wikipedia.org/wiki/Fourier_transform#Uncertainty_...). For example there is a fundamental tradeoff between frequency resolution and time resolution similar to the position/momentum tradeoff in quantum mechanics. The Continuous Wavelet Transform which is alluded to in the article is a way to tune that tradeoff by frequency bin to best align with human sound perception.
- andai 5y agoI've been wondering about the apparent contradiction between the limitations of spectrograms and the remarkable fidelity of MP3 files, which I thought operated along similar lines. When you convert a spectrogram back into sound it sounds like crap, but then how does MP3 store the frequency information (and why can't we use that for visualizations)? The math is beyond my understanding, can anyone give some kind of analogy maybe?
- gugagore 5y agoThe general concepts are described here: https://en.m.wikipedia.org/wiki/Psychoacoustics https://en.m.wikipedia.org/wiki/Psychoacoustics I'm not sure what you mean by converting the spectrogram to sound, but my guess is that the windowing done on the short-time Fourier transform is causing artifacts.
- achillesheels 5y agoMy hypothesis: it is stored magnetically (after all magnetic sinusoidals exist) and converted electrically once the mp3 is activated in time.
- achillesheels 5y agoI’m answering further: it is a result of the software programs limitations to reconstruct the mathematically perfect electromagnetic waveform quantized beforehand which is causing noise. No one writes perfect code.
- akomtu 5y agoThat's a funny thread. Plato distinguishes the world of ideas from the world of matter. The latter is all about magnetic fields, while the former has none of them, as it's the world of archetypes, hilbert spaces and fourier transforms. So mp3 belongs to the world of ideas, it's a concept that expresses a projection of an abstract waveform to a specific basis of functions. Sound represents a particular embodiment (aka a mechanical wave) of an abstract waveform in the world of matter. The fact that this mechanical wave is driven by magnetic fields is irrelevant and is also a tautology, because in the world of matter everything is driven by magnetic fields.
- achillesheels 5y agoWhat do you suppose is frequency content or information? It must principally exist in an analog continuous-time domain before being quantized and cached, correct? So how do these quantities, when processed, not correspond to what is ultimately electromagnetic? I keep reading by commenters that a purely mathematical space which operates on physical movement has no relevance to physical movement vis-a-vis spectral analysis and particularly in how this analyzer - again embedded with recursive electromagnetic charges (software code) - does not affect the original time-sample. I am aware now that MP3 encoding is capturing more information than the analyzer - by design - but how can electromagnetic resonances not be considered in the discussion to warrant continual downvotes? (I'd like to remind readers that downvotes are not made available for all users, but only those who meet a certain criteria, i.e. those which reflect the general Hacker News community.)
- pixelbro 5y ago> but how can electromagnetic resonances not be considered in the discussion to warrant continual downvotes Because everyone's been telling you the same thing in a multitude of different ways, but you refuse to get the point. Yes, the operation of the software we write can be affected by the physical realities of the hardware we use to execute them, but any time this causes the software to behave differently than it would in a purely hypothetical computer, this is considered an error and the results invalid. We even have hardware that automatically corrects for when rare events such as freakin' cosmic rays cause the value of a bit to flip (ECC memory). The domain of software engineering is abstracted from physical reality. There's nothing useful to be gained from such discussions, because the whole job of hardware engineers is to enable us to operate at a higher level where we don't have to concern ourselves with the quantum electrodynamics necessary to make the transistors do their jobs.
- achillesheels 5y agoCan someone please explain to me why a hypothesis of magnetic frequencies contained in an electromagnetic cache (i.e. MP3) are not plausible to be transformable in real electrical output as opposed to downvoting me? It would be much more productive dialogue.
- drdeca 5y agoThis seems a little word-salad-y ? MP3s are stored in bitstrings. It doesn't matter what the medium these bitstrings are stored in. The question being asked is a question about information, not about physics. So, your response is inapplicable.
- achillesheels 5y agoYes, but what are bitstrings but elementally electrical action? What do you think information is fundamentally if it is not contained in a physical paradigm?
- detaro 5y agoYou could write the bitstring of MP3 on a piece of paper with a pen and it wouldn't change a thing.
- achillesheels 5y agoSir, surely you don’t imply that a computer can process that.
- pixelbro 5y agoWhat he's implying, and what you're missing here, is that the discussion is about the FFT and MP3 algorithms, and how that signal processing affects the signal. The physical substrate which executes the algorithm is irrelevant to the discussion. Indeed, even the fact that the signal represents sound waves is irrelevant. We could build a computer that executes the same instructions on the same inputs mechanically, using gears and valves and whatnot, or perform the instructions manually on paper, and the algorithm would result in exactly the same output.
- bad_username 5y agoI implemented a simple clone of mp3 and it was not that hard. If you do a discrete Fourier transform of the audio (in small overlapping windows), quantize the resulting coefficients, and compress them losslessly using the Huffman codes, you will end up with something not that far from mp3. The human ear is quite forgiving to the effects of quantization in frequency domain. MP3 does not have remarkable fidelity though. MP3, and my clone of it, suffers from time domain artifacts. Quantization in the frequency domain causes distortion in the time domain as well, negatively affecting high frequency transient sounds like cymbals. That is more noticeable. Newer generation codecs like AAC handle transients much better, but they are considerably more advanced, and often use different transforms like wavelet transform.
- jcelerier 5y ago> When you convert a spectrogram back into sound it sounds like crap fft gives you the spectrum + the phase. if you only use the spectrum to resynthesise you're missing half the information. temporal domain <-> spectral domain is a 99.9999999% lossless (not 100% I believe because of floating-point shenanigans, but enough to not matter at all) transform in both directions.
- electriccello 5y agoI think the trouble you're running into is that a spectrogram discards phase information so it's not informationally complete, and impossible to perfectly invert. Basically, a Fourier Transform represents a sound as a series of many sound waves at different frequencies added together. In order to make a pretty picture, the phase is thrown away, and only the magnitude of each wave is shown. The trouble is, to go back to a pleasant/accurate sound, we need that phase information that is missing.
- gbh444g 5y agoI was thinking this is the case, until I stumbled upon a stackoverflow question that explains how to recover the phase data from overlapping FFT frames. The key word here is "overlapping".
- crazygringo 5y ago> there is a fundamental tradeoff between frequency resolution and time resolution I've always found it interesting that while that's fundamentally true in terms of information, my understanding is that we perceive things with far more resolution than the uncertainty principle would allow. Specifically, we're able to judge frequencies with far more accuracy than a fuzzy spectrogram would suggest. From what I understand, our brain essentially performs a kind of "deconvolution" on the fuzzy frequency data to identify a far "sharper" and defined frequency, which is relatively straightforward since the frequency "spread" is a known quantity. This works well most of the time because we correctly assume we're dealing with relatively isolated sound sources emanating a distinct fundamental with a distinct series of overtones. Our perception can become innacurate when that assumption fails to hold, and so sounds merge or become indistinguishable, we hear beat tones that don't technically exist, our brain gives up trying to hear frequencies and classifies it all as noise, etc. I've never come across audio spectrogram software that attempted to perform a frequency deconvolution in a way that roughly simulates what our own ears do, but I'd love to know if anyone else has and could point me to it.
- carlosf 5y agoSearch for voice coding. That's how many voice coding algorithms work, you try to find a digital filter that generates a sound that is as close as the original according to a perception based metric, then transmit the filter coefficients. I don't remember the exact details, but if I'm not mistaken generating this sort of metric is really time consuming.
- crazygringo 5y agoThanks! That's definitely along the same lines, although that's for the special case of only a single fundamental frequency (with overtones). And I'm not sure it uses deconvolution -- I've heard of maximum likelihood estimation being used, in order to get a single frequency, though the ideas are closely related. Using deconvolution would be more for the purposes of cleaning up a general-purpose spectrogram for human eyes -- for analysis and for sound editing, whether a single voice or a band of several instruments.
- 5y ago
- gbh444g 5y agoI did experiment with CWT in past [1] and was disappointed, to be honest. Not only it's grossly slow and complicated, it hardly gives more fidelity than plain FFT, it has the "window problem" which makes the low freqs too blurred and the high freqs too sharp, and it has the "wrapping ends" problem that makes it necessary to pad the input (about 1 million samples at least) with sufficient zero padding on both ends, as otherwise the two ends will interfere with each other. Below is my GPU-based CWT that's 50x slower than the JS-only version in the post above. [1] https://soundshader.github.io/?s=cwt https://soundshader.github.io/?s=cwt