4 ms·
As others have said, you appear to be confusing bit depth dynamic range for total loudness. They are related in a way, but not in the way you seem to think. (I
by electrograv 7y ago
As others have said, you appear to be confusing bit depth dynamic range for total loudness. They are related in a way, but not in the way you seem to think. (I will demonstrate to you below how even 20 bits per sample can be extremely insufficient.)
First: Bits per sample just describe how many discrete amplitude values (2^bits) are possible at each sample of a recording. The waveform is quantized to these values.
To understand quantization in the context of dynamic range, imagine how many bits per sample are needed to recreate a very quiet sound without loss, and then check how many bits you need to extend that to reach very loud sounds in the same recording file.
For example: How much precision would you need to accurately record the sound of a pin dropping (10db)? 4 bits? 8 bits? 10 bits? 12 bits?
Let’s be really absurd and say we can use 4 bits — just 16 discrete values — to represent a pin dropping sound (10db) cleanly and indistinguishable from the real thing. This is so obviously impossible, given how terribly quantized the waveform would be, but let’s be generous and assume it works.
Now, for the same audio file to reach all the way up to 110db (not uncommon for bass drum hits in an orchestra for example) is an extra 100db of dynamic range, which is 100,000x the amplitude, which is a little over 16 bits in addition to the original 4. So, rounding down, we’d need 20 bits to represent 10db sounds (with quantization down to only 16 discrete amplitudes) and 110db sounds in the same recording.
I think it’s extremely obvious that even 20 bits in this example is far from sufficient. In fact, even 24 bits would be insufficient if 8 bits per sample are not good enough to record a pin dropping at 10db!
- NobodyNada 7y ago> How much precision would you need to accurately record the sound of a pin dropping (10db)? 4 bits? 8 bits? 10 bits? 12 bits? The absolute minimum threshold of human perception is about -8dB. 1 bit corresponds to about 6 dB of dynamic range, so the sound of a pin dropping at 10 dB can be represented perfectly within the limits of human hearing using only 3 bits. Additionally, with proper dithering, the effective dynamic range of a 16-bit digital signal is around 120 dB. That's sufficient to reproduce everything from the -8 dB absolute quietest audible sound all the way up to your 110 dB bass drum. See https://people.xiph.org/~xiphmont/demo/neil-young.html#toc_1bv2b https://people.xiph.org/~xiphmont/demo/neil-young.html#toc_1... for a more detailed analysis.
- electrograv 7y agoClaiming that a 10db pin drop can be accurately reproduced with only 3 bits per sample without dithering is a rather extraordinary claim, given that human hearing is sensitive to far more than 8 levels of amplitude <= 10db absolute SPL. A pin drop contains a continuous and gradual decay of the ringing over time, which most certainly is audible (at least subjectively) in more gradations than just 8 levels before reaching 0. But this isn’t just subjective: Studies have confirmed that humans can hear decibel differences of 1db quite reliably, and can even hear as little as 0.2db subconsciously (this can be objectively measured)! So long as you can hear a pin drop ringing at 1db, and also at 10db, it’s therefore obviously true that 3 bits per sample (undithered) is insufficient to express this (keep in mind, these samples are linear amplitudes!) So I’m not sure how you can possibly claim that 3bits is enough to replicate the continuous amplitude decay of a pin drop’s ringing sound. What am I missing? That’s said, I will now go read your link. I agree dithering is one viable way to expand the dynamic range, but you also seem to be claiming this 3bit pin drop is true without dithering.
- NobodyNada 7y agoI may be misunderstanding some of the math, but it's important to keep in mind that decibels are logarithmic. Adding 6 dB doubles the amplitude, which is why adding another bit to a digital audio format adds 6 dB of dynamic range. Since we can't hear any sound quieter than -8 dB SPL, then (I think) we can't hear a difference smaller than the 6 dB between -8 and -2 dB because that difference is quieter than the threshold of human perception. Similarly, the next perceptible change in amplitude would be to +1.5 dB (3x the threshold of perception), then 4 dB (4x), 6 dB (5x), 7.6 dB (6x), 8.9 dB (7x), and 10 dB (8x). So, since 10 dB is 8 times the threshold of human perception, I think eight discrete values should be enough to represent everything we can hear up to that point. Of course, I'm assuming we can only hear in discrete increments of the minimum threshold of perception, -- it makes sense to me that we would be unable to hear a difference in amplitude quieter than the threshold -- but this could very well be a mistaken assumption. (It's also worth pointing out that 10 dB is very, very quiet -- the article I linked says 20 dB is the standard noise floor for a soundproof room/recording studio.) Another way of thinking about this is in terms of quantization error. Our 3-bit pin drop signal can be interpreted as the original analog signal plus a quantization error, and that error is clearly always quieter than the minimum threshold of perception. Edit: This article [0] contains a table for the minimum detectable difference in sound level. It shows that, at a level of 5 dB SPL, the just-noticeable difference in amplitude is 2.5 dB at the frequency range of peak sensitivity -- which is pretty close to the numbers I found based on discrete multiples of the absolute threshold. So it looks like 3 bits is indeed sufficient for our pin drop (unless I'm mistaken somewhere). [0]: https://www.sciencedirect.com/topics/engineering/just-noticeable-difference https://www.sciencedirect.com/topics/engineering/just-notice...