7 ms·
How the cochlea computes (2024)
- p0w3n3d 11mo agoTbh I used to think that it does. For example, when playing higher notes, it's harder to hear the out-of-tune frequencies than on the lower notes.
- fallingfrog 11mo agoI haven't noticed that effect, to be honest. Actually I think its the really low bass frequencies that are harder to tune- especially if you remove the harmonics and just leave the fundamental. Are you perhaps experiencing some high frequency hearing loss?
- jacquesm 11mo agoIt's even more complex than that. The low notes are hard to tune because the fundamentals are very close to each other and you need to have super good hearing to match the beats, fortunately they sound for a long time so that helps. Missing fundamentals are a funny thing too, you might not be 'hearing' what you think you hear at all! The high notes are hard to tune because they sound very briefly (definitely on a piano) and even the slightest movement of the pin will change the pitch considerably. In the middle range (say, A2 through A6) neither of these issues apply, so it is - by far - the easiest to tune.
- TheOtherHobbes 11mo agoSee also, psychoacoustics. The ear doesn't just do frequency decomposition. It's not clear if it even does frequency decomposition. What actually happens is lot of perceptual modelling and relative amplitude masking which makes it possible to do real-time source separation. Which is why we can hear individual instruments in a mix. And this ability to separate sources can be trained. Just as pitch perception can be trained, with varying results from increased acuity up to full perfect pitch. A component near the bottom of all that is range-based perception of consonance and dissonance, based on the relationships between beat frequencies and fundamentals. Instead of a vanilla Fourier transform, frequencies are divided into multiple critical bands (q.v.) with different properties and effects. What's interesting is that the critical bands seem to be dynamic, so they can be tuned to some extent depending on what's being heard. Most audio theory has a vanilla EE take on all of this, with concepts like SNR, dynamic range, and frequency resolution. But the experience of audio is hugely more complex. The brain-ear system is an intelligent system which actively classifies, models, and predicts sounds, speech, and music as they're being heard, at various perceptual levels, all in real time.
- jacquesm 11mo agoYes, indeed, to think about the ear as the thing that hears is already a huge error. The ear is - at best - a faulty transducer with its own unique way of turning air pressure variations into nerve impulses and what the brain does with those impulses is as much a part of hearing as the mechanics of the ear, just like a computer keyboard does not interpret your keystrokes, it just turns them into electrical signals.
- deleted 11mo ago[deleted]
- fallingfrog 11mo agoWelll. On guitar you cant really use the "matching the beats" or the thing where you play the 4th on the string below and make them sound in unison, because if you do that all the way up the neck your guitar will be tuned to Just intonation instead of equal interval intonation and certain chords will sound very bad. A series of perfect 4ths and a perfect 3rd does not add up to an octave. Its better to reference everything to the low e string and just kind of know where the pitches are supposed to land. That's a side note, the rest of what you wrote was very informative!
- philip-b 11mo agoNo, it's vice versa. If two wind instruments play unison slightly out of tune from each other, it will be very noticeable. If the bass is slightly out of tune or mistakenly plays a different note a semitone up or down, it's easy to not notice it.
- bloppe 11mo agoMan, I've been spreading disinformation for years.
- rolph 11mo agothe closest i have been, was acoustic phase discrimination by owls. there appears to be no software for this, its all hardware, the signal format flips as it travels through the anatomy.
- nakulgarg22 11mo agoThis might be interesting for you - https://nakulg.com/assets/papers/owlet_mobisys2021_nakul.pdf https://nakulg.com/assets/papers/owlet_mobisys2021_nakul.pdf Owls use asymmetric skull structure which helps them in spatial perception of sound.
- rolph 11mo agothat was the start of it. the offset otic openings result in differential arrival times of the acoustic peaks, thus phase differential. neurosynaptically, there is no phase, there is frequency shift corresponding to presynaptic intensity, and there is spatio-temporal integration of these signals. temporal integration is where "phase" matters its all a mix of "digital" all or nothing "gates" and analog frequency shift propagation of the "gate" output. its all made nebulous by the adaptive, and hysteretic nature of the elements in neural "circuitry"
- lukeinator42 11mo agoalso, the common ancestor of mammals and birds did not have a tympanic ear, so sound localization evolved differently in the avian vs. mammalian hearing systems. A good review is here: https://journals.physiology.org/doi/pdf/10.1152/physrev.00026.2009 https://journals.physiology.org/doi/pdf/10.1152/physrev.0002.... How the brain calculates interaural time delays is actually an interesting problem as the time delays are so short, that it is less time than a neuron has to fire an action potential.
- 11mo ago
- rolph 11mo agoFT is frequency domain representation. neural signaling by action potential, is also a representation of intensity by frequency. the cochlea is where you can begin to talk about bio-FT phenomenon. however the format "changes" along the signal path, whenever a synapse occurs.
- xeonmc 11mo agoNit: It’s an unfortunate confusion of naming conventions, but Fourier Transform in the strictest sense implies an infinite “sampling” period, while the finite “sample” period counterpart would correspond to Fourier Series even though we colloquially refer to them interchangeably. (I had put “sampling” in quotes as they’re actually “integration period” in this context of continuous time integration, though it would be less immediately evocative of the concept people are colloquially familiar with. If we actually further impose a constraint of finite temporal resolution so that it is honest-to-god “sampling” then it becomes Discrete Fourier Transform, of which the Fast Fourier Transform is one implementation of.) It is this strict definition that the article title is rebuking, but it’s not quite what the colloquial usage loosely evokes in most people’s minds when we usually say Fourier Transform as an analysis tool. So this article should have been comparing to Fourier Series analysis rather than Fourier Transform in the pedantic sense, albeit that’ll be a bit less provocative. Regardless, it doesn’t at all take away from the salient points of this excellent article which are really interesting reframing of the concepts: what the ear does mechanistically is applying a temporal “weigting function” (filter) so it’s somewhere between Fourier series and Fourier transform. This article hits the nail on the head on presenting the sliding scale of conjugate domain trade offs (think: Heisenberg)
- meowkit 11mo agoI was a bit peeved by the title, but I think its a fair use of clickbait as the article has a lot of little details about acoustics in humans that I was unfamiliar with (i.e. a link to a primer on the the transduction implementation of cochlear cilia) But yeah there is a strict vs colloquial collision here.
- BrenBarn 11mo agoYeah, it's sort of like saying the ear doesn't do "a" Fourier transform, it does a bunch of Fourier transforms on samples of data, with a varying tradeoff between temporal and frequency resolution. But most people would still say that's doing a Fourier transform. As the article briefly mentions, it's a tempting hypothesis that there is a relationship between the acoustic properties of human speech and the physical/neural structure of the auditory system. It's hard to get clear evidence on this but a lot of people have a hunch that there was some coevolution involved, with the ear's filter functions favoring the frequency ranges used by speech sounds.
- tryauuum 11mo agoman I need to finally learn what a Fourier transform is
- TobTobXX 11mo ago3Blue1Brown has a really good explanation here: https://www.youtube.com/watch?v=spUNpyF58BY https://www.youtube.com/watch?v=spUNpyF58BY It gave me a much better intuition than my math course.
- garbageman 11mo agoIt's an absolutely brilliant bit of maths that breaks a complex waveform into the individual components. Kind of like taking an orchestral song and then working out each individual instrument's contribution. Learning about this left me honestly aghast and in shock that it's not only possible but that someone (Joseph Fourier) figured it out and then shared it with the world. This video does a great job explaining what it is and how it works to the layman. 3blue1brown - https://www.youtube.com/watch?v=spUNpyF58BY https://www.youtube.com/watch?v=spUNpyF58BY
- jama211 11mo agoHahaha, I was working on learning these in second year uni… which was also exactly when I switched from an electrical engineering focussed degree to a software one! Perhaps finally I should learn too…
- adzm 11mo agothe very simplest way to describe it: it is what turns a waveform (amplitude x time) to a spectrogram like on a stereo (amplitude x frequency)
- lala_ 11mo ago[flagged]
- edbaskerville 11mo agoTo summarize: the ear does not do a Fourier transform, but it does do a time-localized frequency-domain transform akin to wavelets (specifically, intermediate between wavelet and Gabor transforms). It does this because the sounds processed by the ear are often localized in time. The article also describes a theory that human speech evolved to occupy an unoccupied space in frequency vs. envelope duration space. It makes no explicit connection between that fact and the type of transform the ear does—but one would suspect that the specific characteristics of the human cochlea might be tuned to human speech while still being able to process environmental and animal sounds sufficiently well. A more complicated hypothesis off the top of my head: the location of human speech in frequency/envelope is a tradeoff between (1) occupying an unfilled niche in sound space; (2) optimal information density taking brain processing speed into account; and (3) evolutionary constraints on physiology of sound production and hearing.
- AreYouElite 11mo agoDo you believe it might be possible that the frequency band of human speech is not determined by such factors at all but more of a function of height? kids have higher voices adults have deeper voices. Similar to stringed instruments: viola high pitched and bass low pitched. I'm no expert in these matters just speculating...
- fwip 11mo agoIt's not height, but vocal cord length and thickness. Longer vocal cords (induced by testosterone during puberty) vibrate more slowly, with a lower frequency/pitch.
- matthewdgreen 11mo agoIf you take this thought process even farther, specific words and phonemes should occupy specific slices of the tradeoff space. Across all languages and cultures, an immediate warning that a tiger is about to jump on you should sit in a different place than a mother comforting a baby (which, of course, it does.) Maybe that even filters down to ordinary conversational speech.
- 11mo ago
- kazinator 11mo ago> A Fourier transform has no explicit temporal precision, and resembles something closer to the waveforms on the right; this is not what the filters in the cochlea look like. Perhaps the ear does someting more vaguely analogous to a discrete Fourier transforms on samples of data, which is what we do in a lot of signal processing. In signal processing, we take windowed samples, and do discrete transforms on these. These do give us some temporal precision. There is a trade off there between frequency and temporal precision, analgous to the Pauli exclusion principle in quantum mechanics. The better we know a frequency, the less precisely we know the timing. Only an infinite, periodic signal has a single precise frequency (or precise set of harmonics) which are infinitely narrow blips in the frequency domain. The continuous Fourier transform deals with periodic signals only. We transform an entire function like sin(x) over the entire domain. If that domain is interpreted as time, we are including all of eternity, so to speak from negative infinite time to positive.
- xeonmc 11mo ago> analgous to the Pauli exclusion principle Did you mean the Heisenberg Uncertainty Principle instead? Or is there actually some connection of Pauli Exlusion Principle to conjugate transforms that I was’t aware of?
- kvakkefly 11mo agoThey are not connected afaik.
- HarHarVeryFunny 11mo ago> There is a trade off there between frequency and temporal precision Sure, and the FFT isn't inherently biased towards one vs the other. If you take an FFT over a long time window (narrowband spectrogram) then you get good frequency resolution at the cost of time resolution, and vice versa for a short time window (wideband spectrogram). For speech recognition ideally you'd want to use both since they are detecting different things. TFA is saying that this is in fact what our cochlea filter bank is doing, using different types of filter at different frequency ranges - better frequency resolution at lower frequencies where the formants are (carrying articulatory information), and better time resolution at the high frequencies generated by fricatives where frequency doesn't matter but accurate onset detection is useful for detecting plosives.
- adornKey 11mo agoThis subject has bothered me for a long time. My question to guys into acoustics was always: If the cochlea performs some kind of Fourier transform, what are the chances, that it uses sinus waves as a base for the vector-space? - if it did anything like that it could just as good use any slightly different wave-forms as a base for transformation. Stiffness and non-linearity will for sure take care that any ideal rubber model in physics will in reality be different from the perfect sinus.
- FarmerPotato 11mo agoI find it beautiful to see the term "sinus wave."
- empiricus 11mo agowell, cochlea is working withing the realm of biological and physical possibilities. basically it is a triangle through which waves are propagating, and sensors along the edge. smth smth this is similar to a filter bank of gabor filters that respond to rising freq along the triangle edge. ergo you can say fourier, but it only means sensors responding to different freq becasue of their location.
- adornKey 11mo agoYeah, but not only the frequency is important - the wave-form is very relevant. For example if your wave-form is a triangle, listerners will tell you that it is very noisy compared to a simple sinus. If you use sinus as a base of your vector space triangles really look like a noisy mix. My question is, if the basic elements are really sinus, or if the basic Eigen-Waves of the cochlea are other Wave-Forms (e.g. slightly wider or narrower than sinus, ...). If physics in the ear isn't linear, maybe sinus isn't the purest wave-form for a listener. Most people in Physics only know sinus and maybe sometimes rectangles as a base for transformations, but mathematically you could use a lot of other things - maybe very similar to sinus, but different.
- kragen 11mo agoBut if you apply a frequency-dependent phase shift to the triangle wave, nobody will be able to tell the difference unless the frequency is very low.
- gowld 11mo agoWhy is there no box diagram for cochlea "between wavelet and Gabor" ?
- anticensor 11mo agoWould look still too much like wavelet.
- shermantanktop 11mo agoThe thesis about human speech occupying less crowded spectrum is well aligned with a book called "The Great Animal Orchestra" (https://www.amazon.com/Great-Animal-Orchestra-Finding-Origins/dp/0316086878 https://www.amazon.com/Great-Animal-Orchestra-Finding-Origin...). That author details how the "dawn chorus" is composed of a vast number of species making noise, but who are able to pick out mating calls and other signals due to evolving their vocalizations into unique sonic niches. It's quite interesting but also a bit depressing as he documents the decline in intensity of this phenomenon with habitat destruction etc.
- HarHarVeryFunny 11mo agoBirds have also evolved to choose when to vocalize to best be heard - doing so earlier in urban areas where later there will be more traffic noise, and later in some forest environments to avoid being drowned out by the early rising noisy insects.
- HaroldBrill 11mo agoMy friend, through all these voluminous deafening comments, we certainly heard your loaded minimalist beautiful tune by you simply broadcasting your marvelous name!
- kulahan 11mo agoProbably worth mentioning that as evolutions that allow them to compete well in nature die out, ones that allow them to compete well in cities takes their place. Evolution is always a series of tradeoffs. Maybe we don't have sonic variation, but temporal instead.
- brcmthrowaway 11mo agoOT: Does anyone here believe in Intelligent Design?
- xeonmc 11mo agoAs low-level physical mechanistic processes? Absolutely not. As higher-order, statistically transparent abstract nudges of providence existing outside the confines of causality? Metaphysically interesting but philosophically futile.
- jibal 11mo agoHopefully not ... it has been thoroughly debunked, whereas the theory of evolution is supported by massive amounts of data and is the foundation of the entire science of biology.
- superb-owl 11mo agoThe title seems a little click-baity and basically wrong. Gabor transforms, wavelet transforms, etc are all generalizations of the fourier transform, which give you a spectrum analysis at each point in time The content is generally good but I'd argue that the ear is indeed doing very Fourier-y things.
- anyfoo 11mo agoAgree on the click-baity part, but as for being wrong... not if we're really pedantic. As you've said, Gabor and wavelet are basically generalizations of the Fourier Transform, not actually Fourier Transforms. Just like FS/DFT/DTFT aren't really Fourier Transforms either. On one corner of the square, you have Fourier Transforms, which are essentially contiguous and infinite. On the opposite corner, you have the DFT, which is both finite (or periodic) and discrete. Hearing is more akin to a Fourier Series, which is finite/periodic but contiguous. That's probably not what the article aims at addressing, though. But then wavelet transforms are different from Fourier Series again, because you have shifted and stretched shapes (some of them quite weird) instead of sinusoids. But yeah, colloquially, I agree, the ear is indeed doing very Fourier-y things.
- fat_cantor 11mo agoIt's a graduate student writing a journal club article about the Lewicki 2002 paper, which is very good, and whose abstract states the idea more precisely: "The form of the code depends on sound class, resembling a Fourier transformation when optimized for animal vocalizations and a wavelet transformation when optimized for non-biological environmental sounds"
- debo_ 11mo agoFourear transform
- antognini 11mo agoIf you want to get really deep into this, Richard Lyon has spent decades developing the CARFAC model of human hearing: Cascade of Asymmetric Resonators with Fast-Acting Compression. As far as I know it's the most accurate digital model of human hearing. He has a PDF of his book about human hearing on his website: https://dicklyon.com/hmh/Lyon_Hearing_book_01jan2018_smaller.pdf https://dicklyon.com/hmh/Lyon_Hearing_book_01jan2018_smaller...
- ddingus 11mo agoThank you! This is an excellent work. Much appreciated
- dr_dshiv 11mo agoNeither this article nor this book discuss the fact that hair cells phase lock to sound pulses. While individual neurons can fire no more than 200hz, populations of neurons are capable of phase locking to frequencies up to thousands of hertz. Because cochlear implants only rely on stimulating the places in the cochlea related to particular frequencies but do not play the actual frequencies themselves (for reasons unknown), people with cochlear implants can detect frequency differences but lose appreciation for music.
- antognini 11mo agoYes this is another important difference between human auditory perception and classical signal processing algorithms. Typically when processing audio we take a Fourier transform and then throw away the phase information. Mostly the amplitude information is all you need to understand a sound, but the ear actually is capable of picking up phase information. (I thought this was discussed at some point in Lyon's book but it's admittedly been many years since I read it, so I can't remember for sure.)
- fluoridation 11mo agoWhat does that mean, though? If you invert the sign of a waveform it sounds the same, so if it's not picking that up, what phase and relative to what does it pick up?
- javier_e06 11mo agoThis is fascinating. I know of vocoders in the military hardware that encode voices to resemble something more simple for compression (a low-tone male voice), smaller packets that take less bandwidth. This evolution of the ear to must also have evolved with our vocal chords and mouth to occupy available frequencies for transmission and reception for optimal communication. The parallels with waveforms don't end there. Waveforms are also optimized for different terrains (urban, jungle). Are languages organic waveforms optimized to ethnicity and terrain? Cool article indeed.
- rolph 11mo agosupplemental: Neuroanatomy, Auditory Pathway https://www.ncbi.nlm.nih.gov/books/NBK532311/ https://www.ncbi.nlm.nih.gov/books/NBK532311/ Cochlear nerve and central auditory pathways https://www.britannica.com/science/ear/Cochlear-nerve-and-central-auditory-pathways https://www.britannica.com/science/ear/Cochlear-nerve-and-ce... Molecular Aspects of the Development and Function of Auditory Neurons https://pmc.ncbi.nlm.nih.gov/articles/PMC7796308/ https://pmc.ncbi.nlm.nih.gov/articles/PMC7796308/
- fennec-posix 11mo ago"It appears that human speech occupies a distinct time-frequency space. Some speculate that speech evolved to fill a time-frequency space that wasn’t yet occupied by other existing sounds." I found this quite interesting, as I have noticed that I can detect voices in high-noise environments. E.g. HF Radio where noise is almost a constant if you don't use a digital mode.
- deleted 11mo ago[deleted]
- amelius 11mo agoWhat does the continuous tingling of a hair cell sound like to the subject?
- dfboyd 11mo agotinnitus
- xmcqdpt2 11mo agoMany versions of this article could be written: The computer does not do a Fourier transform (FFT computes the discrete Fourier transform) Spectroscope dont do a Fourier transform (it's actually the short time FT) The only thing that actually does Fourier transform is a mathematician, with a pen and some paper.
- tim333 11mo agoNice to see a video for the tip links and ion channels. I spent a while reading up on that stuff because I was trying to figure what causes my tinnitus. My best guess is if the hairs over bend, that stuff can break and an ion channel get stuck open causing the cell to fire continually. Another fun ear fact is they incorporate active amplification. You can hook an electrical signal to the loudspeaker type cell to make it vibrate around https://youtu.be/pij8a8aNpWQ https://youtu.be/pij8a8aNpWQ
- deleted 11mo ago[deleted]
- Cadwhisker 11mo agoJust a warning that the video ends with a loud, high pitched tone that will make you want to rip your headphones off. Ironic for a video about hearing.
- tsoukase 11mo agoAs the auditory associative cortex in parietal lobe discriminates frequencies, there must be some time-frequency transform between the ear and the brain. This must be discrete (as neurons fire in bursts and there is a finite frequency resolution capacity) and finite time. The poor man's conversion of finite to equivalent infinite time is if you assume an infinite signal where the initial finite one is repeated infinately to the past and the future.
- dboreham 11mo agoSpoiler: yes it does, but the author isn't familiar with how the term Fourier Transform is used in signal processing.
- shannifin 11mo agoI've always thought the basilar membrane was a fascinating piece of biological engineering. Whether or not the difference between its behavior vs FT really matters depends on the context. Audio processing on a computer, FFT is often great. Trying to understand / model human sound perception, particularly in relation to time, FFT has weaknesses.
- hamonrye 11mo ago[dead]
- rattan12138 11mo agoWow, this discussion about how our ears work is mind-blowing! It's amazing how complex sound processing is, and the comparison to signal processing concepts is really illuminating.
- hbarka 11mo agoSomewhere here must lie the cure to tinnitus.
- Traubenfuchs 11mo agoNo, tinnitus beyond the cochlea, beyond the hearing nerve even, in some part of the audio processing brain tissue. Cutting the hearing nerve does not cure tinnitus. It develops due to a destruction of hearing cells that leads the brain to upregulate gain to catch a weak/absent signal, when the deprivation pattern is just right. (no tinnitus develops when the hearinf nerve is cut -> deprivation pattern matters)