The Physics of Hearing: How Your Ear Performs a Real-Time Fourier Transform

Your ear is a physics instrument of staggering sophistication. It detects pressure variations smaller than the diameter of a hydrogen atom, distinguishes frequencies across a thousand-fold range, processes sounds arriving microseconds apart to locate their source, and does all of this in real time, continuously, for decades. Here's how.

Table of Contents

The Most Sensitive Instrument You Own

Right now, as you read this, your ears are detecting pressure variations in the air around you. Not large pressure variations — the quietest sound you can hear corresponds to an eardrum displacement of about 10 picometres. That’s less than the diameter of a hydrogen atom. It’s roughly one tenth the width of a single atomic nucleus of a heavier element.

At the threshold of hearing, your ear is detecting vibrations on the atomic scale. If it were slightly more sensitive, you’d hear the random thermal motion of air molecules — the constant, meaningless noise of Brownian motion. Evolution has pushed the ear to essentially the physical limit of what a mechanical sensor can detect.

And that’s just the sensitivity. The ear also covers a frequency range from about 20 Hz to 20,000 Hz — a thousand-fold range. It handles an intensity range from the threshold of hearing to the threshold of pain — a factor of about one trillion (10¹²) in power, which we compress into the 0–120 decibel scale. It can distinguish two frequencies that differ by less than 0.3%. It can detect the arrival time difference of a sound between your two ears to within about 10 microseconds. And it does all of this simultaneously, in real time, without pause, for your entire life.

No microphone we’ve built matches all of these specifications at once. Your ear is, by several measures, the most impressive physics instrument you carry around.

How Sound Reaches the Cochlea

Let’s trace a sound wave from the outside world to the point where it becomes a nerve signal.

The outer ear. The pinna — the visible, cartilaginous part of your ear — isn’t decorative. Its folds and ridges subtly modify the frequency spectrum of incoming sound depending on the direction the sound is coming from. Sound from above is shaped differently from sound from below, giving your brain cues about vertical location. The ear canal (about 2.5 cm long) acts as a resonant tube, amplifying frequencies near 3 kHz by about 10 dB — which happens to be the frequency range most important for understanding speech.

The eardrum. At the end of the canal, the tympanic membrane (eardrum) vibrates in response to pressure variations. It’s a thin, cone-shaped membrane about 10 mm in diameter. Despite its small size, it’s mechanically sophisticated: its conical shape and non-uniform thickness ensure that it vibrates efficiently across a wide range of frequencies without significant resonance peaks that would distort the sound.

The middle ear. Behind the eardrum, three tiny bones — the malleus (hammer), incus (anvil), and stapes (stirrup) — form a mechanical linkage that transmits the eardrum’s vibrations to the cochlea. The stapes is the smallest bone in the human body, about 3 mm tall.

These bones do two critical things. First, they act as an impedance-matching transformer. The cochlea is filled with fluid, and sound passing directly from air to fluid would lose about 99.9% of its energy (30 dB) due to the impedance mismatch — most of the sound would reflect off the fluid surface. The middle ear bones solve this by concentrating the force from the large eardrum (area ≈ 55 mm²) onto the much smaller oval window (area ≈ 3.2 mm²). This area ratio of about 17:1, combined with the lever action of the bones (about 1.3:1), produces a total pressure amplification of roughly 22:1 — enough to overcome the air-to-fluid impedance mismatch and transmit sound into the cochlea with minimal loss.

Second, two tiny muscles attached to the middle ear bones — the tensor tympani and the stapedius — contract reflexively in response to loud sounds, stiffening the chain and reducing transmission. This acoustic reflex protects the cochlea from damage, reducing sound transmission by about 10–15 dB. But it has a latency of about 25–150 milliseconds, so it can’t protect against sudden impulse sounds like gunshots — which is one reason impulse noise is particularly damaging.

The Cochlea: A Biological Fourier Analyser

The cochlea is where the physics gets remarkable.

It’s a fluid-filled, snail-shaped tube about 35 mm long when uncoiled, coiled into roughly 2.75 turns. Running along its length is the basilar membrane — a thin strip of tissue that varies systematically in its mechanical properties: narrow and stiff at the base (near the oval window), wide and flexible at the apex.

When the stapes pushes on the oval window, it creates a pressure wave in the cochlear fluid. This wave travels along the basilar membrane as a travelling wave. Here’s the key: because the basilar membrane’s stiffness decreases along its length, different frequencies cause maximum vibration at different positions.

High frequencies (say 20,000 Hz) peak near the base, where the membrane is stiff. Low frequencies (say 20 Hz) peak near the apex, where the membrane is flexible. Mid frequencies peak somewhere in between. The basilar membrane maps frequency to position — it performs a spatial frequency decomposition of the incoming sound.

This is, functionally, a Fourier transform. Any complex sound — a chord, a voice, traffic noise — is decomposed into its constituent frequencies, with each frequency activating a specific location along the basilar membrane. The Hungarian biophysicist Georg von Békésy won the 1961 Nobel Prize for demonstrating this travelling wave mechanism in cadaver cochleae.

But there’s a twist. Von Békésy’s passive travelling wave produced relatively broad, poorly resolved frequency peaks. Real, living ears achieve much sharper frequency resolution — about 100 times better than the passive mechanics predict. How?

Outer Hair Cells: The Cochlea’s Built-In Amplifier

Sitting on the basilar membrane is the organ of Corti, containing two types of hair cells: about 3,500 inner hair cells (which send signals to the brain) and about 12,000 outer hair cells (which amplify the basilar membrane’s vibration).

Each hair cell has a bundle of tiny projections called stereocilia on its top surface. When the basilar membrane vibrates, the stereocilia are deflected sideways. This deflection opens mechanically gated ion channels — literally, the mechanical movement pulls open tiny protein channels in the stereocilia membranes. Potassium and calcium ions rush in, depolarising the cell.

In inner hair cells, this depolarisation triggers the release of neurotransmitter, which activates the auditory nerve fibres. This is the signal that reaches the brain.

Outer hair cells do something different and extraordinary. When depolarised, they change length — contracting by about 4–5% of their length. This electromotility is driven by a unique motor protein called prestin, embedded in the outer hair cell membrane. When the cell’s voltage changes, prestin molecules change shape, and the cumulative effect of millions of prestin molecules causes the entire cell to contract or expand.

This contraction is fast — outer hair cells can change length at frequencies up to at least 70,000 Hz, making prestin one of the fastest known biological motors. And the effect is precisely tuned: each outer hair cell amplifies the basilar membrane’s vibration at its specific location, boosting the response to quiet sounds by about 40–60 dB (100 to 1,000-fold) while providing almost no amplification for loud sounds.

The result is a nonlinear amplifier built into the cochlea. Quiet sounds are amplified enormously; loud sounds are not. This active process explains both the ear’s extraordinary sensitivity at low sound levels and its ability to handle a trillion-fold range of intensities without saturating. It also sharpens the frequency resolution of the basilar membrane, giving the ear its ability to distinguish frequencies differing by less than 0.3%.

The outer hair cells are the ear’s most vulnerable component. They’re the first to be damaged by noise exposure, aging, and ototoxic drugs. When they die, the cochlear amplifier fails: hearing sensitivity drops by 40–60 dB, frequency resolution degrades, and the ability to understand speech in noisy environments collapses. This is the most common form of sensorineural hearing loss, and since mammalian hair cells don’t regenerate, the damage is permanent.

The Decibel Scale: Compressing a Trillion-Fold Range

The intensity range of human hearing — from the threshold of hearing to the threshold of pain — spans a factor of about 10¹² (one trillion). Working with numbers that large is impractical, so acoustics uses the decibel (dB) scale:

dB SPL = 20 × log₁₀(p/p₀)

where p is the sound pressure and p₀ = 20 µPa (the reference threshold). The decibel scale is logarithmic: every 20 dB represents a tenfold increase in pressure (100-fold increase in intensity). Every 10 dB represents a roughly twofold increase in perceived loudness.

Some reference points: a quiet room is about 30 dB. Normal conversation is about 60 dB. A busy street is 70–80 dB. A rock concert or power tool is 100–110 dB. Pain begins at about 120 dB. A jet engine at 30 metres is about 140 dB.

The logarithmic scale isn’t just mathematical convenience — it reflects how the ear actually works. The nonlinear compression by the outer hair cells means that perceived loudness grows roughly logarithmically with intensity. A sound 10 times more intense doesn’t sound 10 times louder — it sounds about twice as loud. The decibel scale captures this perception.

Sound Localisation: 10 Microseconds of Precision

One of the most impressive feats of auditory physics is sound localisation — your ability to determine where a sound is coming from.

The brain uses two primary cues, each effective for different frequency ranges.

Interaural time difference (ITD). For frequencies below about 1,500 Hz, the brain compares the arrival time of a sound at the two ears. Sound from the right reaches the right ear first. The maximum time difference (for a source directly to one side) is about 650 microseconds — the time it takes sound to travel the roughly 22 cm between ears (0.22 m / 343 m/s ≈ 640 µs). The brain can detect time differences as small as 10 microseconds, corresponding to angular resolution of about 1–2 degrees.

How does a neural system operating on millisecond timescales detect 10-microsecond differences? The answer involves specialised neurons in the medial superior olive that act as coincidence detectors — they fire only when signals from both ears arrive simultaneously. Different neurons have different built-in delays, so each neuron is “tuned” to a specific interaural delay. The pattern of firing across the array encodes the sound’s direction.

Interaural level difference (ILD). For frequencies above about 1,500 Hz, the head casts an acoustic shadow. High-frequency sounds have wavelengths shorter than the head’s diameter (the head is about 22 cm; at 1,500 Hz, the wavelength is about 23 cm), so the head significantly attenuates sound reaching the far ear. The brain compares loudness between the ears. At 6 kHz, the head shadow can produce a 20 dB difference between ears for sounds arriving from the side.

The crossover at about 1,500 Hz is not coincidental — it’s the frequency at which the wavelength roughly equals the head’s diameter. Below this, sound diffracts around the head easily (wavelength > head), making level differences negligible. Above this, the head blocks sound effectively (wavelength < head), making level differences significant. The physics of wave diffraction determines which localisation cue works at which frequency.

Vertical localisation uses a third cue: the direction-dependent filtering by the pinna. The folds of the outer ear create small echoes and resonances that modify the frequency spectrum differently depending on whether the sound comes from above, below, or behind. The brain learns these spectral cues — if you wear moulds that reshape your pinnae, your vertical localisation degrades for days until your brain recalibrates.

Cochlear Implants: Engineering a Replacement Ear

When the hair cells are destroyed — by noise, disease, or genetics — no hearing aid can help, because there’s nothing left to amplify. The hair cells were the transducers, and they’re gone.

Cochlear implants bypass the destroyed hair cells entirely. An electrode array is surgically threaded into the cochlea, with electrodes positioned at different locations along the basilar membrane. An external processor captures sound, decomposes it into frequency bands (typically 12–22 channels), and sends the information wirelessly to the implant. Each frequency band drives a specific electrode, which electrically stimulates the auditory nerve fibres at that location.

The principle exploits the cochlea’s frequency-to-place map: stimulating the base triggers high-frequency perception; stimulating the apex triggers low-frequency perception. The brain receives a pattern of nerve impulses that roughly mimics what it would receive from healthy hair cells.

The resolution is crude. Twelve to twenty-two electrodes replace 3,500 inner hair cells. It’s like replacing a high-resolution photograph with a grid of coloured blocks. But the brain is remarkably adaptive — most implant recipients learn to understand speech within months of activation. Many can use the phone. Some enjoy music, though pitch perception remains limited.

Over one million people worldwide use cochlear implants. The technology works because the cochlea’s physics — frequency mapped to position — provides a framework that even crude electrical stimulation can exploit. You don’t need 3,500 channels. You need enough channels, in the right places, for the brain to extract the pattern.

What the Ear Teaches Us About Physics

I find the ear endlessly instructive because it’s a physics instrument that evolution built, and the design choices it made reveal what matters physically.

The impedance-matching transformer of the middle ear solves a textbook wave physics problem — coupling energy between media with different impedances. The basilar membrane performs a spatial Fourier transform — the same mathematical operation that underpins signal processing, quantum mechanics, and image analysis. The outer hair cells implement a nonlinear amplifier with automatic gain control — the same function engineers build into radio receivers and audio processors. The binaural system exploits interference and diffraction for spatial localisation.

None of this was designed. It was evolved, over hundreds of millions of years, by organisms that needed to detect predators, locate prey, communicate with mates, and navigate in the dark. The physics problems the ear solves — impedance matching, frequency analysis, signal amplification, source localisation — are the same problems any acoustic engineer would face. Evolution converged on solutions that we recognise because the physics is the same whether the engineer is human or Darwinian.

Your ear is performing a real-time Fourier transform right now, decomposing the acoustic world into its constituent frequencies, amplifying the quiet ones, compressing the loud ones, and feeding the result to a brain that extracts meaning from the pattern. It does this with a few thousand cells, a membrane the length of a fingernail, and three bones smaller than grains of rice.

Physics at its most elegant doesn’t always happen in a laboratory. Sometimes it happens inside your head.

Frequently Asked Questions

How does the ear convert sound into nerve signals?

Sound waves enter the ear canal and vibrate the eardrum (tympanic membrane). Three tiny bones in the middle ear — the malleus, incus, and stapes (hammer, anvil, stirrup) — amplify this vibration and transmit it to the oval window, a membrane-covered opening to the fluid-filled cochlea. Inside the cochlea, the vibration creates a travelling wave along the basilar membrane — a flexible structure that varies in width and stiffness along its length. Different frequencies cause maximum vibration at different positions: high frequencies near the base (narrow, stiff), low frequencies near the apex (wide, flexible). At the point of maximum vibration, tiny hair cells on the basilar membrane are deflected. This deflection opens ion channels in the hair cell membranes, triggering an electrical signal that is transmitted via the auditory nerve to the brain. The entire process — from sound wave to nerve impulse — takes less than a millisecond.

What is the quietest sound a human can hear?

The threshold of human hearing at its most sensitive frequency (around 2-4 kHz) is about 0 decibels SPL, corresponding to a pressure variation of roughly 20 micropascals — about 0.00000002% of atmospheric pressure. At this threshold, the eardrum moves less than 10 picometres — smaller than the diameter of a hydrogen atom, and approaching the scale of thermal molecular motion. The ear is essentially operating at the physical limit of sensitivity: if it were any more sensitive, you would hear the random Brownian motion of air molecules hitting the eardrum. The ear achieves this sensitivity through mechanical amplification in the middle ear (about 20-fold) and active amplification by the outer hair cells in the cochlea, which act as biological motors that boost the vibration of the basilar membrane by about 100-fold (40 dB) at low sound levels.

How can we tell where a sound is coming from?

The brain determines sound direction using two main cues from the two ears. For low frequencies (below about 1,500 Hz), it uses interaural time difference (ITD) — the tiny difference in arrival time between the two ears. Sound from the right arrives at the right ear a few microseconds before the left ear; the brain detects timing differences as small as 10 microseconds (corresponding to about 1-2 degrees of angular resolution). For high frequencies (above about 1,500 Hz), it uses interaural level difference (ILD) — the difference in loudness between the ears, caused by the head casting an acoustic 'shadow' that attenuates high-frequency sound on the far side. The pinna (outer ear) also shapes the frequency spectrum of incoming sound depending on its elevation angle, providing cues for up-down localisation. The brain combines all these cues in the superior olivary complex and inferior colliculus, creating a three-dimensional spatial map of the sound environment in real time.

What causes hearing loss?

Hearing loss has several causes depending on which part of the auditory system is damaged. Conductive hearing loss results from problems in the outer or middle ear (earwax blockage, fluid behind the eardrum, otosclerosis — abnormal bone growth around the stapes) that prevent sound from reaching the cochlea efficiently. Sensorineural hearing loss results from damage to the cochlea's hair cells or the auditory nerve. The most common cause is noise exposure: sounds above 85 dB sustained over time damage hair cells, which in mammals do not regenerate. A single exposure above 140 dB (gunshot, explosion) can destroy hair cells instantly. Age-related hearing loss (presbycusis) involves gradual loss of hair cells starting at the base of the cochlea (high frequencies), which is why older adults typically lose high-frequency hearing first. Once hair cells are destroyed, that frequency range is permanently lost — there is currently no way to regenerate human cochlear hair cells, though research is active.

How do cochlear implants work?

Cochlear implants bypass damaged hair cells by directly stimulating the auditory nerve with electrical signals. An external microphone and processor capture sound, analyse it into frequency bands (typically 12-22 channels), and transmit the information wirelessly to an implanted receiver. The receiver sends electrical pulses through an electrode array inserted into the cochlea. Each electrode corresponds to a different position along the cochlea (and therefore a different frequency), mimicking the frequency-to-place mapping of normal hearing. High-frequency signals stimulate electrodes near the base; low-frequency signals stimulate electrodes near the apex. The auditory nerve fibres near each electrode fire in response to the electrical pulses, sending signals to the brain. The brain learns to interpret these signals as sound. With 12-22 electrodes replacing roughly 3,500 inner hair cells, the spectral resolution is crude compared to normal hearing, but sufficient for speech comprehension in most recipients after training. Over 1 million people worldwide use cochlear implants.

Read Next