The Physics of Vision: How Your Eye Captures Photons and Your Brain Builds Reality

Your retina detects single photons. Your lens adjusts focal length in milliseconds. Your brain processes 10 million bits of visual data per second — all using optics, photochemistry, and neural computation that no camera can match. Here's the physics of seeing.

Table of Contents

Catching Light

Right now, as you read these words, your eyes are doing something that no camera on Earth can replicate.

Photonselectromagnetic radiation in a narrow band between 380 and 700 nanometres — are entering your pupil at 300,000 kilometres per second. Your cornea and lens bend them to a focus on your retina. There, 130 million photoreceptor cells convert photon energy into electrochemical signals through a photochemical reaction so sensitive that a single rod cell can detect a single photon.

From this stream of neural impulses, your brain constructs a three-dimensional, full-colour, motion-tracked, object-recognised model of the world — in real time, with a processing delay of about 100 milliseconds.

And you do all of this without thinking about it. Which, when you stop to consider the physics involved, is almost absurd.

The Optics: A Two-Lens System in 24 Millimetres

The human eye is an optical instrument. Not a metaphorical one — a literal one, with refracting surfaces, a variable aperture, and a curved focal plane. And it fits in a sphere roughly 24 mm in diameter.

Light enters through the cornea — the transparent, curved front surface. The cornea is where most of the eye’s focusing happens. When light crosses the boundary from air (refractive index n = 1.00) to corneal tissue (n = 1.376), it bends sharply. The cornea provides about 43 dioptres of refracting power — roughly two-thirds of the eye’s total.

Behind the cornea is the aqueous humour (a clear fluid, n ≈ 1.336), then the iris (the coloured diaphragm that controls pupil diameter from about 2 mm in bright light to 8 mm in darkness), and then the crystalline lens.

The lens is a remarkable structure. It’s a transparent, biconvex body about 9 mm in diameter and 4 mm thick, made of tightly packed protein fibres (crystallins) arranged in concentric layers like an onion. It has a gradient refractive index — higher in the centre (n ≈ 1.41) than at the edges (n ≈ 1.38) — which reduces spherical aberration and improves image quality in a way that a uniform lens cannot match.

The lens adds about 15–20 dioptres of power, for a total of about 60 dioptres in the relaxed eye. That gives a focal length of approximately 17 mm — remarkably short. A camera lens of similar image quality would be considerably larger.

Accommodation: Auto-Focus in Milliseconds

The lens does something no fixed camera element can: it changes shape.

Accommodation is the process by which the eye adjusts focus for objects at different distances. The lens is suspended by thin fibres (zonules) attached to the ciliary muscle — a ring of smooth muscle encircling the lens. When the ciliary muscle relaxes, the zonules pull taut, flattening the lens for distant vision. When the ciliary muscle contracts, the zonules slacken, and the elastic lens bulges into a more convex shape — increasing its power by up to 12 dioptres in a young eye. This shifts the focal point forward, bringing near objects into focus.

The response time is about 350 milliseconds — fast enough that you barely notice the adjustment when you glance from a distant tree to the book in your hand. The control loop involves blur detection by the retina, processing in the visual cortex, and motor commands back to the ciliary muscle. It’s a closed-loop feedback system, and it works astonishingly well.

Until it doesn’t. Presbyopia — the gradual loss of accommodation with age — happens because the lens proteins cross-link and stiffen over decades. By age 45, most people have lost enough elasticity that close reading becomes difficult. By age 60, accommodation is nearly zero. The physics hasn’t changed; the material has.

Aberrations and Corrections

No optical system is perfect, and the eye is no exception. The eye exhibits several optical aberrations:

Spherical aberration: rays passing through the edge of the lens focus at a different point than rays through the centre. The gradient-index structure of the lens partially compensates for this — one of evolution’s neater tricks.

Chromatic aberration: different wavelengths focus at different distances (because the refractive index depends on wavelength). The eye has about 2 dioptres of longitudinal chromatic aberration — blue light focuses about 0.5 mm in front of red light. The brain compensates by using mainly the green channel (M-cones) for fine detail, since green is near the middle of the visible spectrum and closest to the focal plane.

Astigmatism: the cornea is not perfectly spherical. Most people have slight corneal astigmatism, and if it exceeds about 0.5 dioptres, a cylindrical correction (glasses or contact lenses) sharpens the image noticeably.

The fact that vision works as well as it does, given these uncorrected aberrations, is a testament to neural processing. The brain doesn’t just passively receive an image — it actively cleans it up.

The Retina: A Neural Computer at the Back of Your Eye

The retina is not a passive film. It’s a thin (about 0.2 mm), multi-layered neural tissue lining the back of the eye, and it performs serious computation before any signal reaches the brain.

Counterintuitively, the photoreceptors are at the back of the retina, facing away from the incoming light. Light must pass through several layers of neurons and blood vessels before reaching the light-sensitive cells. This inverted arrangement exists because the photoreceptors need to be in contact with the retinal pigment epithelium (RPE) — a layer of cells that recycles the photopigment and removes metabolic waste. It’s a design that works, but it’s the kind of design that screams evolutionary history rather than intelligent engineering.

Rods and Cones: Two Systems in One Retina

The retina contains two types of photoreceptor:

Rods — about 120 million of them — are responsible for dim-light (scotopic) vision. They are extraordinarily sensitive: a single rod can detect a single photon. Rods are concentrated in the peripheral retina and are absent from the very centre (the fovea). They provide no colour information — all rods contain the same photopigment (rhodopsin, peak absorption at 498 nm). Rod vision is monochromatic, which is why you can’t distinguish colours in very dim light.

Cones — about 6–7 million — are responsible for daylight (photopic) vision and colour perception. They are concentrated in the fovea — a small pit about 1.5 mm in diameter at the centre of the retina, directly behind the lens. In the very centre of the fovea (the foveola, about 0.35 mm across), only cones are present, and the overlying neural layers are pushed aside so light reaches the cones with minimal scattering. This is why you see fine detail best when you look directly at something — you’re aiming the fovea.

The numbers are striking: 120 million rods and 7 million cones, but only about 1.2 million nerve fibres in the optic nerve. The retina compresses the signal by a factor of roughly 100:1 before sending it to the brain. This compression is not simple averaging — it involves sophisticated neural processing including edge detection, contrast enhancement, and motion sensitivity, all performed by the retinal circuitry before the signal leaves the eye.

The Photochemistry: How a Photon Becomes a Signal

The conversion of light into a neural signal — phototransduction — is one of the most elegant signal amplification cascades in biology.

Each rod cell contains about 100 million molecules of rhodopsin — a protein consisting of the transmembrane protein opsin bound to a small molecule called 11-cis-retinal (a derivative of vitamin A). The 11-cis-retinal is the chromophore — the part that actually absorbs light.

When a photon of the right energy (around 500 nm wavelength, approximately 2.5 eV) is absorbed by 11-cis-retinal, it undergoes a conformational change: the molecule isomerises from the bent cis configuration to the straight trans configuration. This happens in about 200 femtoseconds — one of the fastest chemical reactions known. The quantum yield is about 0.67, meaning roughly two out of three absorbed photons trigger the isomerisation.

This single molecular change sets off a biochemical amplification cascade:

The isomerised rhodopsin (now called metarhodopsin II) activates a G-protein called transducin. One activated rhodopsin molecule catalyses the activation of about 500 transducin molecules over the course of a second.

Each activated transducin activates a molecule of phosphodiesterase (PDE), an enzyme that breaks down cyclic GMP (cGMP).

In the dark, cGMP keeps ion channels in the rod cell membrane open, allowing sodium and calcium ions to flow into the cell (the “dark current”). When PDE destroys cGMP, the channels close, and the cell hyperpolarises — its membrane voltage drops from about −40 mV to about −70 mV.

The amplification is extraordinary. One photon → one rhodopsin → 500 transducin → 500 PDE → destruction of about 100,000 cGMP molecules → closure of hundreds of ion channels → a measurable voltage change of about 1 millivolt. A single photon produces a detectable electrical signal. That’s what it means when physicists say the eye operates at the quantum limit.

Recovery: Resetting the System

After activation, rhodopsin must be deactivated and the 11-cis-retinal must be regenerated. This happens through a multistep process: rhodopsin kinase phosphorylates the activated rhodopsin, arrestin binds to it and blocks further transducin activation, and the all-trans-retinal is enzymatically converted back to 11-cis-retinal in the RPE. The whole recovery cycle takes seconds to minutes, which is why you see afterimages when you stare at a bright light — the photoreceptors in that patch of retina are temporarily desensitised.

Dark adaptation — the gradual increase in sensitivity when you move from bright light to darkness — takes about 30–40 minutes for full rod adaptation. This is largely the time needed to regenerate rhodopsin from the bleached (all-trans) state. Cone adaptation is faster (about 5–7 minutes) because cones have a more rapid pigment regeneration cycle.

Colour Vision: Three Filters, Millions of Colours

Colour is not a property of light. Colour is a property of perception — a construction of the brain based on the relative signals from three types of cone photoreceptor.

The three cone types are:

S-cones (short wavelength): peak absorption at ~420 nm (blue-violet). These make up only about 2% of all cones and are absent from the very centre of the fovea.

M-cones (medium wavelength): peak absorption at ~530 nm (green).

L-cones (long wavelength): peak absorption at ~560 nm (yellow-green, perceived as contributing to “red” perception).

Notice that the L-cone peak is not at 600+ nm (what we’d call red). The naming convention — “red, green, blue” — reflects the perceptual categories, not the actual peak wavelengths. The M and L cone spectra overlap enormously; the difference in their peaks is only about 30 nm. Yet from this small spectral difference, the brain extracts the entire red-green colour axis.

Trichromacy: Why Three Is Enough

Thomas Young proposed in 1802 that three receptor types are sufficient for full colour vision. Hermann von Helmholtz developed the theory further. The logic is beautifully simple: any spectral distribution of light stimulates the three cone types in a particular ratio (S:M:L). Any other spectral distribution that produces the same ratio will be perceived as the same colour. These physically different but perceptually identical stimuli are called metamers.

This is why your phone screen works. The screen produces only three wavelengths (red, green, blue phosphors), but by mixing their intensities, it can produce any S:M:L ratio that natural light produces. Your brain can’t tell the difference between “white” from a mixture of all wavelengths (sunlight) and “white” from three narrow spectral peaks (RGB display) — because the cone ratios are the same.

Trichromacy also explains why some colours have no single wavelength equivalent. Purple, for example, stimulates both S-cones and L-cones without stimulating M-cones — a pattern that no single wavelength can produce. Purple exists only as a neural construct.

Opponent Processing: The Second Stage

Raw cone signals are not sent directly to the brain. The retina and lateral geniculate nucleus (LGN) recode them into opponent channels:

Luminance channel: L + M (the sum of long and medium cone signals — essentially brightness).

Red-green channel: L − M (the difference — sensitive to the red-green axis).

Blue-yellow channel: S − (L + M) (blue-cone signal versus the sum of the other two).

This opponent processing, proposed by Ewald Hering in 1892 and later confirmed neurophysiologically, explains several perceptual phenomena: why you can’t perceive “reddish green” or “yellowish blue” (these are opponent pairs and cancel), why afterimages are in the complementary colour (fatigue in one opponent channel shifts the balance to the other), and why yellow looks like a primary colour despite being coded by the combination of L and M signals.

Visual Acuity: The Limits of Resolution

How fine a detail can you see? Visual acuity is limited by three factors: the optics (diffraction and aberrations), the photoreceptor spacing, and neural processing.

Diffraction limit. For a pupil diameter of 2 mm (bright light), the Airy disc angular diameter is about 1.1 arcminutes at 550 nm. This sets the theoretical maximum resolution. At a 5 mm pupil, diffraction improves but aberrations worsen, so the optimum is around 2–3 mm.

Photoreceptor spacing. In the foveola, cones are packed at a centre-to-centre spacing of about 2.5 micrometres, corresponding to about 0.5 arcminutes per cone. The Nyquist sampling theorem requires at least two receptors per cycle of a grating to resolve it, so the cone spacing limits acuity to about 1 arcminute — which is almost exactly the standard definition of 20/20 (or 6/6) vision. The match between the optical diffraction limit and the receptor spacing is not coincidental — evolution has optimised both to approximately the same resolution.

Neural processing can enhance effective acuity beyond the raw receptor limit through hyperacuity mechanisms. Vernier acuity — the ability to detect that two lines are slightly offset — can reach 5–10 arcseconds, about 5–10 times better than the cone spacing would predict. This is possible because the brain analyses the spatial pattern of cone signals across the boundary, effectively interpolating between receptors.

The Blind Spot and Saccades: What You Don’t See

Your visual system has some dramatic gaps that you never notice.

The blind spot is a region about 6° × 8° in the temporal visual field of each eye (about 15° from the fovea) where the optic nerve exits the retina. There are no photoreceptors here — it’s literally a hole in the image. You don’t see it because the brain fills in the gap using information from surrounding areas and from the other eye. This is not a metaphor: the visual cortex fabricates visual content for the blind spot region, and you cannot distinguish the fabricated content from the real thing.

Saccades are rapid eye movements (up to 500°/sec) that jump the fovea from one fixation point to another about 3–4 times per second. During a saccade, the retinal image is a useless blur — yet you don’t perceive any blurring or blackout. The brain suppresses visual processing during saccades (saccadic suppression) and stitches together the pre- and post-saccade images into a seamless visual experience. You are functionally blind for a significant fraction of your waking hours, and you have no idea.

From Retina to Cortex: Building the World

Signals leave the retina through the optic nerve (about 1.2 million axons per eye), cross partially at the optic chiasm (where fibres from the nasal half of each retina cross to the opposite side, so each brain hemisphere receives information from the opposite visual field), and arrive at the lateral geniculate nucleus (LGN) of the thalamus.

From the LGN, signals travel to the primary visual cortex (V1) in the occipital lobe. V1 contains about 200 million neurons — compared to 1.2 million optic nerve fibres — which tells you that the cortex is expanding and reprocessing the signal, not simply receiving it.

V1 neurons are organised into orientation columns: each neuron responds best to a line or edge at a specific angle. David Hubel and Torsten Wiesel discovered this organisation in the 1960s (Nobel Prize 1981), showing that the cortex decomposes the visual image into oriented edges and spatial frequencies — a kind of local Fourier analysis of the image.

Beyond V1, visual information splits into two major processing streams:

The ventral stream (“what pathway”) runs from V1 through areas V2 and V4 to the inferior temporal cortex. It processes object identity, colour, and form. Damage to the ventral stream can cause visual agnosia — the inability to recognise objects despite intact optics.

The dorsal stream (“where pathway”) runs from V1 through area MT/V5 to the posterior parietal cortex. It processes motion, spatial relationships, and visually guided action. Damage here can cause optic ataxia — the inability to reach accurately for objects despite seeing them clearly.

The total amount of visual processing in the brain is enormous. Estimates suggest that about 30% of the cerebral cortex is involved in visual processing in some capacity — more than any other sense. The visual system doesn’t just detect light. It constructs a model of the world.

What the Eye Teaches Us About Physics

The human eye is not the best optical instrument we’ve built. Modern telescopes, microscopes, and cameras exceed it in resolution, sensitivity, spectral range, and recording capability. What the eye demonstrates is something different: the power of integrated design under constraints.

The eye packs a complete optical system, a focal plane array of 130 million detectors, a preprocessing neural computer, and an auto-focus mechanism into 7 cubic centimetres, powered by about 10 milliwatts. It operates continuously for decades, self-repairs most minor damage, and adapts its sensitivity over a dynamic range of about 10 billion to one — from starlight to tropical noon.

The physics involved spans optics (refraction, diffraction, aberration), photochemistry (rhodopsin isomerisation), electromagnetism (ion channel currents, membrane potentials), thermodynamics (dark noise from thermal isomerisation), signal processing (retinal neural computation), and quantum mechanics (single-photon detection at the physical limit).

What I find most remarkable is the single-photon sensitivity. The rod cell operates at the absolute boundary set by physics. You can’t detect less than one photon. And a rod cell detects exactly one. Every improvement beyond that point — from retinal neural processing to cortical object recognition — is software, built on a hardware platform that already operates at the quantum limit.

The next time you glance at a sunset and see ten million colours blending across the sky, remember: it starts with one photon, one molecule, one conformational change taking 200 femtoseconds. And 100 milliseconds later, you see the world.

Frequently Asked Questions

How does the human eye focus light?

The eye uses a two-element optical system. The cornea — the curved, transparent front surface — provides about two-thirds of the eye's total refracting power (roughly 43 dioptres) because light crosses a large refractive index boundary from air (n = 1.00) to corneal tissue (n = 1.376). Behind the cornea, the crystalline lens (about 15-20 dioptres) provides fine focus adjustment through accommodation: ciliary muscles change the lens shape, increasing its curvature for near objects (up to +12 dioptres of additional power in young eyes) and relaxing it for distant objects. Together, cornea and lens focus light onto the retina — a curved image surface at the back of the eye, about 17 mm behind the lens. The total optical power of the relaxed eye is about 60 dioptres, giving it a focal length of approximately 17 mm. This is a remarkably compact system: achieving equivalent image quality in a camera requires significantly more complex optics.

Can the human eye really detect single photons?

Yes. Experiments dating back to Hecht, Shlaer, and Pirenne in 1942 showed that human subjects could reliably detect flashes containing as few as 5-7 photons reaching the retina. Since not every photon is absorbed by a rod photoreceptor (roughly 10% absorption efficiency for a flash), this means individual rod cells can respond to a single photon. Modern experiments have confirmed this: a single photon triggers the isomerisation of one rhodopsin molecule, which activates a biochemical amplification cascade that produces a measurable electrical signal. The signal-to-noise ratio of a single rod responding to one photon is about 3:1 — enough for the cell to distinguish signal from thermal noise. However, the brain applies a threshold filter requiring several rods to fire nearly simultaneously before registering a conscious perception, which reduces false alarms from thermal isomerisation events (about one per rod per 400 seconds).

How do we see colour?

Colour vision relies on three types of cone photoreceptors in the retina, each containing a different photopsin protein tuned to absorb light in a different wavelength range: S-cones (short wavelength, peak ~420 nm, perceived as blue-violet), M-cones (medium wavelength, peak ~530 nm, green), and L-cones (long wavelength, peak ~560 nm, yellow-green, perceived as red). The brain determines colour not from the absolute signal of any single cone type but from the ratio of signals across the three types — a process called trichromatic colour vision, first proposed by Thomas Young in 1802 and developed by Hermann von Helmholtz. Any visible colour can be matched by a suitable mixture of three primary colours because any spectral distribution stimulates the three cone types in a particular ratio, and any other distribution that produces the same ratio looks identical (a metamer). This is why RGB displays work: three phosphor wavelengths can simulate millions of perceived colours.

What causes colour blindness?

Most colour blindness results from genetic variations in the photopsin genes on the X chromosome, which is why it predominantly affects males (about 8% of men vs. 0.5% of women). The most common form is red-green colour blindness: the L-cone and M-cone photopsins are encoded by adjacent genes on the X chromosome, and their high sequence similarity (96% identical) makes them prone to unequal crossover during meiosis. This can delete one gene entirely (dichromacy — the person has only two cone types) or produce a hybrid gene whose protein has a shifted absorption spectrum (anomalous trichromacy — all three cone types present but with abnormal spectral separation). Protanopia (missing L-cones) and deuteranopia (missing M-cones) are the dichromatic forms. In both cases, the physics of photon absorption is unchanged — the same wavelengths arrive at the retina — but the neural comparison between cone types is impaired, reducing the dimensionality of colour space from three to two.

Why can't we see ultraviolet or infrared light?

The visible range (roughly 380-700 nm) is set by the absorption properties of the photopigments and the transmission properties of the eye's optics. The cornea and especially the crystalline lens absorb strongly below about 400 nm, blocking most ultraviolet light before it reaches the retina — this protects the retina from UV damage. People who have had their lens removed (aphakia, often after cataract surgery) can perceive light down to about 310 nm because the UV filter is gone, confirming that the retina itself has some UV sensitivity. At the long-wavelength end, the photopsins simply do not absorb efficiently beyond about 700 nm: the photon energy is too low to trigger the conformational change in retinal (the chromophore) that initiates the visual signalling cascade. Infrared photons lack the roughly 1.8 eV minimum energy needed to isomerise 11-cis retinal to all-trans retinal. So the visible range is bounded by optical filtering on the short end and photochemistry on the long end.

Read Next