How do film sound and Foley work?

Most of what you hear in a film was never recorded on set. Hang a boom mic just out of shot, snap celery for breaking bones, build a monster's roar from three layers, re-voice a line in sync, and fly a helicopter over your head in Dolby Atmos.

FilmSoundClearOpened 31 Aug 202615 min to playFree · no sign-up

In 60 seconds

  1. Voices on set

    A shotgun mic on a boom hangs just above the frame, as close to the mouth as the shot allows. Sound power falls with the square of distance: twice as far is 6 dB quieter, while the background noise stays the same. Actors also wear hidden lavalier mics, and the crew records 30 seconds of room tone.

  2. Foley

    Footsteps, rustling clothes and small props are added later by Foley artists who watch the film and perform in sync, named after Jack Foley at Universal from 1929. Celery snaps make breaking bones, coconut shells make hooves, and a punched cabbage makes a punch.

  3. Sound design

    Sounds that cannot be recorded are built from layers: a low rumble for weight, a mid growl for character and a high screech for bite. Designers pitch-shift, reverse and move them. A passing ship rises then falls in pitch: the Doppler effect.

  4. ADR and dubbing

    Noisy or weak lines are re-recorded in a booth to the picture: three beeps, then speak on the fourth. Viewers notice sound more than about 45 ms early or 125 ms late. Dubbing re-voices a whole film in another language, and India's biggest hits are released in five or more languages.

  5. The mix

    On a dubbing stage, mixers balance five stems: dialogue, music, effects, Foley and ambience, ducking the music under the words. Streaming targets about −27 LUFS of dialogue; cinemas are calibrated instead. 5.1 and 7.1 feed fixed speakers, while Dolby Atmos places sound objects anywhere, even overhead.

  6. In the cinema

    The film arrives as a Digital Cinema Package with up to 16 uncompressed channels. A processor and amplifiers drive speakers that stand behind a screen full of holes about 1.2 mm wide. Sound takes about 3 ms per metre, so the back row hears a little later, and phone speakers lose the deep bass.

The history

A century from pianists in the dark to helicopters flying over your head in Dolby Atmos.

Read the full history
  1. 1927The Jazz Singer talks
  2. 1931India's first talkie: Alam Ara
  3. 1977Star Wars and Dolby Stereo
  4. 1992Digital sound in the cinema
  5. 2012Dolby Atmos: sound as objects

The full explanation

FilmSoundClear, chapter by chapter

Chapter 1

Recording the voices on set

Aim a boom mic, clip on a lavalier, and see why the mic hangs just out of frame.

On a film set the production sound team records the actors’ voices. The picture and sound go to separate machines, and FilmClear shows how the clapperboard lines them up again. Here we look at the microphones.

The workhorse is a shotgun mic on a boom pole. A shotgun hears best straight ahead and ignores much of the sound from the sides: its polar pattern is a long lobe, drawn here in 3D. Aim it at the mouth. Aim it at the chest and the voice drops. Actors also wear a tiny lavalier (lav) mic hidden in their clothes, sending sound by radio. It is always close, but it can catch rustling cloth.

Distance is everything. Sound spreads out as it travels, so its power falls with the square of the distance: twice as far, a quarter of the power, 6 dB quieter. The background noise stays the same, so the signal-to-noise ratio (SNR) drops. That is why the boom operator hangs the mic as low as it can go without appearing in the shot. A close-up lets the mic come close. A wide shot pushes it far away, which is when lavs earn their keep.

Before a scene wraps, everyone freezes for 30 seconds so the recordist can capture room tone: the “silence” of that room. Editors use it to fill gaps so the background never drops out. Noise has always been the enemy: India’s first sound film, Alam Ara (1931), was shot between about 1 and 4 in the morning to dodge the trains running past the studio. How we hear it all is in EarClear, and what a sound wave is in WaveClear.

Try “On set” in the interactive model →

Chapter 2

Foley: sounds made by hand

Walk in the pits, snap celery for bones, and try to perform footsteps in sync.

Most of the small sounds in a film were not recorded on set. The boom mic was pointed at the actors’ mouths, so footsteps, rustling clothes and the clink of a cup are faint or missing. They are added later by Foley artists, who watch the film on a big screen and perform the sounds live, in sync with the picture.

The craft is named after Jack Foley, who did this at Universal Studios from 1929. A Foley stage has pits of gravel, wood, concrete and more, and shelves of odd props. The trick is that the real thing often sounds wrong on film. A snapped stick of celery makes a better breaking bone than a bone. Coconut shells knocked together make hooves. A fist into a cabbage is a punch, and a twisted leather jacket gives creaks of tension.

A Foley team records one layer at a time: footsteps for each character, then a cloth pass for every movement, then the props. Each is placed to the exact frame. At 24 frames a second that is 41.7 milliseconds, and viewers notice when a step lands more than about 45 ms early or 125 ms late.

Try “Foley” in the interactive model →

Chapter 3

Designing big sounds

Build a monster’s roar from three layers, bend its pitch, and fly a spaceship past your ears.

Dragons, spaceships and giant robots do not exist, so nobody can record them. A sound designer builds their sounds from pieces. The secret is layering: a low layer for weight, a mid layer for character, and a high layer for bite. Each alone sounds like nothing much. Stacked, they become a beast.

Designers bend sounds too. Play a sound faster and its pitch rises, like a speeded-up recording: 12 semitones up is twice as fast and an octave higher. Slow a small animal down and it becomes huge. Play a sound backwards and a crash becomes an eerie swell. The pieces come from sound libraries of thousands of ready-made clips, or are recorded fresh for the film. Walter Murch was first credited as sound designer on Apocalypse Now (1979), and Ben Burtt made Star Wars’ lightsaber from a projector’s hum and a TV’s buzz.

Movement needs physics. When a ship rushes towards you, each wave crest leaves from a little closer, so the crests arrive squashed together: higher pitch. As it flies away they are stretched: lower pitch. That drop is the Doppler effect (more in WaveClear). Want to build tones from scratch? That is SynthClear.

Try “Sound design” in the interactive model →

Chapter 4

Re-recording the dialogue: ADR

Loop a scene, count the beeps, and find out how far lips and words can drift apart.

Sometimes the words recorded on set cannot be used. A plane flew over, a generator hummed, or the director wants the line said differently. Then the actor goes into a quiet booth for ADR, automated dialogue replacement, also called looping or dubbing.

The scene plays on a screen in a loop. The actor hears three beeps, evenly spaced, while a line called a streamer wipes across the picture. On the silent fourth beat they speak, watching their own lips and matching them. Good ADR artists can hit a line within a frame or two.

How close is close enough? Broadcast tests found most people notice when sound is about 45 ms early or 125 ms late. We forgive late sound more, because in real life light always reaches us before sound. The clapperboard in FilmClear exists for the same reason.

The biggest ADR job is dubbing a whole film into another language. India is a giant of dubbing: Telugu and Tamil films like Baahubali and RRR were released in Hindi, Tamil, Telugu, Malayalam and Kannada, and Hollywood films arrive in Hindi, Tamil and Telugu. Dubbing writers hunt for words that fit the lips: when the actor’s lips close, the new line needs a p, b or m there too.

Try “Dialogue & ADR” in the interactive model →

Chapter 5

The final mix: stems, loudness and Atmos

Ride five faders, duck the music under the words, and fly a helicopter around the room.

All the sound comes together on a dubbing stage: a small cinema with a huge mixing console in the middle. The mixers keep the sound in groups called stems: dialogue, music, effects, Foley and ambience. Keeping them apart lets them make a music-and-effects version with no words, ready for dubbing into other languages.

The golden rule is that you must hear the words. When someone speaks, the music is turned down a few decibels and comes back up after: ducking. Loudness is measured in LUFS, a scale built to match how loud things feel. TV in Europe aims for −23 LUFS. Netflix aims for −27, measured only while people talk. Cinemas are different: the whole room is calibrated to a standard level, so a film plays as loud as the mixers heard it.

Then the sound is placed in space. 5.1 means five speakers (left, centre, right and two surrounds) plus .1, a subwoofer channel for rumbles. 7.1 splits the surrounds into sides and rears. Dolby Atmos, first used for Brave in 2012, adds speakers in the ceiling and treats a sound as an object with a 3D position. The cinema’s processor decides which speakers play it. How speakers make sound at all is in MagnetismClear.

Try “Music & the mix” in the interactive model →

Chapter 6

Hearing it in the cinema

Follow the sound from the digital package to your seat, and see why the screen is full of holes.

A finished film reaches the cinema as a Digital Cinema Package (DCP): a set of files. Its sound is up to 16 channels of uncompressed audio, 48,000 samples a second, plus an Atmos track. A cinema processor decodes it, places the Atmos objects on this room’s speakers, and corrects for the room. Amplifiers then drive the speakers. The room is calibrated so a test tone plays at a standard level, and full volume can reach about 105 decibels per speaker.

Where do the front speakers go? Behind the screen, so voices come from the actors’ mouths. The screen is perforated: tens of thousands of holes about 1.2 mm wide per square metre, only around 5% of its area. The rest still reflects the picture, and the sound passes through the holes. From your seat they are too small to see.

Sound is slow. At 343 metres a second, it takes about 3 ms to cross each metre, so the back row hears the screen later than the front. Big halls stay under the 125 ms that viewers notice. At home the same film plays on TV, laptop or phone speakers that cannot make deep bass, so mixers often make a gentler home mix. How your ears catch it all is in EarClear.

Try “In the cinema” in the interactive model →

Test yourself

Frequently asked

The boom mic moves from 50 cm to 1 m from the actor. The voice at the mic gets…

about 6 dB quieter. Twice the distance spreads the sound over four times the area, a quarter of the power: 6 dB less.

Why does the boom operator hang the mic just above the frame line?

It is the closest the mic can get without being in the shot. Closer means a louder voice against the same noise. The frame edge is the limit.

What is room tone for?

Filling gaps in the edit with the room’s own background sound. Every room has its own quiet hum. Editors lay it under cuts so the background never drops out.

Why are footsteps usually added after filming?

The boom mic points at the voices, so steps are faint, and Foley gives clean control. Production sound focuses on dialogue. Foley recorded later is clean and can be shaped for each moment.

What does a Foley artist often use to make the sound of breaking bones?

A stick of celery. Snapped celery gives a crisp crack and tearing fibres that sound like bone on screen.

At 24 fps, a footstep that is 2 frames late is late by about…

83 ms. One frame is 1/24 s ≈ 41.7 ms, so two frames ≈ 83 ms, still under the ~125 ms most people notice.

Why build a monster roar from several layers?

Each layer adds a different part: weight, character and bite. Low, mid and high layers fill different parts of the spectrum and together sound huge.

A sound is played twice as fast. What happens?

Pitch rises an octave and it lasts half as long. Every frequency doubles (an octave, 12 semitones) and the sound takes half the time.

A ship flies past you. Its engine sounds…

higher coming, lower going. Waves bunch up ahead of a moving source and stretch out behind it: the Doppler effect.

Why might a perfect acting take still need ADR?

Noise, like a plane overhead, buried the words. The picture is kept, and the words are re-recorded clean in a quiet booth.

Which sync error do viewers notice sooner?

Sound 60 ms early. We notice sound about 45 ms early, but tolerate up to about 125 ms late, because light beats sound in real life.

What do the three beeps in an ADR loop tell the actor?

When to start: the line begins on the silent fourth beat. They count the actor in, so the first word lands exactly on the lips.

Why is the music ducked under dialogue?

So the words can be heard clearly. Dialogue carries the story. Dipping the music a few dB while people speak keeps every word clear.

What does the “.1” in 5.1 mean?

A low-frequency effects channel for deep rumbles. The LFE channel feeds the subwoofers with sounds below about 120 Hz.

What makes Dolby Atmos different from 7.1?

Sounds are objects with a 3D position, and there are ceiling speakers. Each object carries its position; the cinema’s processor picks which speakers, including overhead ones, play it.

Why is a cinema screen full of tiny holes?

So sound from speakers behind it can pass through. The left, centre and right speakers sit behind the screen, so voices seem to come from the actors. The holes let the sound through.

You sit 30 m from the screen. The sound reaches you about…

87 ms after the picture. Sound travels about 343 m a second: 30 / 343 ≈ 0.087 s. Light takes a tenth of a microsecond.

Why does a big explosion sound weak on a phone?

Tiny phone speakers cannot make the deep bass that gives it weight. A coin-sized speaker produces little below a few hundred hertz, where most of the rumble lives.

Words worth knowing

Production sound
The dialogue and sound recorded on set while filming.
Polar pattern
How sensitive a microphone is in each direction; a shotgun hears mostly straight ahead.
Inverse-square law
Sound power falls with the square of distance: 6 dB less for each doubling.
Foley
Everyday sounds performed live in sync with the picture after filming.
Layering
Stacking several sounds so they play as one bigger sound.
ADR
Automated dialogue replacement: re-recording lines in a studio to match the lips on screen.
Stem
One group of sounds in a mix, such as all the dialogue or all the music.
LUFS
A measure of how loud a mix feels, used to set loudness targets.
Dolby Atmos
Cinema sound with ceiling speakers where each sound is an object with a 3D position.

Fork it. Teach with it.

This box is plain HTML, CSS and JavaScript, with no build step and no accounts. Run it yourself and it sends nothing anywhere. The code is MIT. The words, images and videos are CC BY 4.0, so you can reuse them anywhere if you credit “Glassbox, glassbox.how/e/filmsoundclear”.

git clone https://github.com/bdeeps/filmsoundclear.git

Built with three.js (MIT), Geist, Instrument Serif (SIL OFL 1.1).

←→ previous / next box · / search

Would you like to see the full page, with the interactive model?