Updated: September 10, 2026
Viewers forgive a soft picture and leave over bad sound
It is one of the more reliable observations in video: an audience will tolerate a shot that is slightly out of focus, poorly framed or unevenly graded, and will click away within seconds from audio that is muffled, uneven or too quiet. Sound is processed differently. Bad audio feels like effort, and effort is what makes people stop watching.
Despite that, audio post is routinely the stage that gets compressed when a deadline tightens. The picture edit runs long, the grade takes a day, and the mix becomes an hour of nudging levels before export. This is a look at what audio post should actually involve, in what order, and what standards the finished file needs to meet.
What audio post covers
Audio post is not one task. It is a sequence, and doing them out of order wastes time because each stage changes what the next one is working with.
1. Editing and cleanup
Removing unwanted noise between words, tightening pauses, taking out lip smacks and breaths that distract, and dealing with obvious problems such as a bumped microphone. This is manual work and it is where most of the improvement happens.
2. Noise reduction
Reducing continuous background noise: air conditioning, traffic hum, room tone. Modern tools are remarkably good, and the temptation is to over-apply them. Aggressive noise reduction produces a hollow, watery voice that sounds worse than the original noise.
3. Equalisation
Shaping the tone of each voice. Rolling off unnecessary low frequencies removes rumble, and a gentle lift in the presence range improves clarity. The aim is a voice that sounds like the person, not one that sounds processed.
4. Dynamics
Compression to even out the difference between loud and quiet passages, so a speaker who trails off at the end of sentences stays audible without the loud parts becoming harsh.
5. Balance and music
Setting the relationship between dialogue, music and effects, and ducking music under speech so the words always win. This is where most amateur mixes fail, with music that feels right to the editor and buries the dialogue for everyone else.
6. Loudness and delivery
Bringing the whole programme to the level the destination platform expects, and checking peaks. This is a measurement step, not a taste step.
The problems that show up on Miami shoots
| Problem | Cause | What post can do |
|---|---|---|
| Air conditioning hum | units running during recording | reduces well with careful processing |
| Traffic and sirens | urban exteriors | steady hum reduces; sirens usually need a cut |
| Wind noise | coastal locations | limited; heavy wind is often unrecoverable |
| Room reverb | hard surfaces, glass, tile | partial improvement, never a full fix |
| Clothing rustle on a lav | mic placement | manual removal, sometimes a swap to another mic |
| Level mismatch between speakers | different mics or distances | fully fixable with balancing |
Two rows deserve emphasis. Wind and reverb are the two problems post cannot properly solve, and both are common in Miami. A microphone in a sea breeze without adequate protection produces a recording that no software will rescue. A room with glass on three sides and a tile floor produces reverb baked into the recording, and reverb cannot be removed the way noise can, only reduced at the cost of the voice sounding strange.
The implication is that audio post begins on set. Capturing clean sound is discussed in the production guides on our blog, and everything that goes right there saves hours later.
Loudness: the part with actual numbers
Loudness is the one area of audio post with objective targets rather than taste. Platforms measure the perceived loudness of a programme and adjust it towards their own reference, which means a file mastered far above the target will simply be turned down, often losing punch in the process.
The practical approach is to know where the video is going and master for that destination. Broadcast, streaming platforms and social platforms all publish loudness specifications, and they are not identical. Where the same programme goes to several places, produce a master and adjust for each rather than compromising with a single figure that suits none of them.
Alongside overall loudness there is peak level, which prevents distortion, and loudness range, which describes how much variation exists between the quiet and loud parts. A programme with too wide a range is uncomfortable on a phone in a noisy environment; too narrow and it sounds lifeless.
Mixing for how people actually watch
The mix that sounds excellent on studio monitors may be unusable on a phone speaker, and most viewers are on a phone speaker.
Check the mix on at least three systems: good headphones, a laptop speaker and a phone. If the dialogue is clear on the phone, it will be clear everywhere. Pay particular attention to low frequencies, which phone speakers do not reproduce at all, so a music bed carried by bass will vanish and take its emotional weight with it.
Assume a proportion of viewers watch with sound off. That is an argument for captions rather than for the mix, but it interacts: a video that depends entirely on a music-driven emotional arc will land differently for a silent viewer, and knowing that shapes the edit. Caption formats and accuracy are covered in the post-production material on our blog.
Music and licensing
Music choice is a creative decision with a legal dimension attached. Confirm what the licence covers before the mix is locked: which platforms, which territories, how long, and whether paid advertising use is included.
From a mixing perspective, choose music with space in it. A dense track competing with the same frequencies as a human voice forces you to duck it so far that the music stops working. A sparser track sits comfortably underneath dialogue and needs less fighting.
Practical technique: set the dialogue level first and leave it alone, then bring the music up until it is present but does not compete. Doing it the other way round, setting music first and then trying to fit dialogue over it, produces the buried-dialogue mixes that dominate corporate video.
A workable order of operations
- Lock the picture edit before starting audio post. Recutting after a mix means redoing the mix.
- Clean and edit dialogue manually first, before any processing.
- Apply noise reduction conservatively and listen to the result on headphones.
- Equalise for clarity, then apply compression for consistency.
- Set dialogue level, then place music and effects underneath.
- Check on headphones, laptop and phone.
- Measure loudness against the destination target and adjust.
- Export, then listen to the exported file rather than trusting the timeline.
That last step catches more errors than any other. A file that sounded correct in the edit and wrong after export is a familiar experience, and the only way to find it is to play the actual deliverable.
Room tone and the seams between takes
One habit separates a mix that sounds continuous from one that draws attention to itself: filling the gaps. When dialogue is cut, the background between words changes abruptly, and those small jumps in the noise floor read to the ear as edits even when the picture is smooth.
The fix is room tone, a stretch of the location's own background recorded on set with nobody speaking. Laid under the dialogue track it fills every gap with the correct ambience and the seams disappear. It takes thirty seconds to record and it is the single most useful thing a sound recordist can hand to an editor. Where none exists, the alternative is to build it from quiet passages within the takes, which works but takes far longer.
When to bring in a specialist
An editor with good instincts can handle a straightforward corporate piece. A dedicated audio person becomes worth it when there are multiple speakers with problem recordings, when the piece is going to broadcast with strict delivery requirements, when music and effects are doing significant creative work, or when a recording has a fault serious enough that specialised tools are needed.
The signal to bring someone in is usually time. If an editor has spent half a day fighting a dialogue track and it still sounds wrong, that is a specialist job and continuing is expensive.
How we handle audio post
We treat audio as a scheduled stage rather than as the last hour before delivery, and we clean dialogue manually before applying any processing, because tools applied to messy tracks produce artefacts rather than clarity. We set dialogue first and place music underneath it, and we check every mix on a phone speaker before it goes out.
We master to the loudness target of the actual destination rather than to a general figure, and where a piece goes to several platforms we produce versions rather than a compromise. More on our post-production process is on the blog, background on the about page, and you can reach us through the contact page.
The short version
Audio is what makes viewers leave, so give it a scheduled stage rather than the last hour. Work in order: manual dialogue cleanup, conservative noise reduction, equalisation, compression, then balance with music underneath. Wind and room reverb are the two Miami problems post cannot fix, which makes capture on set decisive. Set dialogue level first and bring music up to it, never the reverse. Check the mix on headphones, a laptop and a phone. Master to the loudness target of the actual destination, produce separate versions for different platforms, and always listen to the exported file rather than the timeline.