Most people describe bad audio with one word. It is almost always three problems wearing a coat, and they need fixing in a specific order.

The three problems people mean by "bad audio"

Room. The steady hiss under everything: a fan, a laptop, traffic through a window, the air conditioning you stopped hearing an hour ago. It has no rhythm, which is exactly why it is removable.

Tone. The voice itself sounds thin, boxy or harsh. Nothing is wrong with the recording; the frequency balance is unflattering. A small mic close to your mouth and a large room both do this, in opposite directions.

An audio waveform shown before and after silences are closed up, the take shorter.

Level. The words are fine, they are just too quiet, or they duck under the music, or one clip is twice as loud as the next.

They sound alike coming out of a laptop speaker and they need completely different fixes.

Apple's Final Cut Pro user guide is the shortest useful explanation of levels and headroom - including why a technically correct peak can still sound bad.

Fix them in this order

Room, then tone, then level. That order is not a preference, it is arithmetic.

Every tone adjustment is a gain change on some part of the spectrum. If you brighten a voice before you remove the hiss, you brighten the hiss too, and now it is a harder thing to remove because it has been shaped along with the words. If you raise the level first, you raise the noise floor with it and give the noise removal a louder problem to solve.

Do it the other way round and each step makes the next one easier.

Remove the room first

Open the recording in Zella and run Remove Noise. It runs on-device, so nothing is uploaded and nothing waits on a queue. What it does is learn the steady part of the signal - the part that is present when you are not talking - and subtract it.

Two things worth knowing. The first is that it works on steady noise, not on events: a fan yes, a door slam no. The second is that if it leaves a faint hiss under your words, that is a specific and fixable thing rather than a failure, and why noise removal still leaves hiss is the whole explanation.

If the noise is loud enough that you can hear it clearly in the room, move the mic closer before you record again. Nothing in software beats twelve inches of distance.

Then shape the tone

Polish Voice is the tone stage, and it ships five voicings on one broadcast-clean foundation:

Tone What it does Use it when
Natural Subtle, true to source The recording is already good and you want it tidied
Warm Fuller low end A small mic or a thin voice, common on laptop and phone mics
Bright Crisp, airy highs A dull room, a mic behind a pop filter, a soft speaker
Podcast Clear, present, even Long-form talking, where consistency matters more than character
Cinematic Big, deep, dramatic A voiceover over picture, not a talking head

Each has an intensity slider, and the honest advice is to use less than you think. Tone processing is flattering at 40 percent and obvious at 100. If a listener can hear that something was applied, it was applied too hard.

One thing Polish Voice deliberately does not do is touch your colour or your pacing. It is an audio stage, and the one-tap Auto-Polish pass is where the other decisions live.

Then set the level

Per-clip volume comes last, because the two stages above changed how loud everything is. Two rules cover most videos:

  1. Match your clips to each other before you match anything to the music. A viewer forgives a quiet video and never forgives a video that changes volume between cuts.
  2. Let the music move, not the voice. Auto-Duck dips the bed under your words and swells it back when you stop, which is the correct direction: the voice is the content and the music is the frame. Audio ducking covers how far to dip it.

What polish cannot fix

Clipping. If the input was too hot and the waveform is squared off at the top, the information is gone and no processing puts it back. The tell is a crackle on your loudest words that survives every setting you try.

Also: a mic across the room. Room tone is subtractable, but the reverb that comes with distance is part of the voice by the time it reaches the file. Get closer.

Frequently asked questions

Should I remove noise or polish first? Remove noise first, always. Polish is gain applied to a frequency range, so polishing first amplifies the noise you were about to remove and makes it harder to take out cleanly.

Which tone should I use for a talking-head video? Warm on a laptop or phone mic, Podcast on a decent USB or XLR mic. Bright is for a dull room and Cinematic is for voiceover over picture rather than a face on camera.

Does any of this upload my audio? No. Both stages run on your device, which is also why they work on a plane.

Why does my voice sound worse after polishing? Almost always intensity. Pull the slider back to around 40 percent and listen again. The second most common cause is polishing before removing noise.

Can I fix a recording that peaked into the red? Not really. Clipping deletes the top of the waveform, so there is nothing left to restore. Re-record if you can, and drop the input gain a few dB.

Does this work on the iPhone app too? Yes. Remove Noise and Polish Voice both run on-device on Zella for iPhone, with the same five tones.