FOR PODCASTERS
One recording, a week of clips.
Turn a long conversation into short vertical clips with captions, without scrubbing the timeline hunting for the good bit.
THE PAIN
- The best 40 seconds are buried in 90 minutes
- Clipping by hand costs longer than recording did
- Audio-first content dies without captions
THE WORKFLOW
Import the episode recording
Cut the strongest moments on the timeline
Captions and a vertical reframe on each, export
FROM YOUR PHONE
Record a solo episode or a field interview on the phone, then let Auto-Polish clean the voice before you cut clips.
Zella for iPhone →WHEN THIS OUTGROWS THE FREE TIER
Unlimited export length. Podcast cuts routinely run past the free ceiling, and the multi-aspect queue renders each one for every platform in a single pass.
Everything else here is free forever: the complete editor, all the AI, unlimited recording, and watermark-free export up to 2 minutes at 1080p. Pro is $89 once and covers Mac and iPhone on a single purchase. See pricing →
Podcast audio is the easy part. The reason most shows stay small is that nothing about an audio file travels on a feed built for video, and cutting clips by hand is slow enough that it is the first thing to get dropped in a busy week.
Zella turns an episode recording into the clips that actually get discovered. It transcribes on-device, cuts the silences and the filler words that sound fine in audio and look terrible on screen, and captions the result so it works on mute. Cut the strongest moments on the timeline and auto-reframe turns each one into a vertical, captioned cut.
Voice cleanup runs locally too. The mic noise, the room, and the hiss come out on your own machine rather than being uploaded to a service that charges by the minute.
WHAT YOU ACTUALLY SHIP
A captioned clip from the best two minutes
Cut the best stretch on the timeline - captions come from your on-device transcript, so the clip reads on mute.
An audiogram that is actually a video
Captions, a title card and a waveform, exported vertical for Reels and Shorts rather than as a static image.
A cleaned-up voice track
On-device noise removal and voice polish, with music that ducks under your voice automatically rather than by hand.
The audio on its own
Export M4A when you want the episode itself, from the same project as the clips.
THE HONEST ANSWERS
My episodes are an hour long.
Zella is not trying to be your DAW. It is for the clip layer on top: the two-minute cut that goes on social. Recording length is unlimited on every tier, and the export cap applies to the clip you actually publish, which is usually well under two minutes.
Will the transcription handle two people talking?
It transcribes what it hears and gives you a word-level transcript you can correct by tapping a line. It does not label speakers, so if your workflow depends on diarisation, that is a real gap worth knowing before you commit.
I record remotely, so audio quality varies.
That is where the on-device noise removal earns its place. It handles room noise and hiss on the machine you are sitting at, and the expander is tuned to avoid chopping the front of words, which is the usual failure mode of aggressive gates.
FREQUENTLY ASKED
Can Zella make video clips from a podcast?
Yes. It transcribes the episode on-device, you cut the strongest moments on the timeline, and auto-reframe plus captions turn each one into a vertical clip.
Does Zella remove filler words like um and uh?
Yes, on both Mac and iPhone. On iPhone the first caption run downloads a one-time speech model over Wi-Fi; until it finishes, fillers may not appear as removable.
Can I export just the audio?
Yes. M4A export gives you the audio track on its own, from the same project.
Is my audio uploaded anywhere for transcription?
No. Transcription, noise removal and voice polish all run on your own device. There is no account and no upload path.
CLOSE TO THIS