ZELLA / Docs Download for macOS

On-Device AI on iPhone & iPad: Auto-Polish, Make Vertical, B-Roll & Assistant

On-device AI on iPhone — Auto-Polish, Make Vertical, background removal and the typing assistant

Quick answer: On the phone, Auto-Polish is a full one-tap retention edit — it cuts silences and fillers, keeps a steady cut cadence with varied transitions and a matched sound effect under every cut, adds Beast keyword captions with emoji, removes noise and polishes the voice, opens with a hook title and punch-in, places emphasis zooms, grades the color, lays a mood-matched music bed with an arc, and breaks up any static stretch. Clean Up does just the silences and fillers, Make Vertical does a face-aware 9:16 reframe, Reel Pack turns a clip into a reel in one tap, and transcript editing ripple-cuts the video as you delete words. The camera’s live background (blur, cut out or replace) is baked straight into the recording as you film. Everything runs on-device — nothing ever leaves the phone.

On this page: Auto-Polish · Make Vertical · background removal · Auto B-Roll · AI Assistant · privacy · FAQ


Auto-Polish, in one tap

Auto-Polish is the flagship one-tap edit — a full retention-style edit of a raw take, built from what viral shorts actually share. Point it at a raw recording and it runs the whole pipeline, on-device:

  1. Removes silences — cuts the dead air between sentences, with word-safe padding so it never clips a word.
  2. Removes filler words — trims the um, uh and like.
  3. Adds a hook — an opening title built from your first spoken line, canvas-fit.
  4. Keeps a cut cadence — spreads cuts across the whole take (roughly every five seconds) so no stretch drags, then places a varied transition on every join — not the same dissolve twelve times.
  5. Sound at every cut — one matched whoosh, impact or riser under each transition, picked by transition type, so cuts feel edited instead of spliced. Tap once to remove them all; sounds you placed yourself stay.
  6. Removes noise & polishes the voice — neural denoise under the speech, then the studio polish: warmth, compression, −14 LUFS broadcast loudness.
  7. Captions in the Beast style — big word-pop captions with on-device keyword emphasis and one emoji per line, synced to the final timeline.
  8. Frames the speaker & punches the opening — zooms to your face, then starts the video punched in (~1.3×) on the face and settles out over ~1.7 s — the first 1–3 seconds decide retention, so they get their own move.
  9. Emphasis zooms — small punch-ins on your strongest spoken words through the whole video, capped so it stays tasteful and never stacks on a transition.
  10. A color grade — a subtle warm-cinematic grade, skipped if you already graded anything.
  11. Mood-matched music with an arc — picks a loop to match your pace and energy, snaps its start to the take’s first beat, then ducks under speech, swells in pauses, and fades out at the end.
  12. Pattern interrupts — a final pass that scans the finished layout and breaks up any window that’s still visually static for six seconds or more.

If you’ve picked a platform preset, it finishes with a face-aware reframe to that aspect. Every step is undoable, and re-running Auto-Polish never stacks doubles — it replaces its own previous work.

It’s the mobile version of the Mac’s AI cleanup — a raw phone recording becomes a tight, scored, captioned cut in a single tap.

Why “um” and “uh” actually get cut on iPhone. iOS transcribes with a verbatim on-device Whisper model, which writes disfluencies out in full — so filler removal can really see and cut them. (On the Mac, Apple’s speech engine returns a cleaned transcript, which is why the Mac doc warns that plain “um”/“uh” may survive.) The honest caveat: the model is downloaded once, about 145 MB, the first time you transcribe, with a progress note. After that it’s fully offline. If it hasn’t downloaded yet — say you’re on a plane on first launch — Zella falls back to Apple’s on-device recognizer, so captions still appear; they’re just from the cleaned transcript, and stray “um”s may need a manual trim.


Make Vertical

Make Vertical reframes your video to 9:16 for TikTok, Reels and Shorts, using face-aware auto-tracking so you stay centered as you move. One tap turns a horizontal or square recording into a vertical clip that’s ready to post. See the Mac’s reframe for the same idea on desktop.


Background removal & replacement

Beyond the live-while-you-record background from the recording tools, you can remove and replace the background in the editor too. It’s segmented live on the Neural Engine — no green screen, nothing uploaded — so you can drop yourself onto a different scene after the fact.

The cutout is locked to the main subject: the segmenter picks the largest person in frame and gates everyone else out, so someone walking behind you doesn’t get cut out with you, and the edge is choked so no pale halo crawls along your hairline over solid backgrounds.


Auto B-Roll

Auto B-Roll reads your captions, picks the most visual keyword per line on-device, and places a cutaway at each one automatically. On iPhone and iPad it works with no key and no account:

  1. Add captions first (Auto B-Roll works from them).
  2. Tap Add cutaways.
  3. Zella generates a cutaway card per keyword — a designed gradient card with the word on it — and drops each onto the timeline near the line that mentions it.

These are generated motion-graphic cards, not licensed stock footage — which is exactly why they need no key and never touch the network. The whole run is 100% on-device.

Optional: use real stock footage

Want actual stock clips instead? Open Use my own stock footage (Pexels / Pixabay), paste a free key from either provider, and tap Add stock b-roll. Zella stores the key in the Keychain on the device. Only the keyword leaves the phone to search; your video never does. This is the path the Mac’s Auto B-Roll uses.


The AI Assistant: edit by typing

The AI Assistant lets you edit by typing plain English51 editing actions, driven by chat. Tell it “remove the dead air, add captions and make it vertical,” and it runs the whole chain. Ask it to add music, drop a sound effect, place a title, or export a specific format.

  • Bring your own key — use a free language-model provider; the key stays on the device.
  • On-device and privateonly your typed prompt is sent to your chosen provider. Your video is never uploaded.

It’s the mobile version of the Mac’s AI Assistant.


What actually leaves the phone?

Zella’s AI is on-device by default. Only two things are opt-in and network-based, and neither is on unless you turn it on:

  • Stock b-roll (the optional Pexels / Pixabay path) sends only a keyword from your captions to the provider — never your video. The default cutaway cards are generated on the device and send nothing at all.
  • The AI Assistant sends only your typed prompt to the language-model provider whose key you added — never your video.

Everything else — Auto-Polish, captions, Make Vertical, background removal, caption translation, and the default Auto B-Roll cutaways — runs entirely on the device.


iOS AI FAQ

What does Auto-Polish do in one tap? The full retention edit: silence and filler cuts, a steady cut cadence with varied transitions and a matched sound effect under every cut, Beast keyword captions with emoji, noise removal and voice polish, a hook title with an opening punch-in, emphasis zooms on your strongest words, a color grade, mood-matched music that ducks under speech and swells in pauses, and pattern interrupts so nothing stays static for six seconds.

Can I re-run Auto-Polish after making changes? Yes — it replaces its own previous work instead of stacking doubles, so a re-run is always safe.

Does Make Vertical follow my face? Yes — it’s a face-aware 9:16 reframe, so you stay centered as you move.

Do I need a green screen for background removal? No. It’s segmented live on the Neural Engine, on-device, with nothing uploaded.

What key does Auto B-Roll need? None. The default Add cutaways generates gradient cutaway cards on-device with no key and no account. A free Pexels or Pixabay key is only needed for the optional real stock footage path, and it’s stored in the Keychain on the device — only the keyword ever leaves the phone.

Is my video sent anywhere by the AI Assistant? No. Only your typed prompt goes to your chosen provider. Your footage never leaves the phone.

Does Zella actually remove my “um”s on iPhone? Yes. iOS transcribes with a verbatim on-device Whisper model, so plain “um”/“uh” are written out and can be cut — unlike the Mac, where Apple’s cleaned transcript often drops them. The model downloads once (~145 MB) on first transcription, then works fully offline.

Do I need a network connection for captions? Only for that one-time model download. If it hasn’t happened yet, Zella falls back to Apple’s on-device recognizer so captions still work offline — just from a cleaned transcript.

Pro tips & gotchas

  • Start with Auto-Polish, then fine-tune by hand — it does the boring 80% in one tap.
  • Captions before b-roll. Auto B-Roll reads your captions for keywords, so add captions first.
  • No key needed to start. Cutaway cards work out of the box; add a free Pexels/Pixabay key only if you want real stock footage. The AI Assistant does need a free provider key, which stays on the device.
  • Offline? Only stock b-roll and the Assistant need the network. Cutaway cards, captions, translation and everything else run with no network at all.

Related: Captions & titles → · The iOS editor → · Export & handoff → · Mac AI cleanup → · Mac AI Assistant → · Auto B-Roll →