For years, editing a video meant learning where every button lives. That is changing. With an AI editing assistant you can just type what you want and let the app do the clicking - cut the dead air, add captions, reframe for TikTok, drop in music, export.

The short version

Open your recording in Zella, open the Assistant tab, add a free AI key, and type your request. For example:

"Remove the silent gaps, add captions, and make it vertical for TikTok."

How to Edit a Video by Typing on Mac (AI Assistant)

The assistant transcribes and trims the dead air, generates captions, reframes to 9:16, and tells you what it changed. Every step is undoable with ⌘Z. See the AI Assistant docs for the full setup.

Transcription of this kind runs through Apple's Speech framework, which supports on-device recognition, so the audio does not have to leave your device.

Face-aware reframing is built on Apple's Vision framework, which does the detection locally so the crop can follow a speaker without a cloud round trip.

Why type instead of click?

Two reasons. First, speed: one sentence can trigger a five-step edit that would take a beginner ten minutes of menu-hunting. Second, discoverability: most people never find half of an editor's features. If you can describe the outcome - "make this look cinematic," "cut the part where I fumble the intro" - you do not need to know the feature's name.

What you can say

Zella's assistant drives more than 70 editing actions. A sample of prompts that work in one line:

You type What happens
"Remove all the silences and filler words" On-device cleanup ripple-deletes dead air and "um / uh / like"
"Add captions in the Word Pop style" Transcribes locally and adds animated captions
"Make it cinematic and auto-enhance" Applies a filmic look and a one-click color lift
"Add background music and a whoosh on the first cut" Adds a music track under your voice plus a built-in sound effect
"Make it vertical for TikTok and export in 4K" Reframes to 9:16 and renders (4K needs Pro)

You can chain them into a single request and the assistant runs the whole sequence.

Cutting by what you said, not by timestamp

The most useful trick: the assistant can read your transcript. So you can edit by content:

"Cut the part where I ramble about the old pricing."

It searches the spoken words, finds the span, and ripple-deletes it - no scrubbing the waveform to find the spot.

The assistant can look at the frame too

Ask "is my face centered?" or "is this too dark?" and the assistant inspects the actual video frame on-device (faces, framing, brightness) and can fix it. Nothing is uploaded to do this.

Is it private?

Zella never uploads your video. The only thing that leaves your device is the text of your request plus a small summary of the project, sent to the AI provider you chose with your own key - exactly like typing into that provider's app. Want zero network use? Run a local model with Ollama and the whole loop stays on your machine.

What the assistant costs

Nothing. It is not behind the Pro unlock, and it works with free AI providers (Groq, Gemini, Cerebras, or local Ollama), so there is no per-edit fee either. You paste your key once and it lives in your Mac's Keychain.

The catch (and when to still click)

Prompt editing is fantastic for the 80% of edits you can describe. For frame-exact tweaks - nudging a title two pixels, dialing a specific curve - the panels are still faster. The good news: the assistant and the panels share the same timeline and the same undo history, so you can mix both freely.

FAQ

Does the assistant upload my video? No. Only the text of your request and a short summary of the project - durations, the clip list, how many captions exist - go to the AI provider you picked. The frames and the audio stay on your Mac.

Which AI provider do I need? Whichever you like. Groq, Google Gemini and Cerebras all have free tiers, and Ollama runs a model locally with no key and no internet at all.

What happens if it does the wrong thing? Every action it takes is an ordinary timeline edit, so ⌘Z undoes it. The assistant and the manual panels share one timeline and one undo history, so you can always finish the job by hand.

Can it cut by what I said instead of by timestamp? Yes. It reads the transcript, so "cut the part where I ramble about the old pricing" finds the spoken span and ripple-deletes it.

Is there a per-edit fee? No. Zella runs no paid AI service of its own, so you pay your provider directly - on a free tier, that is usually nothing.

Try it

  1. Download Zella (free, no watermark).
  2. Record or import a clip.
  3. Open the Assistant tab, add a free key, and type your first edit.

Prompt-based editing will not replace a timeline for everyone - but for turning a raw take into something postable in one sentence, it is the fastest thing on a Mac today.