A video call to action works when it asks for one specific thing in the last 3 to 5 seconds, spoken out loud and shown on screen at the same moment. Comment the word GUIDE and I will send it beats like and subscribe, because it names a single action and pays the viewer for taking it. Match the ask to the goal: follow for reach, comment for the algorithm, one link for sales.
A call to action works when it is specific, singular, and placed where attention is highest: the last 3 to 5 seconds, said out loud and shown on screen at the same time. Like and subscribe fails because it is vague and asks for a favour. Comment the word GUIDE and I will send it works because it is one clear action with an obvious payoff.
Most videos either forget the CTA or bolt a generic one on after the viewer has already left. The ask is not a formality. It is the whole reason a view turns into a follow, a comment or a sale.
What separates a CTA that converts from one people ignore?
Three things, and all three are cheap to fix:

- Specific beats generic. Save this for later outperforms engage with this, because the viewer knows exactly what to do. Name the button, name the word, name the action.
- One ask, not five. Like, comment, subscribe, share and check the link gives the brain a decision to make, so it makes none. One action per video.
- A reason to act now. Give a payoff (comment GUIDE and I will DM the template) or a stake (this comes down tomorrow). A CTA with no reason is a wish.
Get those three right and even a plain CTA converts. Get them wrong and no amount of design saves it.
Which CTA should you use?
Match the ask to what you actually want from this video, then use only that one.
| Ask | What it is for | Wording that works |
|---|---|---|
| Follow / subscribe | Reach and channel growth | If this helped, follow. I post one of these a week |
| Comment | Algorithm signal, doubles as lead capture | Comment CHECKLIST and I will send the full list |
| Save | Reference content. Saves signal high value | Save this before your next recording |
| Share | Distribution beyond your followers | Send this to the person who edits your videos |
| Link in bio | Traffic and sales. Weakest by default, so pay for the click | The free template is in my bio. Grab it before the video ends |
| Watch next | Watch time and session length | If you liked this, the one on captions goes deeper |
Buy and sign up sit at the far end of that list: reserve them for warm audiences, and make the offer concrete and low-friction. Do not stack any of these. Pick the single action that matches this video's job.
Where does the CTA go in the video?
Placement matters as much as wording. In order:
- Deliver the value first. Nobody acts on an ask the video has not earned.
- Put the primary CTA in the last 3 to 5 seconds. Intent peaks right after the payoff, when the viewer is deciding what happens next. Say it and show it in the same moment.
- Plant a soft setup early, optionally. A quick line in the first 10 seconds (stick around, the template is at the end) opens a loop without asking for anything yet.
- Never bury it after the ending. If the video visually ends and then the CTA appears, retention has already fallen off. The CTA is part of the ending, not an epilogue.
- Close loop-friendly. On short-form, a last line that flows back into the opening keeps the replay running, which lifts the same signals the CTA is chasing.
Why every CTA has to be on-screen text too
Around 85% of social video is watched muted, so a spoken-only CTA reaches a fraction of your audience. The rule: every CTA is also on-screen text. Burn the exact action into a caption or a callout in the final seconds.
Four things that lift response:
- Put the CTA word in a different colour or a box so it separates from your regular captions.
- Aim an arrow or pointer at the follow button or the comment field.
- Hold it on screen for the full final 3 to 5 seconds, long enough to read twice.
- Match the on-screen words to the spoken words exactly, so muted and sound-on viewers get the same instruction.
Someone watching with no sound should still know precisely what to do.
A phrase bank you can steal
Same goal, two versions. The difference is always specificity plus a reason.
| Goal | Weak line | Stronger line |
|---|---|---|
| Followers | Like and subscribe | Follow if you want part 2 tomorrow |
| Comments | Let me know what you think | Comment GUIDE and I will send the checklist |
| Saves | Bookmark this | Save this for the next time you record |
| Traffic | Link in bio | The 1-page template is in my bio, free, no email |
| Watch time | Check out my other videos | The captions one is next. It fixes the mistake at 0:12 |
| Shares | Share with a friend | Send this to whoever edits your videos |
Frequently asked questions
Where should a CTA go in a short video? In the last 3 to 5 seconds, as spoken audio and on-screen text together, after the value has landed. Optionally plant a soft tease in the first 10 seconds, but keep the actual ask for the close.
How many CTAs should one video have? One primary ask. A soft setup earlier is fine, but the video should drive toward a single action. Competing CTAs split attention and lower response on all of them.
Why is like and subscribe not working? It is generic, it asks for a favour with no payoff, and everyone has tuned it out. Replace it with a specific action tied to a reason: follow if you want the next one, or comment GUIDE and I will send the template.
Do CTAs hurt retention? Only when they are early, long, or interrupt the value. A tight CTA in the final seconds costs nothing, because the viewer already got what they came for. Slow openings hurt retention. Clear endings do not.
What is the best CTA for followers versus sales? For followers, ask for the follow immediately after overdelivering, or use a comment-a-word mechanic to spike the algorithm. For sales, drive to a single low-friction link with a concrete reason to click now. Never do both in one video.
Do it faster in Zella
Zella records and edits on-device on Mac, iPhone and iPad, with no account and no subscription, which makes a clean readable CTA a two-minute job:
- Record and let Auto-Polish run. It fires the moment a recording opens and generates captions automatically, so your baseline text layer already exists.
- Add the CTA as a callout. Drop on-screen text, an arrow or a popup into the last 3 to 5 seconds naming the exact action. There are 19 callout types, and arrow, label, popup and sticky note all read well at the end of a clip.
- Style it apart from your captions. Colour the CTA word differently from your caption preset so muted viewers cannot miss it. Zella ships 9 caption presets with active-word highlight.
- Match the spoken line. Make your final spoken sentence and the on-screen text say the same words.
- Export watermark-free. Free exports run up to 2 minutes at 1080p. Longer and 4K are a one-time $89 Pro unlock, one payment covering Mac, iPhone and iPad. See pricing.
Pair a strong CTA with a tighter edit using AI cleanup, and if talking-head is your format, make it land harder with less boring talking-head videos. The Mac app is on the Mac App Store, and Zella for iPhone and iPad is coming to the App Store: see Zella and give your next video an ending that actually asks.
Make your next video with Zella.
Record, edit and ship on Mac, iPhone and iPad - local, private, free to start.
RELATED