Stack the two videos on a 9:16 canvas: the clip you are reacting to on top, you underneath, both draggable and resizable. Give the source clip more height than your face - roughly two thirds to one third works - because that is what the viewer came for. In Zella, Reaction Split does the layout in one tap and needs both videos in the Clips panel first.
A reaction video is two pieces of footage competing for one screen. Almost everything that makes one good or bad is how you resolve that competition.
The layout
Vertical, stacked, not side by side. On a 9:16 canvas two stacked panels each get a usable width; two side-by-side panels each get half a phone screen of width, which is unreadable for anything with detail in it.
Source clip on top, you underneath. This is close to universal in the format and it is not arbitrary: the top of a vertical frame is where the eye lands first, and the source is the reason the viewer is there.

Proportions: roughly two thirds to one third. The clip being reacted to gets the larger share. Your face is a reaction shot, and a reaction shot does not need to be big to work. Giving yourself half the screen is the most common mistake in the format and it makes the video feel self-regarding before anyone has decided whether it is.
There is an exception: if you are talking over the clip for long stretches rather than reacting to moments, you are making commentary and not a reaction, and the proportions can move toward even.
The tracking behind a face-aware crop is Apple's Vision framework, detecting on the device so the frame can follow a speaker with no server in the loop.
AVFoundation is the layer underneath - Apple's media framework, and what a Mac editor uses to assemble a timeline and encode the finished file.
Doing it in one tap
Zella has Reaction Split in the Viral Templates tab. It stacks both videos on a 9:16 canvas and both panels stay draggable and resizable afterwards, so the automatic layout is a starting point rather than a fixed frame.
It needs both videos in the Clips panel first: your reaction and the clip you are reacting to. That requirement is real rather than incidental, because a layout template with only one video has nothing to lay out.
Once the layout is there, picture in picture and split screen covers the manual sizing controls if you want to depart from the default.
The audio, which is the actual hard part
Two videos means two audio tracks and they will fight.
Duck the source under your voice. When you talk, the source clip should drop. When you stop, it comes back up. Doing this by hand is tedious and doing it with automatic ducking is one toggle. Music ducking explained is the same mechanism applied to a bed rather than a clip.
Do not mute the source entirely. Reaction videos where the source is silent while you talk feel broken, because the viewer loses the thing they are watching you react to.
Watch your own levels. You are close to a mic in a quiet room and the source clip is compressed, loud broadcast audio. Without correction you will be either buried or shouting. How to change video volume covers per-clip levels.
Cutting it down
The temptation is to react to the whole thing. Almost every good reaction video is a tenth the length of its source.
Keep the moments where you genuinely responded, cut everything between them, and let the source clip jump. Viewers accept discontinuity in the source panel far more readily than in yours, because they understand they are watching an edited reaction rather than a screening.
Cutting the dead air does most of this pass mechanically.
Two things that get videos taken down
Worth being direct about. Reacting to someone else's footage is a copyright question, and the answer depends on how much you use, how transformative your commentary is, and the platform's own tolerance. Short excerpts with substantial original commentary sit in much safer territory than long uncut playback with occasional laughing.
The practical version: use less of the source than feels natural, talk over more of it, and never post a reaction where the source runs uninterrupted for minutes.
Captions
Caption yourself, not the source. Two sets of captions on a split screen is visual noise, and the source usually has its own audio that the viewer can hear.
Put your captions in your panel, not across the middle of the frame where they cross the boundary between the two videos. See how to move captions so they do not cover your face.
Common questions
Which panel goes on top? The clip you are reacting to. The top of a vertical frame is where the eye lands first, and the source is the reason the viewer is there.
How big should my own panel be? Around a third of the height. Giving yourself half the screen is the most common mistake in the format and it reads as self-regarding.
Should I mute the source clip while I talk? Duck it, do not mute it. A silent source while you speak feels broken because the viewer loses the thing they are watching you react to.
Make your next video with Zella.
Record, edit and ship on Mac and iPhone - local, private, free to start.
PART OF A SERIES
One-tap templates
The seven passes that restructure an edit rather than recolour it.
- Video Templates That Actually Change the Edit (Not Just the Look)
- How to Make a Reaction Video With a Split Screenyou are here
- How to Add Photo Pop-Ins to a Talking Head Video
All of it lives under One-Click.
RELATED