The effect where a big word sits behind the person, and they walk in front of it, is one cutout and one layer order. It looks expensive and it is neither difficult nor slow.

What is actually happening

Three layers, bottom to top:

  1. The video frame.
  2. The title.
  3. The same video frame again, masked to just the subject.

Layer three is the whole trick. The person is drawn twice - once as part of the background, once as a cutout on top - and the text is sandwiched between the two copies. The eye reads that as depth, because in the real world a thing you can see part of is a thing behind something else.

A video frame with captions burned in, one word highlighted as it is spoken.

That subject mask is the same cutout used for background removal, running on your device, so no green screen is involved and nothing is uploaded.

Reframing that keeps a face centred uses Apple's Vision framework; the detection happens locally, which is what lets the crop keep up without a round trip.

Adding it

  1. Add a title to your clip in Zella and set the text.
  2. Position it where the subject will cross it. Behind the head and shoulders is the classic placement.
  3. Turn on Text Behind Subject.
  4. Play it back. The mask is generated per frame, so a moving subject reveals and covers the text as they move.

It is a switch on the title rather than a separate feature, which means everything else about that title still works: the font, the gradient fill, the stroke, the shadow, the kinetic entrance. A word can slide in, then get walked in front of.

What makes it read

Big text. This effect is about overlap, and there is nothing to overlap if the word is small. One or two words at a size that spans a third of the frame is the shape that works. A sentence does not.

Static text, moving subject. If both move, the eye cannot tell which one is in front. Pin the title and let the person cross it.

Contrast against the background, not against the person. The text is read where it is visible, which is the part of the frame the subject is not covering. Pick a colour that stands off the wall behind you.

Placement behind the shoulders. Head height is the natural spot because that is where a person's outline is most recognisable, so the overlap is unmistakable.

Where it breaks, and what to do

Fast motion. A cutout is generated per frame, and a fast-moving limb is a motion-blurred edge with no clean boundary. Hands crossing the text at speed can flicker. Move the text away from the hands, or slow the shot.

Hair against a busy background. Fine detail against a background with similar colour is the hardest case for any segmentation, and it is the same failure described in why background removal looks bad when you move. A plainer wall fixes more than any setting.

Loose clothing and props. A scarf, a mug, a held phone: these are things the cutout may or may not consider part of the subject, and the answer can change between frames. Keep the text clear of anything the person is holding.

A subject that leaves the frame. With nobody to cut out, the text is just a title. That is fine, and often a good way to end the shot - but it is worth knowing the effect vanishes rather than degrades.

Where it earns its place

  • The opening word of a short. One noun, big, behind you, entering as you start talking. It states the topic without a voiceover explaining that you are about to state the topic.
  • A name or a number. A price, a year, a chapter marker. Something the viewer should hold in mind while you talk over it.
  • A section break in a long video. Same word treatment, used two or three times, becomes a structure the viewer can feel.

Used once it is a moment. Used on every title in a video it is a tic, and the video stops having an emphasis level above ordinary text.

Frequently asked questions

Do I need a green screen for this? No. The subject cutout is generated on your device from the footage itself, so an ordinary room works. A plainer background makes a cleaner edge, but it is not a requirement.

Does it work on any title style? Yes. Text Behind Subject is a switch on the title rather than a separate object, so the font, fill, stroke, shadow and entrance animation all still apply.

Why does the text flicker around my hands? Fast-moving limbs blur, and a blurred edge has no clean boundary for the cutout to follow. Move the text away from where your hands travel, or use a slower shot.

Can I do this with a whole sentence? You can, but it rarely reads. The effect depends on overlap, and small text has too little of itself hidden to register as being behind anything.

Does it slow the export down? It adds a per-frame mask to the render, so a long clip takes longer than the same clip without it. On a short title it is not something you will notice.

Will it work if the subject is not a person? The cutout is trained on people, so a person is the reliable case. A pet or an object may or may not be found, and it is quicker to test on a few seconds than to guess.