There are two ways to show a picture in a talking head video, and one of them is much better than the other.

Cut away versus pop in

Cutting away replaces your video with the photo. The viewer loses you entirely for the duration. It works when the photo IS the subject and you are narrating it.

Popping in puts the photo over your video as a card, usually tilted a few degrees, sized to maybe a third of the frame. You stay on screen. The photo illustrates the sentence you are saying and then leaves.

An editing timeline: video clips, a zoom keyframe and an audio track below.

For a talking head video the pop-in wins almost every time, and the reason is continuity of attention: the viewer is watching a person talk, and a full-frame cut asks them to switch context and switch back for something that was only ever an aside.

For what a viewer will and will not sit through, Nielsen Norman Group's video usability work is more useful than any amount of instinct.

Why tilted, and why not full frame

The tilt is doing real work rather than being a style choice.

A perfectly rectangular image sitting square over a video reads as a screenshot, a mistake, or an ad. Two or three degrees of rotation plus a slight shadow reads as an object placed on top, which the eye files as "supporting material" rather than "the video changed".

Size it to about a quarter to a third of the frame. Big enough to see, small enough that you are still the subject. If the photo needs to be bigger than that to be legible, it has too much detail to be a pop-in and you should cut to it properly.

Timing, which is the whole thing

A pop-in that appears half a second after you say the word is worse than no pop-in, because the viewer has already moved on and now has to reconstruct why it is there.

In on the word, out two to three seconds later. Not held for the rest of the sentence. The card exists to connect a word to an image, and once that connection is made the card is in the way.

Never two at once. Two cards on screen simultaneously means neither is connected to anything you are saying.

Doing this by hand means scrubbing to each keyword and setting an in and out point, which is why it is one of the more tedious manual edits and one of the more useful things to automate.

Doing it automatically

Zella has a Photo Cards template. It places tilted photo-card pop-ins on your keywords along with word-pop captions, matching images against the caption text so the placement follows what you actually said.

The requirement is that the photos are in the Clips panel first. It matches them to keywords, so it needs both halves: a transcript, and images to match.

The captions and the cards are placed by the same pass, which is why they agree on timing. If you already have captions you like, they are what the matching reads.

Choosing the photos

One clear subject per image. A photo with six things in it cannot be parsed at a quarter frame size in two seconds.

Crop before you import. The card is small. Whatever the picture is of should fill it.

Avoid text-heavy images. Screenshots of documents, dense charts and anything with small type are illegible at pop-in size. If it needs reading, cut to it full frame and hold it, which is a different edit.

How many

For a sixty second video, four to six cards is a comfortable ceiling.

Past that the technique becomes the format, and the video starts to feel like a slideshow with a presenter in it. The same density logic applies as with sound effects: repeated too often and the eye stops treating each one as an event.

Where they should sit in the frame

Opposite your head. If you are framed left, cards go right.

Also above the caption zone, because a card that overlaps your captions makes both unreadable and the viewer gets neither. The two systems have to share the frame, which is easier to arrange than it sounds once you have decided that you occupy one side and the supporting material occupies the other.

If your captions and your cards keep colliding, moving the captions is usually the easier of the two to relocate.

Common questions

Why tilt the photo card? A perfectly square image over video reads as a screenshot or an ad. Two or three degrees plus a soft shadow reads as an object placed on top, which the eye files as supporting material.

How long should a pop-in stay? In on the word, out two to three seconds later. Holding it for the rest of the sentence puts it in the way once the connection is made.

How many photo cards is too many? Four to six in a sixty second video. Past that the technique becomes the format and the video turns into a slideshow with a presenter in it.