Agentic Services

Covers and Thumbnails: The Frame That Decides Everything Else

By Editorial Team — reviewed for accuracy Published
Last reviewed:

An episode nobody clicks is an episode nobody watched, and the click is decided by one still frame and a few words. On a feed-driven platform that frame competes with everything else on the screen, and it competes for about as long as it takes to scroll past.

This is the highest-leverage twenty minutes in the entire pipeline, and it is routinely the part that gets ten.

Pull the frame from the clip, do not generate a new one

There is a tool for this: Frame Extractor takes a single still out of a video and keeps it as an image. A clip contains hundreds of frames; open the video, scrub to the moment, and save it.

Doing it this way rather than generating a fresh image matters for a reason that is easy to miss: the cover has to be a promise the episode keeps. A separately generated cover shows a moment that is not in the video, which converts a click into a bounce and teaches the platform’s ranking system that your content disappoints. Extracting guarantees the moment is real.

The extracted frame arrives in your library as an ordinary image. From there it can be upscaled, cleaned up, or published like anything else.

Choosing which frame

Scrub the whole clip rather than taking the frame you remember. The best cover is frequently not the best moment — it is the most legible one at thumbnail size.

What to look for:

  • A face with a readable expression. Faces outperform everything else in a feed, and an expression that raises a question outperforms a neutral one.
  • A clear silhouette. At thumbnail size, shape survives and detail does not. Squint at it; if you cannot tell what is happening, neither can anyone scrolling.
  • High contrast between subject and background. A dark figure on a dark street reads as a smudge.
  • Space for text. If you are adding a caption, the frame needs an area where it will not sit on top of the face.
  • A moment mid-action. A person about to do something is more interesting than a person who has finished.

Avoid the first and last fractions of a second of a generated clip, where motion is starting or settling and detail is at its softest.

Clean it up

Two tools cover almost every problem with an extracted frame.

Magic Eraser removes something that should not be there — a distraction in the corner, an artefact, an object you cannot use. Paint over what you want gone and it is replaced with something that fits the surroundings.

It is invention, not recovery: nothing is hidden behind the object waiting to be revealed, so the tool makes up a plausible background. Getting good results is mostly technique:

  • Zoom in before you paint. Small objects are far easier to cover cleanly at a larger size.
  • Cover the whole object plus a margin. A missed edge leaves a smear.
  • Include the shadow and any reflection. Removing an object and leaving its shadow reads as obviously wrong.
  • One thing at a time, checking after each.
  • Small objects on plain backgrounds are close to perfect; large objects on busy backgrounds are the hard case and may take a couple of attempts.

Each attempt costs, since a new picture is being made either way. Results are saved as new items — your original is never overwritten.

Upscale increases resolution, reconstructing detail as it grows rather than stretching. Use it at the end, on the version you have decided to keep — not on every draft.

What it will not do is rescue a picture whose composition is wrong. Upscaling a bad frame gives you a large bad frame. If the cover is not working, extract a different frame rather than sharpening the wrong one.

Sizing it for where it is going

Pick the shape for the destination before you do anything else, because cropping throws away part of the picture.

DestinationShape
Phone-first feeds (TikTok, Reels, Shorts)Portrait
YouTube standard thumbnailsLandscape
Most social posts and previewsSquare

If a piece is going to more than one place, extract the frame once and crop per destination — but check each crop separately. A composition that works in portrait frequently puts the subject’s head against the top edge in landscape.

Text on the cover

Add it in an image editor, not in the generation. Models produce convincing-looking nonsense where readable text is concerned, and a cover with garbled type reads as low effort at a glance.

What works, in the order it matters:

  • Few words. Three to five. It is read in motion.
  • Large. Sized to survive a thumbnail, which is smaller than the version you are looking at while you design it.
  • High contrast, with a shadow or a solid backing shape so it stays readable over a busy frame.
  • A question or an incompleteness. The cover’s job is to open a loop the video closes.
  • Off the face. Text over an expression wastes the strongest asset in the frame.

Consistency across a series

Once a format works, the cover treatment should be recognisable — same text position, same typeface, same colour treatment, same framing convention.

This is not decoration. A returning viewer scrolling a feed identifies your series by its cover before they read the title, and a recognisable cover converts a passive follower into a repeat viewer. It is also free: it costs one decision, made once, and applied thereafter.

Build a template in whatever image editor you use, with the type and layout fixed and only the extracted frame changing.

A short checklist

Before publishing anything:

  1. Frame extracted from the actual clip, not generated separately.
  2. Scrubbed the whole clip; chosen for legibility, not sentiment.
  3. Distractions erased; shadows and reflections included.
  4. Upscaled once, at the end.
  5. Cropped for the destination and checked in that shape.
  6. Text added in an editor: few words, large, high contrast, off the face.
  7. Matches the series template.
  8. Viewed at actual thumbnail size before committing.

Step eight catches more problems than the other seven combined.

Last reviewed: · Editorial policy · Report an error