A worked thumbnail example
This is a fictional design created for this tutorial. The map is an actual UNISAL v1 prediction using MIT1003 weights, calculated from the image shown here.


Download the example input to try the workflow. Your comparison should use variants of this same format and purpose.
Start with the subject, not a formula
There is no requirement that every thumbnail contain a face, a fixed number of words or a particular colour. The right composition depends on the video's subject and the promise made to the viewer. A clear object or scene can be more appropriate than a face.
View the original at a realistic small size. Check whether the subject and text remain understandable. A heatmap estimates visual saliency; it cannot judge whether the thumbnail accurately represents the video.
Upload and review
Use a PNG, JPEG or WebP up to 8 MB. The initial public allowance is limited. Processing happens on the server, and this tool stores the uploaded original and result with result URLs.
- Upload the final composition, including its text and background.
- Inspect the stronger predicted regions and identify which elements they overlap.
- Compare the original at a small display size to check legibility.
- Try one change, such as the crop or amount of background detail, and compare matched variants.
References: Privacy and result storage
Use the ranking as a review input
A higher score does not establish that viewers prefer the thumbnail or that it will get more clicks. A detail can attract predicted attention without communicating the video. Keep your editorial judgement and the viewer's expectations in the decision.
The tool's tips are rule-based prompts, not a semantic evaluation of the uploaded subject. Treat them as questions to investigate, especially when the advice does not fit the genre or intended audience.
Validate with viewers
YouTube provides an A/B testing workflow for eligible titles and thumbnails. Its own documentation explains the available testing conditions and how results are evaluated. That audience evidence is different from a model prediction on an image.
References: YouTube: A/B test titles and thumbnails