Most short-form video on mobile is played without sound. Autoplay in feeds is muted by default on most platforms, and a significant share of viewers never unmute. If a clip's text caption is the first moving element visible, it is doing the work of the hook before the viewer decides to tap audio. The styling of that text -- size, position, contrast, timing -- determines whether the viewer follows along or scrolls past.
Burn-in vs. overlay
Burn-in captions are embedded into the video frame as pixels -- they are part of the image itself and display identically on every device and every player. Overlay captions are delivered as a separate caption track or as JavaScript-rendered text over the video element. They depend on the player supporting and displaying the caption track, which is not guaranteed across all feed contexts.
For short-form clips on TikTok, Reels, and Shorts, burn-in is the correct choice. These platforms do not consistently expose a caption track toggle in their in-feed players, and the native auto-caption features these platforms offer are inconsistent in accuracy and styling. Burn-in gives the creator control over appearance and guarantees the text is visible in every viewing context, including when the video is downloaded and reshared by another account.
Word-by-word vs. sentence display
Word-by-word captioning highlights one word at a time as the speaker says it. Full-sentence captioning displays the complete sentence and holds it until the next sentence begins. Neither approach is universally better -- performance depends on the pacing of the specific speaker and the density of information in the content.
For fast-paced commentary and gaming content, word-by-word captioning keeps the viewer's eye moving with the audio at the same rhythm as the speech. For explanatory or analytical content where each sentence carries a complete idea, sentence-level display is easier to read because the viewer does not have to piece together meaning word by word. The common mistake is applying word-by-word timing to slow, deliberate speech -- it produces a staccato caption effect that feels out of sync with the speaker's cadence.
Contrast and placement
Caption contrast is a readability issue that varies based on the background content of each clip. White text over bright footage is unreadable. The standard solution is a semi-transparent dark background behind the text, which provides consistent contrast regardless of what is behind it. Some creators prefer a text stroke (a dark outline on light text) for a less opaque look, but this approach fails on complex or busy backgrounds.
Placement within the safe zone matters more than many creators expect. Captions placed too low get covered by the platform UI. Captions placed too high sit in the area where the platform tends to show creator information on certain views. The practical sweet spot is the lower-middle of the safe zone -- low enough to not compete visually with the subject's face, high enough to clear the interaction buttons. This positioning also works with natural reading direction: the viewer's eye tracks downward from the subject's face to the caption text.
Font weight and legibility
The caption font that appears in many automated clip exports is whatever the underlying tool defaults to. This is often a thin sans-serif that reads well on a desktop monitor but loses legibility at the small sizes typical in a mobile feed. For short-form captions, the font needs to be heavy enough to read at 375px wide. Weights below 600 tend to disappear at small sizes, especially on OLED screens in outdoor conditions. A bold sans-serif at a size large enough to read without zooming is more important than matching the aesthetic of the creator's brand kit -- if the caption is unreadable, it does not matter how on-brand it looks.