11 September 2026 · 4 min read
Captions that get read with the sound off
Most people will see your clip muted. That makes captions the primary channel, not an accessibility add-on, and it changes what good captions look like.
Assume the sound is off. Not as a concession, as the default: a large share of feed video is watched muted, at least until something earns the tap. That makes your captions the main way the clip communicates for the first few seconds.
Word-level timing beats line-level
Captions that appear a full line at a time force the viewer to read ahead of the speaker, which is a strange experience and kills any sense of timing. A joke lands before it is said.
Word-by-word highlighting, where the whole line is visible but the current word is emphasised, keeps the reading pace tied to the speaking pace. It is the reason that style has taken over short-form: it preserves delivery.
Size for a phone held badly
Preview your captions at the size they will actually be seen — a phone at arm's length, in a feed, possibly in sunlight. Text that looks generous on a laptop is often marginal there.
Two practical rules: no more than about four or five words visible at once, and enough contrast that the text survives a bright frame. A drop shadow or a slight outline does more for legibility than any font choice.
Where they sit matters more than people think
Every platform overlays its own interface on your video. Captions in the bottom fifth of the frame end up behind usernames, descriptions and buttons depending on where the clip is posted.
Keeping text in the middle third vertically is the safest choice if a clip might be posted to more than one platform, which it almost always is.
Check the transcript, especially for names
Automatic transcription is very good and still wrong in the places that embarrass you: proper nouns, product names, technical terms, anything said over a laugh. A misspelled guest name in burned-in captions is not fixable after posting.
Read the first line of every clip before it goes out. It takes seconds and it catches the errors that matter most, because the first line is the one everybody reads.
- Word-level highlighting, not whole lines at once
- Four or five words visible at a time
- Middle third of the frame, out of the platform interface
- High contrast, with a shadow or outline
- Proofread the first line of every clip, and every proper noun
A note on burned-in versus platform captions
Platform auto-captions are free and improving, but they are styled by the platform, can be switched off by the viewer, and are not present when somebody downloads and reposts the clip. Burning them in means the clip carries its own words wherever it ends up.
Keptbits cuts on sentence boundaries.
Paste a link or upload an episode, say how many clips you want, and get them back captioned and vertical. Seven days free on any plan.
Start your free trialRead next
Burned-in captions or uploaded subtitles?
Two ways to put words on a video, and they are not interchangeable. Which to use for short-form, which for long-form, and why the answer is sometimes both.
Why a transcript's timing decides where the cut lands
The difference between a transcript that knows when each sentence started and one that knows when each word started is the difference between a clean clip and a clipped syllable.
9:16, 1:1 or 16:9: which ratio for which platform
Three shapes, and the choice is less about platform rules than about how much of the frame you are willing to throw away. A practical guide to picking one.