Most people meet your video with the volume already down, which means the text on screen is doing the work the audio was supposed to do. Captions here are generated from the speech, timed word by word, styled eleven different ways, and rendered into the video itself rather than attached beside it.
🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
⏱ That video's a bit long for the demo
Words highlight in time with the voice
Eleven caption styles to pick from
Rendered into the video, not a side file
Filler words removed before captioning
Applied to every clip automatically
One long video in, a week of posts out
Paste a link. Get the best moments.
Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.
Captioning is not a separate task here. It happens as part of producing the clip.
1
Send in the video
Paste a link or upload the file. The audio is transcribed as part of processing, and there is no separate captioning step to start, no transcript to approve first and no timing file to import. If you only want to sample the output, the free demo runs on anything under half an hour.
2
Speech becomes timed text
The transcript is aligned to the audio at the level of individual words, not blocks of a few seconds. That granularity is what allows a word to light up exactly as it is spoken, and it also means the text never drifts behind a fast talker or hangs on screen after a sentence has ended.
3
Pick a style and export
Eleven styles are available, from clean sans-serif blocks to the bold high-contrast look that dominates short-form feeds. Switch between them and re-render to compare. The captions are drawn into the frames of the exported file, so what you download is what every viewer sees on every platform.
What the caption engine gets right
Captions look simple until you have shipped a hundred of them. These are the details that decide whether people keep watching.
⏱
Word-level timing, not block timing
Traditional captions push two lines at a time and swap them every few seconds, which is fine for a film and dull in a feed. Aligning each word individually lets the caption track the rhythm of the speech, so the emphasis on screen matches the emphasis in the voice. The active word is highlighted as it lands.
🎨
Eleven styles, all built for small screens
A caption competing with footage behind it needs weight, contrast and an outline or shadow to stay legible. Every style here is designed against that constraint rather than being a font choice with a colour picker attached, which is why they hold up on a phone in daylight.
🔒
Burned into the pixels
The text becomes part of the image. Nothing to upload alongside the video, nothing that a platform can strip, nothing that breaks when the clip is downloaded and reposted by somebody else. Wherever the file travels, the captions travel inside it.
🧹
Filler is cut before the text is drawn
Verbatim transcription is a terrible caption. Real speech is full of "um", "you know", restarted sentences and half-finished clauses, and putting all of that on screen makes an articulate person look scattered. Those are removed from the audio and never reach the caption track.
📱
Placed where a phone can read it
The bottom of a vertical video is covered by the caption box, the handle, the music strip and the button column, and text placed there is simply lost. The styles position the text in the band of the frame that stays visible, which is a boring detail that decides whether the captions were worth generating at all.
🎬
Captions on every clip, without asking
There is no checkbox to remember. Any clip the engine produces comes back captioned, including clips cut from a live YouTube, Twitch or Kick broadcast in real time. Consistency here is worth more than configurability — the clip you forgot to caption is the one that underperforms.
🖼
Text that survives the vertical crop
Reframing widescreen to 9:16 while burning in captions is where a lot of tools produce something with the words half off the edge. Because the crop and the caption layout are decided together, the text fits the final frame rather than the original one.
🏷
Consistent across everything you post
The Brand Kit keeps your fonts, colours and logo applied across exports so a run of clips looks deliberate. A viewer scrolling past their third clip of yours in a week should recognise the look before they read the name on the account.
Where automatic captions matter most
🎙 Talking-head and interview clips
Nothing but speech carries the clip, so the caption is effectively the entire visual argument for staying.
🎮 Gaming and reaction content
Chaotic audio, overlapping voices and game sound make text the only reliable channel. Picture-in-picture keeps the face and the captions both readable.
📈 Marketing and sales clips
A product point delivered on mute has to land in text or it does not land. Captions are the difference between a scroll and a stop.
🏫 Teaching and explainer video
Technical terms are easier to absorb read than heard, particularly for anyone watching in a second language.
♿ Accessibility-conscious publishers
On-screen text is the baseline for viewers who are deaf or hard of hearing. The video subtitle generator page covers what burned-in text does and does not satisfy.
🛍 Creators posting in noisy places
Commutes, offices, waiting rooms. Your audience is frequently somewhere they cannot turn the sound on even if they want to.
The caption is the video for a large share of your audience
Short-form feeds autoplay silently by default, and plenty of viewers never unmute anything. The practical result is that your opening line is read before it is heard, if it is heard at all. A clip with no text on screen is asking a stranger to make an extra decision in the first second, and most of them decline.
This changes what a caption is for. It is not a transcript service or an accessibility afterthought bolted onto the export — it is the primary interface of the video. Once you accept that, the styling decisions stop being cosmetic. Size, weight, contrast and position are the difference between someone reading your first sentence and someone scrolling past it.
It also changes what you write and how you speak. Creators who caption everything tend to front-load the interesting clause, because they can see it appearing on screen at second zero rather than arriving in the middle of a sentence nobody heard.
What makes a caption readable at arm's length
Contrast comes first. Text over video is text over an unpredictable background, so the styles rely on heavy weights with outlines or plates behind them rather than thin type that vanishes the moment the shot brightens. A caption that is legible in a dark studio and invisible against a window is not legible.
Then pace. Words appearing faster than they can be read produce a subtly stressful clip that people leave without knowing why. Because the timing here is tied to the actual speech rather than to fixed intervals, the caption naturally slows when the speaker does, and the reading load stays proportional to the delivery.
Then position, which is the one most people get wrong. Every platform stacks its own interface over the lower portion of a vertical video, and captions parked in that zone are covered by the description, the audio strip and the row of buttons. Keeping the text above that band is unglamorous and it is the single highest-return caption decision you can make.
Finally, restraint. Every word on screen at once is a wall; one word at a time can feel jittery on a slow talker. The styles here sit between those extremes, showing a readable phrase with the current word emphasised, which reads as movement without becoming noise.
Why the platform's own auto-captions are not enough
TikTok, Instagram and YouTube all offer to caption your upload, and those captions are better than nothing. They are also small, styled to the platform rather than to you, positioned wherever the platform decides, and frequently switched off by default depending on the viewer's settings. You are handing your most important design decision to somebody else.
They also do not travel. A clip downloaded from one app and re-uploaded to another arrives with no captions at all, which is exactly the workflow most creators actually use. Burned-in text is immune to this: the words are in the frames, so they survive every download, repost and cross-post.
The last problem is timing. Platform captions typically chunk text into blocks that appear and disappear on a fixed cadence. That is serviceable for comprehension and useless for pace, which is what you want captions doing in a feed.
A captioning routine worth copying
Pick one style and stay with it. Rotating through all eleven because each looks good in isolation costs you the thing captions can quietly buy — recognition. Somebody who has scrolled past your clips twice should register the third as yours before they read the account name, and consistent type does more of that work than a logo does.
Check the first clip from every new recording setup. Not every clip, just the first one after you change microphone, room or guest, because those are the changes that move transcription accuracy. Thirty seconds of checking catches the mistake that would otherwise repeat across fifteen exports.
Turn the platform's own caption feature off when you upload. If TikTok or Instagram lays its automatic text over a clip that already has words rendered into it, the two sets overlap and the result looks careless. This trips up almost everyone the first time.
And read your captions rather than watching them. Mute the clip, look only at the text, and see whether the opening line still makes an argument for staying. If it does not read well as a sentence on a screen, the clip is relying on delivery that a silent viewer will never receive.
Where automatic captioning gets things wrong
Speech recognition is very good and it is not perfect, and it fails in specific, predictable places. Proper nouns are the worst offender — brand names, product names and unusual surnames are exactly the words a model has the least evidence for and exactly the words you least want mangled. If your show is full of them, check the first clip before you build a habit on the output.
Heavy background music, two people talking over each other, and low-quality recordings all reduce accuracy. So does very fast speech with a strong regional accent. None of these are unique to this tool; they are the boundaries of automatic transcription generally, and any product claiming immunity from them is overselling.
One thing worth being explicit about: the captions are in the language being spoken. Speech is transcribed, not translated, so a French video produces French captions. If you need a different language on screen than the one in the audio, that is a translation workflow and this is not the tool for it.
Troubleshooting captions that came back wrong
A word runs off the edge of the frame. Caption size is stored per clip, and the layout maths is calibrated around the default value rather than around whatever the slider was last dragged to. Above the default the engine allows fewer characters on each row automatically, so enlarging the type gives you bigger words on more rows instead of a line bleeding off the side. When one clip in a batch overflows and the other fourteen are fine, open that clip and read the size control before you suspect the renderer.
A word appears that nobody said. Speech models label audible events as well as speech, and a strong music bed can arrive in the transcript as a literal word. For example, a scored intro under a cold open can put the word "music" on screen, and the giveaway is the same phantom word surfacing in several unrelated clips cut from one recording. Mixing the bed lower under the voice removes it at the source, which is the only place it can be removed.
The bottom line of text is hidden. You almost certainly checked the clip in a plain video player, which carries none of the description, audio strip and button column that a short-form app stacks over the lower third. Post one clip to a private account and look at it there instead. Doing that once tells you where the real safe band sits on the platform you publish to, and you will never have to think about it again.
One name is wrong in every single clip. That is transcription rather than captioning, and it does not vary between clips from the same recording — the model will mishear a surname identically fifteen times. Run thirty seconds containing the name through the demo, see exactly what comes back, and then decide whether to introduce it differently on camera or accept it. Learning this from one clip costs nothing; learning it from a published batch costs a correction.
When not to burn captions in, and the alternatives that fit those jobs better
A long-form upload that somebody sits down to watch is the clearest case against. Imagine a forty-minute recorded lecture going onto a course platform or a television app: that viewer wants the option to switch text off, and a caption track gives them the toggle while also handing the platform something machine-readable to index. Rendered-in words do neither. For that deliverable a separate caption file is simply the better instrument, and the subtitle generator page maps out exactly where the line between the two falls.
Formal accessibility work is the second case. A track that assistive software can read aloud, resize and disable is a different artefact from words painted into a frame, and if you are working to a published standard, a captioning service with human review is what actually satisfies it. Treating on-screen text as a substitute for that is the kind of shortcut that gets discovered during an audit.
The third is any job that needs the words in a language other than the one being spoken. Speech is transcribed and never translated here, so a translation workflow has to sit entirely outside this tool rather than being configured inside it.
And if you honestly publish one video a month, the alternative worth using is the caption feature already sitting in your phone editor. It has become competent. The argument for an engine is volume and consistency across a run of clips, and at one clip a month you have neither of those problems to solve.
What captioning by hand still buys you
Automatic captioning is a transcription problem plus a layout problem, and both are solved well enough now that the output is usually indistinguishable from careful manual work. Where it is not is emphasis that depends on meaning rather than on sound: a caption held back half a beat so a punchline lands after the setup, a word deliberately oversized for a joke, text that moves with a gesture. Those are authored decisions and no timing model is going to make them on your behalf.
The same is true of any typography beyond dialogue — a lower third naming a guest, a figure pulled out and animated because it is the entire point of the clip, a callback appearing in a corner thirty seconds after the line it refers to. That work belongs in an editor and it always will.
The division most people settle on is unglamorous and it holds. Say you cut fifteen clips from one recording and one of them is going behind paid promotion: let the engine caption all fifteen so they exist and match each other, then take that one into an editor and finish it properly. Trying to hand-finish every clip is the habit that dies in week three, and captioning nothing while you wait for time to do it properly is the worst outcome of the three.
Three things captioning a lot of real audio has taught us
Observations from operating the pipeline in production — not general advice.
The same word twice in a row is usually the room, not the speaker
Close-mic talking-head audio reflects off a desk or a wall and returns to the microphone roughly 50 to 200 milliseconds behind the voice, and a speech model can hear that reflection as a second utterance. Rendered literally it ghosts the caption: THAT THAT, WAIT WAIT. The rule that removes it is deliberately narrow — a repeated word is dropped only when it directly follows the first and lands within four tenths of a second — because natural speech that repeats a word almost always has another word between the two. Widen the window and you start deleting somebody genuinely shouting "go go".
Two rows, three words, and a budget that tightens as the type grows
Every style shares one chunking rule rather than each preset carrying its own: never more than two rows on screen, never more than three words on a row, a budget of roughly 16 characters per row, and a word is never broken across rows. The part worth knowing is what happens when you enlarge the type, because the character budget contracts in step with it. A larger font therefore buys you fewer words per row rather than a line running off the frame, which is why pushing caption size up changes the rhythm of the text as much as its scale, and why one clip with the size raised reads noticeably choppier than the rest of a batch.
A chunk of nothing but small words gets no highlight at all
The moving highlight picks the longest meaning-carrying word in each group and steps over the function words — the, of, and, to, pronouns, auxiliary verbs. When a group happens to contain nothing but those, the highlight is suppressed for that group instead of landing on "OF", because a colour parked arbitrarily on a preposition reads as a glitch to anyone watching closely. In practice this is what makes the emphasis look considered rather than metronomic: it follows the words carrying the sentence and goes quiet across the connective tissue between them.
Compared with the other ways to caption
vs. typing captions by hand
Manual captioning on a timeline is among the least rewarding tasks in video work. A minute of speech takes many minutes to type, time and position, and it has to be redone every time you change a cut. The result can be excellent. Almost nobody sustains it past the first few videos.
vs. a transcript tool that exports a file
Transcription services will give you accurate text with timestamps, which is genuinely useful if you want a document. It still leaves you to import the file into an editor, style it, position it and render — the part that actually takes the time. Here the styled, positioned, rendered result is the output.
vs. the caption feature in a mobile editor
Phone editors have decent auto-caption tools now, and if you are producing one clip that is a fine place to do it. The difference shows at volume: fifteen clips means fifteen sessions of importing, captioning and exporting on a small screen, and by clip nine the styling has quietly drifted.
vs. captioning inside a desktop editor
Premiere and Resolve can caption automatically and give you total control over the result. They also require the project to exist, which means you have already cut the clips. If you are generating the clips anyway, captioning them in the same pass removes an entire round trip — see the clip generator for the rest of what happens in that pass.
Frequently asked questions
How do I add automatic captions to a video?
Submit the video by link or upload and the captions are generated as part of processing. There is no separate step to trigger, no transcript to approve first and nothing to import. The clip you download already has the text drawn into it.
What does word-by-word captioning mean?
Each word is timed individually against the audio and highlighted at the moment it is spoken, rather than a block of text appearing and sitting still for three seconds. It reads as movement, which holds attention in a feed, and it keeps the text locked to the delivery.
How many caption styles are there?
Eleven, ranging from restrained sans-serif to the heavy high-contrast look common in short-form. You can switch style and re-render to compare on your actual footage, which is a faster way to choose than looking at previews.
Are the captions burned into the video?
Yes. They are rendered into the frames rather than delivered as a companion file, so nothing has to be uploaded alongside the clip and no platform can strip or restyle them. It also means they survive being reposted by other people.
Can I get a separate subtitle file instead?
The output here is a video with the text rendered into it, which is the right shape for short-form posting. If your workflow specifically needs a sidecar file for a long-form upload, the subtitle generator page explains the difference and when each one is appropriate.
How accurate is the transcription?
Good on clear speech and noticeably weaker on overlapping voices, heavy music beds, strong accents and unusual proper nouns. We deliberately publish no accuracy percentage because the number would be meaningless across those conditions — run one of your own recordings through the demo and judge it on your audio.
Will it caption filler words like "um"?
No. Hesitations, false starts and repeated words are removed from the clip before captions are generated, so they never appear on screen. Verbatim captions make fluent speakers look unsure, which is not what anyone is after.
Does it caption multiple speakers?
It captions everything spoken, and in a back-and-forth conversation the text simply follows whoever is talking. It does not label who is speaking on screen, which is a deliberate choice for short-form where the visual usually makes that obvious.
Can it translate the captions into another language?
No. The text is a transcription of what is actually said, in the language it was said in. Translated subtitles are a genuinely different product and we would rather point that out than let you find it after paying.
What languages can it caption?
The major world languages are supported, with English the most reliable. Quality depends as much on recording conditions as on the language itself, so the honest recommendation is a single test video in the language you work in.
Will the captions be covered by the TikTok interface?
They are positioned to avoid it. The lower portion of a vertical frame is occupied by the description, the audio strip and the button column on essentially every short-form app, so the styles keep text above that region where it stays readable.
Can I change the caption style after generating a clip?
Yes. Pick a different style and re-render, as many times as you want. Because the timing data is already computed, switching styles does not require re-transcribing anything.
Do captions work on clips from a live stream?
They do. Clips cut in real time from a YouTube, Twitch or Kick broadcast come back captioned like any other clip, which matters because the whole point of live clipping is posting before the moment goes cold.
Is there a watermark on captioned clips?
Only on the free demo, and it is small. Paid exports are clean, and the Brand Kit can place your own logo on them instead if you want one.
Is the auto caption generator free?
The demo is free for videos under thirty minutes and needs no card. Continuing costs a dollar for a three-day trial, then twenty-nine dollars a month, cancellable in one click.
Do captions actually improve retention?
The mechanism is straightforward rather than magical: most feed viewing starts muted, so a clip without text asks people to act before they can understand it. We are not going to quote you a percentage, but the reason every serious short-form account captions everything is not aesthetic.
Can I use all capitals for the captions?
Some of the eleven styles are set in capitals and others are not, so choosing the style chooses the case. Capitals read as louder and are common in high-energy content; sentence case tends to suit longer or more considered lines.
Does captioning slow down processing?
Not meaningfully. Transcription happens as part of the same pass that selects and cuts the clips, so it is not additional waiting stacked on the end of the job.
What if a word is transcribed incorrectly?
Most errors cluster around names and jargon. If a specific term matters to your content, check it on a demo clip first — you will find out immediately whether it is being heard correctly, and that is worth knowing before you publish twenty clips containing it.
Can I caption a video without generating clips?
The engine is built around producing clips, so captioning is part of that flow rather than a standalone utility. If your goal is a captioned version of a whole long video, this is not the shape of tool you want.
Do captions work with the screenshare layout?
Yes, and the layout is taken into account when placing them so text does not sit over the shared content. Dense slides are still difficult to read on a phone, which is a property of the source material rather than of the captions.
Will captions stay in sync if I trim the clip?
Yes. Adjusting the in and out points re-renders the clip with the timing recalculated, so there is nothing to nudge back into place. Desynced captions after a trim are a classic sidecar-file problem that burned-in text does not have.
Can I add emoji or highlight keywords?
The styles handle emphasis through the word-level highlight rather than through decoration you place manually. This keeps a run of clips consistent, which matters more for an account than any individual flourish does.
Does it caption profanity?
It transcribes what was said. If you need language cleaned up for a particular platform, that is an editorial decision to make before recording or after export.
Can I caption videos through an API?
Yes, the same engine is available programmatically for teams processing volume, and there is an MCP connector so Claude can run jobs conversationally. Both are documented in the developer reference.
Do captions appear on square and landscape exports too?
They do, laid out for whichever aspect ratio you export. The 9:16 version has the tightest constraints, so if the text works there it works everywhere.
How do I know the captions are good before I pay?
Run the free demo on a recording with your real audio conditions — your microphone, your room, your vocabulary. Marketing samples are always recorded in ideal conditions, and yours are the only ones that tell you anything.
Caption a clip and watch it on mute
Play the result with the volume all the way down, which is how most of the feed will meet it anyway. If the point still lands in silence, the captions are doing their job; if you catch yourself reaching for the volume to follow the sentence, the text on screen is not carrying enough of it yet.