HomeAI Video ToolsVideo Subtitle Generator
Video Subtitle Generator

Subtitles that are always on, everywhere the clip goes

Speech is transcribed, aligned to the audio at word level, and drawn permanently into the exported video. Nobody has to enable anything, no companion file travels with the clip, and the text cannot be switched off by a platform setting or lost when the file is downloaded and posted somewhere else.

🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
  • Generated from the audio, no typing
  • Always visible — nothing to toggle on
  • Rendered into the frames, not a .srt
  • Same language as the speech, not translated
  • Survives reposts and cross-posting
One long video in, a week of posts out

Paste a link. Get the best moments.

Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.

https://youtu.be/z_bX3runikk
Source video: Matthew McConaughey's Concerns Over AI Source video 15:27
Matthew McConaughey's Concerns Over AI
9 clips found  ↓  top 5
Matthew McConaughey speaking, captioned vertical clip
Score 92
9:16
Joe Rogan gesturing, captioned vertical clip
Score 90
9:16
Matthew McConaughey listening, captioned vertical clip
Score 89
9:16
Joe Rogan mid-sentence, captioned vertical clip
Score 89
9:16
Matthew McConaughey talking to camera, captioned vertical clip
Score 88
9:16

How subtitles get onto the video

Three stages, and you are only involved in the first and the last.

  1. 1

    Provide the source

    A YouTube link or an uploaded file both work. Nothing needs preparing first — no transcript, no timing file, no chapter markers. Videos under thirty minutes run on the free demo, and paid plans accept up to two hours per submission.

  2. 2

    Speech is recognised and aligned

    The audio is transcribed and each word is matched to the moment it occurs. That alignment is what separates usable subtitles from a wall of text on a timer: the line on screen corresponds to what is being said right now rather than approximately, and it advances at the pace of the speaker.

  3. 3

    The text is rendered into the export

    Choose one of eleven presentations and the subtitles are drawn into the frames of the finished clip. What comes out is a single video file with the words already in it, which is the version that behaves predictably no matter where it ends up.

How these subtitles behave

Subtitles vary less in how they look than in how they behave. These are the behavioural properties worth understanding before you choose a tool.

Open by default — nothing to enable

Closed subtitles depend on the viewer having them switched on, and most viewers never touch that setting. Open subtitles are part of the picture, so every single person who sees the video sees the text. For social video, where you get one unrepeatable chance in the feed, that certainty is worth more than the flexibility you give up.

They cannot become detached

A sidecar file is a second file, and second files get lost. They are dropped by re-encoding, ignored by apps that do not support the format, and left behind whenever somebody downloads the video and uploads it again. Text rendered into the frames has none of these failure modes because there is nothing separate to lose.

Timed at the word, not the block

Conventional subtitle files hold a line for a fixed span and then swap it. Word-level alignment lets the text move with the speech, which reads more naturally on short vertical video and avoids the mismatch where the subtitle is a beat behind the mouth.

The language spoken is the language shown

Audio is transcribed rather than translated, so a German recording produces German subtitles. This is a limitation and it is stated deliberately: if you came looking for foreign-language subtitling, you need a translation workflow, and being told that now is better than discovering it after signing up.

Eleven presentations to choose from

The set covers restrained editorial styling through to the heavy treatment that suits fast-paced short-form. Switching between them re-renders the clip without re-transcribing, so comparing options on your own footage takes a minute rather than a project.

Trimming never desyncs anything

Move the start or end of a clip and it re-renders with the text recomputed against the new boundaries. The familiar problem of a subtitle file drifting out of alignment after an edit does not arise, because the words are placed during rendering rather than referenced from outside.

Applied to live-cut clips as well

Clips produced in real time from a YouTube, Twitch or Kick broadcast arrive subtitled like everything else. Speed is the whole reason to clip live, and a clip that needs a subtitling pass afterwards has lost the advantage it was created for.

Clean text rather than a raw transcript

Hesitations, restarts and repeated words are stripped before the text is placed, so what appears on screen reads as written language. A literal transcript of natural speech is accurate and almost unreadable, which is a distinction most automatic subtitle tools ignore.

Who needs subtitles on every clip

Publishers with deaf and hard-of-hearing audiences

On-screen dialogue is the baseline requirement, and always-on text means it is never dependent on a setting somebody forgot to turn on.

Anyone with an international audience

Reading a second language is markedly easier than hearing it at speed. Subtitles in the original language widen reach even without translation.

Internal comms and corporate video

Open-plan offices and shared spaces mean a lot of workplace video is watched silently, regardless of what the sender assumed.

Paid social advertisers

An ad that only works with sound wastes most of its impressions. Text on screen is the difference between a played impression and a understood one.

Interview and panel content

Crosstalk, varied microphones and accents make dialogue harder to follow by ear than creators expect. See the podcast clip generator for multi-speaker specifics.

Short-form creators generally

Feed video autoplays muted. Whatever you call the text, having it there is not optional if you want the first three seconds to work.

Subtitles, captions, and why the two words get swapped

Strictly, the terms describe different things. Subtitles render the dialogue for someone who can hear the audio but cannot follow it — historically because it is in another language. Captions are written for someone who cannot hear the audio at all, so alongside the dialogue they identify who is speaking and describe meaningful non-speech sound: a door slamming, music swelling, laughter off camera.

In practice the distinction has eroded, and it eroded differently in different places. British usage calls almost all on-screen dialogue text subtitles. American usage reserves captions, and specifically closed captions, for the accessibility version. Most people searching for a subtitle generator today simply mean "put the words on the screen", and both terms get used for that.

It is worth knowing where ClipSpeedAI actually sits. What it produces is the dialogue, transcribed from the audio and rendered on screen. It does not label speakers by name and it does not describe non-speech sound. By the strict definition that makes these subtitles rather than captions, and if your requirement is genuinely the accessibility variety with sound cues, you should know that before you rely on it.

Open subtitles versus a sidecar file

Closed subtitles live in a separate file — usually SRT or VTT — that the player reads and overlays. The viewer can turn them off, change the size, or choose a different language track if one exists. Search engines and platforms can read the text, which helps with indexing. It is the right architecture for a film, a course platform, or a long-form YouTube upload where the viewer is settled and in control.

Open subtitles are rendered into the picture. They cannot be turned off, resized, or read by a machine, and there is only one version. For a forty-second clip crossing four platforms and being reposted by strangers, all three of those constraints turn into advantages, because the failure mode you are actually defending against is text that silently fails to appear.

ClipSpeedAI produces the open kind. The engine is built around short clips destined for feeds, and in that context an unreadable dependency on viewer settings is a worse problem than the loss of a toggle. If you need a sidecar file for a long-form upload or a compliance workflow, use a dedicated captioning service for that deliverable — the two things can coexist perfectly well in one publishing routine.

What always-on subtitles do and do not solve

They solve the largest practical problem immediately: a viewer who cannot hear the audio can follow what is being said, with no setting to find and nothing to switch on. That covers the muted scroller and it covers a real portion of the deaf and hard-of-hearing audience, and it does so on every platform identically.

They do not make a video formally accessible in the sense a regulated organisation means. Burned-in text cannot be enlarged by someone with low vision, cannot be read aloud by assistive software, cannot be turned off by a viewer who finds it distracting, and does not carry the non-speech information that a proper caption track provides. If you are working to a legal standard, the honest answer is that this tool is not the compliance answer and should not be presented as one.

The reasonable position for most creators is that always-on subtitles on every clip is a large improvement over the common alternative, which is no text at all or platform captions the viewer never enabled. Treat it as a floor you never fall below, not as a certificate.

Subtitle mistakes that quietly cost you viewers

The most common one is placement. Text sitting near the bottom edge of a vertical video is covered by the description, the audio strip and the button column on every major app, and creators rarely notice because they preview their clips in a player that has none of that furniture over the top. Preview inside the app you are posting to, at least once.

The second is doubling up. If the platform adds its own automatic text on top of subtitles already rendered into the file, viewers get two overlapping layers of words. Switch the platform feature off for clips that arrive already subtitled.

The third is trusting the text without reading it. Automatic transcription is strong on ordinary speech and weak exactly where your content is distinctive — the product name, the technical term, the guest's surname. Those are the words a viewer will notice being wrong, and they are the ones most likely to be wrong.

The fourth is treating subtitles as a finishing touch. If you write or speak knowing the words will appear on screen, you tend to front-load the interesting clause and cut the throat-clearing, and the clip gets better before any software touches it.

A publishing routine that covers what pixels cannot

Because burned-in text is an image rather than characters, no platform can index it. A closed caption track on a long YouTube upload contributes text that search can read; subtitles rendered into a clip contribute none. This is a real trade rather than a detail, and it is worth handling deliberately.

The fix costs nothing and takes a moment. When you post, put the strongest line from the clip into the description or the on-platform text field, so the words exist in a machine-readable form alongside the visible one. Each clip here arrives with a suggested title taken from what was actually said, which is usually a serviceable starting point for exactly that.

If you also publish the long-form source, add a proper caption track there. The long upload is where indexable text pays off, and the clips are where always-on visible text pays off. Doing both is not duplicated effort — the two formats are being read by different things.

Troubleshooting a clip whose subtitles are not behaving

A clip arrived with no text on it. Give it a few minutes and reload the library before doing anything else. Rendering the words into a clip is a separate stage from producing the clip, it occasionally loses a race against everything else the machine is doing, and a background pass re-renders any clip that has a usable transcript but no visible text. Resubmitting the whole video is the instinct and it is the wrong move — you would be paying for a second run of a job that is already repairing itself.

The subtitles disagree with what was said. Two different faults look identical here. Either the audio was hard to transcribe, which shows up as plausible wrong words scattered through the clip, or the same term is wrong everywhere, which points at a proper noun the model has no evidence for. Scattered errors improve with a better microphone; a consistently wrong term will not improve at all, so decide up front whether you can live with it.

The words look right but sit in the wrong place. Almost always the frame changed after you last looked. Take the case of a widescreen source reframed to 9:16 — the crop is decided first and the text is laid out against the frame that survives it, so a subtitle you judged against the landscape master is not the one your audience sees. Judge placement on the vertical export only, since that is the file being posted.

Two sets of subtitles are stacked on top of each other. The platform added its own. Nothing in the file can prevent that, because from the outside a clip with subtitles rendered into it is just a video with no track attached, so the app assumes it needs to help. Switch its caption feature off per upload.

Settings and export choices that actually change the result

Text size is the setting with the most consequence and the least obvious behaviour. The layout is calibrated at the default, and the engine responds to a larger size by putting fewer words on each row rather than by letting a row run wide, so raising it makes the subtitles both bigger and more staccato. Nudge it rather than dragging it, and compare on a phone rather than on the monitor you set it from.

Aspect ratio is the next one. Vertical 9:16 is the tightest frame the text has to survive, and 1:1 and 16:9 exports inherit the same layout rule with more room to spend, which is why a subtitle that reads cleanly in the vertical version is safe everywhere else. Check the vertical first and the other two need no separate review.

Presentation style is worth choosing once and then leaving alone. Switching costs a re-render rather than a re-transcription, so experimenting is cheap in time, but a house style that changes every week costs you the recognition that consistent type quietly builds.

The one export decision people get wrong is source quality. Say you have both a compressed download of your own upload and the original master — send the master. Platform compression is applied to the audio as much as to the picture, and everything on screen here is derived from the audio.

When a subtitle file beats rendered text, and the alternatives to reach for

The honest test is whether the viewer is settled. A person who chose your video, pressed play and is watching it on a big screen benefits from control: text they can size, disable, or swap for another language. Everything about always-on rendered text is aimed at the opposite viewer, the one who did not choose the clip and is deciding in under two seconds whether to keep watching. Serving the first viewer with the second viewer's format is a mismatch, not a bug.

So for a long-form upload, an archive, a course platform or anything an institution will store, generate a caption file with a service built for that output. Those alternatives exist, they are inexpensive relative to what they cover, and they produce something this deliberately does not: a text layer separate from the picture.

The same applies to any legally mandated caption, where speaker identification and non-speech sound description are part of the requirement rather than optional extras. This tool writes dialogue and nothing else.

And if your source is a bare audio file with no video at all, there is nothing here to render subtitles onto. That is a different job, and a transcription service that emits a timed file is the sensible route to it — you can bring the result into an editor over a static frame, which is not what this engine is shaped for.

Automatic subtitling compared with doing the timing yourself

Hand-timed subtitling is not merely slower; it is a different discipline with its own conventions — reading-speed limits measured in characters per second, line breaks placed at grammatical joints, a minimum duration per cue so the eye is never rushed. Broadcast subtitlers work to those standards and the results are better than anything automatic, on the specific axis of comfort over a long sitting.

None of that is the constraint on a forty-second vertical clip. The reading load is set by how fast the person spoke, the viewer is not settling in for an hour, and the failure you are defending against is not fatigue but a silent scroll in the first second. Automatic word-level timing is well matched to that problem and the manual craft is aimed at a different one.

Where the manual approach still wins outright is anything requiring judgement about meaning: condensing a rambling sentence into readable text, choosing not to subtitle a deliberate pause, or writing a line that differs from the words spoken because the literal version reads badly. Automated subtitling writes what was said. Sometimes what was said is not what should be on screen, and only a person can decide that.

What shipping subtitles on every clip taught us

Observations from operating the pipeline in production — not general advice.

A renderer reporting success is not evidence that anything was drawn

One caption path composites a transparent overlay over the video. When that overlay is generated but nothing paints into it — a font that never loaded, a frame captured before the page settled — the result is a perfectly valid file of entirely transparent frames. Every check we had passed: the file existed, it had a size, the composite exited cleanly, the database recorded captions as done, and the clip shipped visibly bare. The gate now samples the overlay transparency at the timestamps where words are meant to be, because a flag written by the thing you are checking is not evidence about the thing you are checking.

The straggler clip is a scheduling problem, not a content problem

For a long time, one or two clips per video would come back without text while the rest of the batch was fine, and the natural assumption was something wrong with those particular clips. It was not: their transcripts were clean and the render succeeded immediately when retried. They had simply missed a deadline under render contention and been marked terminal, and every recovery loop skipped terminal work by design. A standing pass now re-renders any clip with a real transcript and no visible text regardless of its recorded state, which is why a bare clip usually fixes itself within minutes.

All eleven presentations run the same layout engine underneath

It is natural to assume each preset carries its own idea of how many words belong on screen, and early on they did. They no longer do: chunking, row count, sizing and position are one shared calculation, and a preset now contributes only visual treatment — typeface, colours, the plate behind the text, the highlight behaviour. The reason is that layout is where subtitles break, so it needed exactly one implementation to audit rather than eleven. In practice it means choosing a style changes how your subtitles look without changing where they sit or how they break.

Compared with other subtitling routes

vs. writing an SRT by hand

Hand-timing a subtitle file gives you complete control and takes a genuinely unpleasant amount of time. For a feature film or a course module it can be justified. For a stream of short clips it is unsustainable, and the version people actually ship is the one that got generated automatically.

vs. relying on the platform to subtitle it

YouTube, TikTok and Instagram will each generate their own text over your upload. You then have three different-looking versions of the same clip, styled by three companies, some of which the viewer has to opt into. Deciding the appearance yourself and rendering it in is the only way the clip looks the same everywhere.

vs. a dedicated captioning service

Professional captioning with human review, speaker identification and sound descriptions is the correct choice when you have a legal obligation or a broadcast standard to meet. It is priced and paced accordingly. This is a different job: subtitles on volume short-form, generated in the same pass that produced the clip.

vs. subtitling inside an editor

A desktop editor gives you frame-level control over every line, if you already have a project open with the clip in it. When the clip does not exist yet, subtitling is the second half of a job whose first half is finding and cutting the moment — which is what the AI clip generator does before the text is ever placed.

Frequently asked questions

How do I generate subtitles for a video?
Paste a link or upload the file and subtitles are produced automatically as part of processing. The audio is transcribed, the words are aligned to the timeline, and the text is rendered into the finished clip. You do not supply a transcript or set any timings.
What is the difference between subtitles and captions?
Subtitles present the dialogue, on the assumption the viewer can hear but needs the words. Captions are written for viewers who cannot hear at all, so they also identify speakers and describe important sounds. In everyday use the terms have blurred, but the distinction still matters if you have an accessibility requirement to meet.
Are the subtitles burned in or a separate file?
Burned in. The text becomes part of the image, so the clip is one self-contained file with nothing to attach or upload alongside it. That is the correct choice for social video and the wrong one if you specifically need a toggleable track.
Can I download an SRT or VTT file?
The deliverable is a video with the subtitles rendered into it, not a sidecar file. If a subtitle file is what your workflow requires — for a long-form upload, an LMS, or a compliance archive — a dedicated captioning service is the right tool for that particular output.
Can viewers turn the subtitles off?
No, and that is intentional. Open subtitles are visible to everyone with no setting to discover, which is the behaviour you want in a feed where the viewer decides in under two seconds. The cost is that someone who dislikes on-screen text cannot remove it.
Will it translate subtitles into another language?
It will not. Speech is transcribed in whatever language it was spoken, so an Italian video comes back with Italian subtitles. Translation is a separate capability and we do not offer it — claiming otherwise would waste your time.
Which languages are supported?
Transcription covers the major languages, and English is the strongest by some distance. Because accuracy is affected as much by microphone quality and crosstalk as by language, one test video through the free demo will tell you more than a support list would.
Are these subtitles ADA or WCAG compliant?
They should not be treated as a compliance solution. Burned-in text cannot be resized, disabled or read by assistive technology, and it carries no non-speech sound information. If you are subject to a formal accessibility standard, use a captioning workflow built for that and treat these subtitles as a separate, additional layer.
Do the subtitles identify who is speaking?
No speaker labels are added. In short vertical clips the speaker is almost always visible in frame, and name tags consume space that the words need. For a panel where identification genuinely matters, the split-screen and grid layouts make the speaker obvious visually.
Are non-speech sounds described?
No. Applause, music and sound effects are not written out. This is one of the specific ways these are subtitles rather than full captions, and it is worth being aware of if that information is part of your requirement.
How accurate is the subtitle text?
Reliable on clean, clearly recorded speech and weaker where there is heavy background music, people talking over each other, or unusual names and technical terms. We publish no accuracy figure on purpose, since a single number across those conditions would be misleading.
Can I edit the subtitle text?
The text is generated from the transcript, so the practical approach is to check the first clip from a new source. If a particular term is being misheard consistently, you will spot it immediately rather than after publishing a run of clips containing it.
Can I change how the subtitles look?
Yes, there are eleven presentations covering everything from understated to very bold. Selecting a different one re-renders the clip. The caption generator page goes into what makes each treatment readable on a phone.
Do subtitles get in the way of the platform interface?
The text is positioned to stay clear of the region where short-form apps stack their own controls. Subtitles placed too low are covered by the description and buttons, which is a common and entirely avoidable mistake.
Will search engines index my subtitles?
They cannot, because the words are pixels rather than text. The straightforward workaround is to put your strongest line into the post description when you publish, giving the platform something readable alongside the visible version.
Can I subtitle an entire long video rather than clips?
The engine produces clips, so subtitling is part of that output rather than a standalone service for full-length files. If you want a two-hour video subtitled as a two-hour video, this is not the right shape of tool.
Do subtitles appear on live-clipped videos?
Yes. Clips cut during a live YouTube, Twitch or Kick broadcast come back with the text already rendered, which preserves the speed advantage that makes live clipping worth doing.
Should I also turn on the platform captions?
Generally no. If TikTok or Instagram adds its own caption layer on top of subtitles that are already in the video, you get two sets of text overlapping. Switch the platform feature off for these uploads.
Is there a watermark?
On free demo output there is a small one. Paid exports carry no ClipSpeedAI marking, and the Brand Kit can apply your own logo across everything you produce.
What does the subtitle generator cost?
The demo is free for videos under thirty minutes with no card needed. After that it is one dollar for a three-day trial and twenty-nine dollars a month for Pro, and cancelling takes a single click.
Is there a limit on how long the video can be?
Up to two hours on a paid plan and thirty minutes on the free demo. Live broadcasts have no equivalent ceiling, since clipping runs continuously for as long as you stay on air, and a file longer than two hours can be split at a natural break.
Will trimming a clip break the subtitle timing?
It will not. Changing the boundaries re-renders the clip and the text is recomputed against the new edit, so there is no realignment step. This is one of the practical advantages of rendering text in rather than referencing it from a file.
Can I generate subtitles through an API?
Yes. The engine is available programmatically and also as an MCP connector, so a script or Claude itself can submit videos and collect the subtitled clips. Full details sit in the developer documentation.
Does subtitling work with screen recordings?
It does. The screenshare layout keeps the shared window readable while the subtitles sit clear of it, though very small on-screen text in the recording remains hard to read on a phone no matter how the frame is arranged.
Can I subtitle a video I did not record?
The tool will process any source you can point it at. Whether the finished clip can be published depends on the rights attached to that material and on the policies of wherever you post it, which is a judgement for you to make.
What happens to my video after subtitles are generated?
It is processed to create your clips and nothing further — we do not publish or redistribute it. The specifics of storage are set out in the privacy policy.
How do I judge the quality before committing?
Use the free demo on a recording with your real conditions: your microphone, your accent, your vocabulary, your typical background noise. Sample footage from any vendor is recorded in a quiet room, and yours probably was not.

Put the words on screen, permanently

Generate a clip and check the text on a phone at arm's length. Legibility at that distance is the whole test.

⚡ Get A.I Clips — $1 trial
3-day trial · just $1 · cancel anytime