HomeAI Video ToolsBurn Subtitles Into Video
Burn Subtitles Into Video

Subtitles baked into the pixels, not attached to the file

A burned-in caption is part of the picture. It cannot be switched off, lost in a re-encode, stripped by a platform, or left behind when somebody downloads your clip and posts it somewhere else. Send in a long video and every clip that comes back has word-by-word subtitles rendered into the frames, timed to the speech, with no sidecar file to attach and nothing to re-sync after a trim.

🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
  • Captions rendered into the frames
  • Word-by-word highlight timing
  • Eleven caption styles
  • Positioned clear of platform UI
  • Nothing to re-upload or re-sync
One long video in, a week of posts out

Paste a link. Get the best moments.

Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.

https://youtu.be/z_bX3runikk
Source video: Matthew McConaughey's Concerns Over AI Source video 15:27
Matthew McConaughey's Concerns Over AI
9 clips found  ↓  top 5
Matthew McConaughey speaking, captioned vertical clip
Score 92
9:16
Joe Rogan gesturing, captioned vertical clip
Score 90
9:16
Matthew McConaughey listening, captioned vertical clip
Score 89
9:16
Joe Rogan mid-sentence, captioned vertical clip
Score 89
9:16
Matthew McConaughey talking to camera, captioned vertical clip
Score 88
9:16

How to burn subtitles into a video

Three steps and no subtitle editor at any point.

  1. 1

    Send the video in

    Paste a YouTube link or upload the file. There is no separate transcription step to run first and no SRT you have to source from somewhere else — the speech is read directly off the audio track of whatever you submit. A paid upload can run to two hours; the free demo takes sources under thirty minutes.

  2. 2

    The speech is transcribed and timed word by word

    Timing is the part that decides whether captions read as professional or amateur. Every word gets its own timestamp rather than a block appearing for a whole sentence, which is what lets the caption highlight track the voice instead of trailing it. Filler words and dead air are cut out of the clip first, so the subtitle never displays the ums that were removed from the audio.

  3. 3

    Choose a style and export it burned in

    Pick from eleven caption looks, or pin one of them as your default in the Brand Kit alongside your brand colours and your logo, so the choice stops coming up on every export. The text is rendered into the final vertical frame at export, which means the captions are composed for the crop you are actually posting rather than scaled down from a landscape master and ending up half off screen.

What the subtitle burner does

Everything here exists because of a specific way subtitles go wrong.

Part of the picture, permanently

A hardcoded subtitle is made of the same pixels as the rest of the frame. Re-encode it, download it, hand it to a client, upload it to a platform that has never heard of your subtitle file — the words are still there, in position, in sync. That permanence is the entire reason people search for this rather than exporting an SRT.

Word-level timing rather than blocks

Traditional subtitles show a full line for two or three seconds. On short vertical video that reads as sluggish, because the viewer finishes the line and then waits. Highlighting one word at a time keeps the caption moving at the speed of speech, which is why the format took over short-form in the first place.

Kept clear of the interface

TikTok stacks a username, a description and a column of buttons over the lower right of every video. Subtitles dropped at the classic bottom-centre position land underneath all of it. Captions here sit inside the region that stays visible across vertical platforms, so you are not discovering the overlap after the post is live.

Eleven styles, one Brand Kit

Style is not decoration on a muted platform; it is legibility. Heavy weights with a stroke survive busy footage, lighter treatments suit clean backgrounds. Choose the style once, save it in the Brand Kit next to your colours and your logo, and every later export inherits all three, which is how a feed of clips starts looking like one channel.

The subtitle never shows what got cut

Filler removal and caption rendering are the same pipeline, not two tools stapled together. When a false start or a three-second silence is removed from the audio, the words come out of the caption track too. Bolt-on subtitle burners routinely leave that mismatch behind, and it is instantly visible.

Rendered after the reframe, not before

Captions burned into a 16:9 master and then cropped to 9:16 lose their edges. Here the vertical crop and the speaker framing happen first and the text is composited onto the final frame, so caption size and position are correct for the shape you are actually publishing.

Live clips arrive already captioned

Clips cut from a running YouTube, Twitch or Kick broadcast come back with subtitles already burned in. The whole value of a live clip is posting it while the moment is current, and stopping to caption it by hand is exactly what destroys that window.

Cuts that do not slice a sentence in half

A caption is only as good as the clip boundary underneath it. Clips are cut on sentence boundaries, so the first subtitle starts on the first word of a thought and the last one completes it. Half a word on screen at second zero is the fastest way to look automated.

Callable from code or from Claude

If burning captions is a step inside something larger — a content pipeline, a client delivery process — the same engine is reachable through a developer API and an MCP connector. Documentation is at the API docs.

Who needs captions burned in

Podcast clippers

Interview clips are dialogue with almost no visual event in them, so the words on screen are the whole hook. The podcast clip generator covers that source in full.

Paid social and ads

Ads autoplay muted in a feed. A hardcoded caption is the only version of your copy that still reaches someone who never turns the sound on.

Agencies delivering to clients

The deliverable is a finished MP4 the client posts themselves. Attaching a subtitle file and hoping it gets used correctly is how captions go missing.

Streamers and gaming channels

Busy backgrounds punish thin lettering. Stroked, high-contrast styles keep the text readable over explosions and UI clutter.

Educators and course creators

A definition or a step list is far easier to absorb as text on screen than as speech alone, especially when the viewer is on a train.

Faceless channels

With nobody on camera, captions carry the entire performance — more on that on faceless video clips.

Hardcoded, soft and platform captions are three different things

Soft subtitles live in a separate file — an SRT or a VTT — that a player reads alongside the video. The viewer can toggle them, translate them and restyle them, and none of that survives being separated from the file. Upload the MP4 without the sidecar and the subtitles simply do not exist.

Platform captions are generated by TikTok, Instagram or YouTube after you upload, and they live inside that platform only. They are convenient and they are trapped. Download your own video back from TikTok and the caption layer either comes with a watermark or does not come at all, which is the exact moment most people discover the problem.

Hardcoded captions are drawn into the video frames during export. There is no toggle and no file, because there is nothing to attach — the words are as much a part of the image as the person speaking. That is the version that travels.

Why burned-in survives re-upload and the alternatives do not

The common workflow is to cut a clip, post it to TikTok, then push the same file to Reels and Shorts a day later. If the captions came from TikTok, that second and third post go out silent and unreadable, or carry a competitor logo across the frame. Everyone learns this once, usually from a clip that was doing well.

The second failure is handoff. You send a file to a client, a colleague, an editor, or a scheduling tool. A sidecar SRT has to travel with it, be recognised at the other end, and be re-synced if anyone trims a second off the front. In practice at least one of those three steps fails, and the version that gets published has no captions at all.

The third is the repost economy. Clips get downloaded and reposted by other accounts, and a burned caption goes with them. Whether you welcome that or not, it decides whether your words survive the trip — an SRT never does.

When burning is not the right answer and a sidecar file is

Burning is wrong for plenty of jobs and pretending otherwise would be dishonest. On a forty-minute YouTube upload, viewers expect to switch subtitles off, and a permanent caption strip across a lecture is an irritation rather than a feature. Long-form belongs to soft subtitles.

Multiple languages need soft tracks too. A single burned track can only ever be one language, whereas a set of sidecar files lets a viewer choose. To be clear about our own scope: we transcribe the language that was spoken and burn that. There is no translation or dubbing in this product, so multi-language delivery is not something to plan around here.

Formal accessibility work is the third case. Where a standard requires the viewer to control caption presentation — size, contrast, on and off — burned text does not satisfy it, because the whole point is that the user cannot change it. Short-form vertical clips are a different context, and there burned captions are the accessible option in practice, since the platforms give viewers no reliable alternative.

How to check a burned caption before you post it

Look at it on a phone rather than the monitor you exported from. Text that feels comfortable at desktop size can be marginal at arm's length on a bright train, and that is the only viewing condition that decides anything.

Watch the first second with the sound off. If the opening word is missing, clipped or arriving late, fix the clip boundary instead of the caption — the boundary is nearly always the real cause, since text cannot start cleanly on a sentence that was already halfway through.

Scan for names and jargon. Speech models handle ordinary language well and mangle proper nouns, product names and industry acronyms far more often, and those tend to be the words the clip exists for. Two seconds of attention each time is a fair price for not publishing a misspelled surname.

Then check the corners. Put your thumb where the platform stacks its buttons and confirm nothing important is hiding underneath. Doing this once on your first export usually settles the question for everything you publish afterwards.

Caption mistakes that make a clip look cheap

Contrast first. Text with no stroke or shadow disappears the moment the footage behind it turns light, and it will do it on exactly the frame that mattered. A heavy weight with an outline is unglamorous and it works everywhere.

Then line length. Word-by-word styles keep only a few words on screen at once, which suits a phone held at arm's length. A full sentence shrunk to fit the width of a vertical frame is technically present and practically unreadable.

Then timing. A caption that lags the audio by even a fraction of a second is worse than no caption, because the viewer notices the seam instead of the content. Word-level timestamps rather than sentence blocks are what keeps the highlight sitting on the syllable being spoken.

What we have learned running this engine

Observations from operating the pipeline in production — not general advice.

The Brand Kit stores a caption style, not a font file

A brand kit in a design tool takes a typeface upload, and this one deliberately does not, which is worth knowing before you plan a house style around it. What gets saved is a named caption style out of a fixed set, your two brand colours as hex values, and a logo with a corner and a size. The reason is legibility rather than laziness: every preset has had its weight, stroke and highlight behaviour tuned to survive a 9:16 frame viewed on a phone, and an arbitrary uploaded face would arrive with none of that tuning and fall apart on the first bright background it met.

The caption safe area is an intersection of three platforms, not one

Bottom-centre is where a television subtitle belongs and it is close to the worst position available on a vertical clip. TikTok stacks a description and a column of buttons over the lower right, Reels puts its furniture somewhere slightly different, and Shorts differs again. Because the same file usually goes to all three, the usable band is the intersection of what every platform leaves alone rather than what any one of them permits, which is why captions here sit noticeably higher than broadcast convention would put them. Looking a little conservative beats losing a word behind a follow button.

Switching caption style costs a render, not a re-analysis

The expensive stage of captioning is not drawing the text, it is transcribing the audio and pinning a timestamp to every individual word. That result is cached against the project, which is why choosing a different style and rendering again never sends the audio back through the speech model. In practice that should change how you pick one: instead of judging from a preview sitting on a clean background, export the same clip twice against your own worst frame — a blown-out window, a busy game HUD — and compare those, because the only footage the decision applies to is yours.

Compared with the other ways to do this

vs. exporting an SRT and uploading it to each platform

That path means a transcription tool, a correction pass, an export, and then a separate upload on every destination that accepts sidecar files — several of which do not. It is more steps for a weaker result on short-form, where nobody is toggling subtitles anyway.

vs. TikTok or CapCut auto captions

They are free and they are quick, and they tie you to one editor. Text added inside a platform app tends not to survive being exported for use elsewhere, and the auto-generated timing is usually per-line rather than per-word. Fine for one post, painful as a workflow.

vs. captioning by hand in Premiere or Final Cut

A professional editor gives you total control over every word, and it costs a real chunk of an afternoon per finished minute once you include correction and styling. That trade makes sense for a flagship piece. It does not make sense fifteen times a week.

vs. a standalone subtitle burner

If you already have one finished thirty-second file and you only want text stamped onto it, a dedicated burner is the simpler tool and you should use it. This is a clipping engine: it takes long video, finds the moments, cuts them, reframes them and captions them. The captions are the last step of a bigger job — see what comes out before deciding which shape of tool you need.

Frequently asked questions

What does burning subtitles into a video actually mean?
It means the text is drawn permanently into the video frames during export, becoming part of the image rather than data attached to it. Once burned, there is no way for a player or a platform to hide the words, and equally no way for them to get lost. The industry term for this is hardcoded, as opposed to soft subtitles which sit in a separate file.
Do burned-in captions survive re-uploading to another platform?
Yes, and that is the main reason to use them. Because the words are pixels, they behave exactly like the rest of the picture through every download, re-encode and re-upload. Take the same file from TikTok to Reels to Shorts and the captions are identical in all three.
Is there any way for a viewer to hide burned captions?
No, and that is the deliberate trade you are making. If being able to hide the text matters for your use case, you want a sidecar subtitle file instead of burning. On short vertical video almost nobody wants them off, which is why the format standardised on burned-in.
Can I get an SRT file as well?
The output here is finished video with the words already rendered into it, because the product is built to hand you something postable rather than a subtitle asset to manage. If your workflow specifically needs a sidecar file for a long-form upload, a dedicated transcription tool is the better fit for that part of the job.
How reliable is the transcription on real audio?
Good on clear speech with reasonable audio, and it degrades the way every speech model does — heavy accents, crosstalk, loud music and low-bitrate recordings all cost accuracy. Rather than quote a number that would not describe your material anyway, run one video you know through the free demo and read the captions against what you remember being said.
Which languages can it caption?
Transcription covers the major languages and English is the strongest of them. The captions are always in the language that was spoken, because there is no translation step in this product at all.
Can it translate the subtitles into another language?
No. We do not do subtitle translation or dubbing, and we would rather say so plainly than let you find out after subscribing. What you get is an accurate rendering of the audio as it was recorded.
Can I use my own font for the burned-in subtitles?
Not a font file of your own, and it is worth saying that plainly rather than letting you subscribe expecting it. What you choose between is eleven caption styles, each one a fixed treatment of weight, stroke and highlight behaviour. Your brand colours and your logo do carry across every export through the Brand Kit, and switching style and re-rendering takes seconds when the first pick does not sit well over the footage.
Will the captions cover the TikTok interface?
They are positioned to stay clear of the areas platforms reserve for their own overlays — the username, description and button column that sit over the lower portion of a vertical video. This is one of the most common self-inflicted problems with burned captions and it is worth checking on your first export.
Are captions word by word or whole sentences?
Word by word, with each word timed individually so the highlight lands on the syllable being spoken. Block subtitles are the right choice for film and television and the wrong one for a thirty-second vertical clip, where the caption needs to move as fast as the speech.
Do I have to re-sync captions if I trim the clip?
No. Trimming and re-rendering regenerates the burned text against the new boundaries, so the drift you get from editing a video after attaching a subtitle file never happens. That drift is one of the main reasons hardcoding a caption at the end of the process is safer than at the start.
Can I burn subtitles into a video I have already edited?
You can upload your own file rather than a URL, so an edited master works. Bear in mind what the tool is: it takes a long source, picks the moments and returns short clips. If you want a single already-finished short video captioned untouched, this is not shaped for that job.
Do the captioned clips have a watermark?
Clips from the free demo carry a small one. Paid exports have none, and if you want a badge on your video it should be your own logo placed through the Brand Kit.
What length of source video can I send?
The ceiling is two hours per upload on a paid plan, and thirty minutes on the free demo. Anything beyond that has to be split first, or you can connect a live channel and have clips cut during the broadcast instead.
Does it caption clips taken from a live stream?
Yes, clips cut from a running YouTube, Twitch or Kick broadcast come back with subtitles already burned in. Captioning by hand is exactly the delay that makes a live clip stale, so it happens as part of the same pass.
Will the captions include filler words?
No, because filler and dead air are removed from the clip before the text is rendered. A caption that faithfully reproduces every um is technically accurate and reads terribly.
What happens when two people talk over each other?
Overlapping speech is genuinely hard for any transcription model and crosstalk is where you will see the most errors. On interview footage it is worth reviewing captions on any clip where people interrupt each other, which is usually also the clip you most want to post.
What if there is no speech in the video?
Then there is nothing to caption. Music-only footage, ambient gameplay with no commentary and silent screen recordings come back without text, since captions are transcribed from audio rather than invented.
Can I fix a word the model heard wrong?
You can change caption style, adjust the clip boundaries, edit the title and re-render. Where a transcription error lands on a critical word, the practical move is usually to shift the clip or pick a different moment, since the caption is generated from the speech rather than typed.
Does burning captions reduce video quality?
Rendering text into frames requires re-encoding the video once, which is unavoidable for any hardcoded caption in any tool. The export is produced at a resolution appropriate for vertical platforms, and the platforms themselves re-compress everything on upload regardless of what you send them.
Can I burn subtitles onto 16:9 or square video?
Yes. Exports are available as 9:16, 1:1 and 16:9, and the captions are composed for whichever frame you choose rather than positioned once and cropped afterwards.
Is this good for accessibility?
For short vertical clips, burned captions are the most reliable way to make speech readable, since platform caption support is inconsistent and viewers rarely enable it. For long-form video where a viewer should control caption display, a soft subtitle track is the correct format and burning is not a substitute.
Can I do this from my own code?
Yes, through the developer API, and there is an MCP connector if you would rather drive it conversationally through Claude. Both run the same pipeline described on this page — start with the developer docs.
What does burning captions cost?
Nothing to run the demo, which takes sources under thirty minutes and returns watermarked clips. Full access begins with a three-day trial billed at one dollar and continues at twenty-nine dollars a month if you keep it. You are emailed before that first monthly charge and cancelling is one click.
How long does it take to process?
A few minutes for a typical video, longer for a two-hour source or a busy queue. Nothing needs you present while it runs, so submitting and coming back later is the normal way to use it.
Can I do this from a phone?
Submitting a video, reviewing captions and downloading finished clips all work in a mobile browser. Checking caption legibility on the device people will actually watch on is a habit worth having.
What happens to my video after captioning?
It is processed to produce your clips and we publish nothing anywhere. The finished files are yours to use. Details are set out in the privacy policy.
How do I stop it renewing?
One click in your account stops it, and an email goes out before the trial converts so the charge is never a surprise. You keep everything you paid for until that billing period runs out.

See your words in the frame

Run one video through and check the captions on a phone. Legibility and timing are the two things you can only judge by looking.

⚡ Get A.I Clips — $1 trial
3-day trial · just $1 · cancel anytime