HomeAI Video ToolsAI Reframe Tool
AI Reframe Tool

Reframing that follows the subject, not the centre of the frame

A landscape frame holds far more width than a vertical one can show, so reframing is a series of decisions about what to throw away. This engine makes those decisions from what is happening on screen: who is speaking, how many people are in shot, whether a screen is being shared, and whether the footage should be cropped at all.

🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
  • Tracks the active speaker
  • Seven layouts, chosen from the footage
  • Outputs 9:16, 1:1 and 16:9
  • Override the layout and re-render
  • Same engine behind the API
One long video in, a week of posts out

Paste a link. Get the best moments.

Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.

https://youtu.be/z_bX3runikk
Source video: Matthew McConaughey's Concerns Over AI Source video 15:27
Matthew McConaughey's Concerns Over AI
9 clips found  ↓  top 5
Matthew McConaughey speaking, captioned vertical clip
Score 92
9:16
Joe Rogan gesturing, captioned vertical clip
Score 90
9:16
Matthew McConaughey listening, captioned vertical clip
Score 89
9:16
Joe Rogan mid-sentence, captioned vertical clip
Score 89
9:16
Matthew McConaughey talking to camera, captioned vertical clip
Score 88
9:16

How the reframe pass runs

Three stages, and the middle one is where the interesting work happens.

  1. 1

    Give it footage

    Paste a link or upload a file. Landscape 16:9 is the usual input, but square, ultrawide and already-vertical sources are all accepted. Paid uploads run to two hours; the demo covers anything under thirty minutes, which is plenty to inspect the framing on your own material.

  2. 2

    It reads the shot before it crops it

    The pass detects faces, works out who is speaking, notices when the shot changes, and identifies footage that is not a talking head at all — a shared screen, gameplay, a wide static camera. Those observations pick the layout. A single presenter and a four-person panel do not get the same treatment, because a crop that works for one destroys the other.

  3. 3

    Review, override, re-render

    The output is a reframed clip at your target ratio with the subject held in frame. If you disagree with a layout choice — you wanted the whole stage in shot, not a crop into the speaker — switch it and render again. The automatic choice is a default, not a verdict.

What the engine does on each clip

Every item here exists because a specific kind of footage broke without it.

Active-speaker tracking

The crop window follows whoever is talking rather than the geometric middle of the picture. In a conversation the framing moves with the exchange, which is the difference between a vertical clip you can watch and one where the person answering is off screen.

Movement that settles instead of chasing

A tracker that reacts to every small head movement produces a frame that will not sit still, and constant micro-drift is more tiring to watch than a slightly imperfect crop. The window holds a position and moves when the subject genuinely moves, so the result reads like a camera operator rather than a nervous algorithm.

Shot changes handled as shot changes

When the source cuts to a different angle, the framing needs to jump rather than glide, because a smooth pan across a hard cut looks like a mistake. Cuts in the source are detected and the crop is re-established on the new shot instead of sliding into it.

Seven layouts, not one crop mode

A tight fill, a whole-frame fit, a two-person split, three-up and four-up grids, a screenshare mode and a gameplay inset. Each is a different answer to the same question about what deserves the vertical frame, and having seven means a panel discussion and a solo monologue are not forced through the same template.

The layout is inferred from the content

Two faces trading lines suggests split. A game window with a webcam corner suggests picture-in-picture. A slide deck suggests screenshare. You are not asked to classify your footage before uploading it, because the footage already says what it is.

Override anything and render again

Automatic decisions are right most of the time and wrong occasionally, so every one of them is editable. Change the layout, nudge the crop, and produce a new file. Nothing about the first pass is final.

Three output ratios from one pass

9:16 for Reels, Shorts and TikTok, 1:1 for square feed placements, and 16:9 if you need a landscape version of the same cut. Each ratio gets its own framing decision rather than one crop being stretched to fit the others.

Applied across a whole batch

Reframing one clip by hand is tedious but survivable. Reframing the fifteen clips cut out of a long recording is where manual work stops being viable, and the engine does the whole batch in the same pass that produced them.

Runs on live footage too

The same framing logic is applied to clips cut from a running YouTube, Twitch or Kick broadcast, which means a live moment arrives already vertical and already framed rather than as a landscape grab you have to fix later.

Available programmatically

If reframing is a step in a pipeline you already run, the engine is reachable through the ClipSpeedAI API and through an MCP connector, so Claude can reframe footage for you inside a conversation.

Where reframing decides whether the clip works

Interview and panel footage

The single hardest case, and the one that exposes bad reframing immediately. Two or more people spread across a wide frame cannot survive a fixed crop, which is why split and grid layouts exist at all.

Gaming and streaming

Gameplay plus a webcam is two sources competing for one vertical frame. Picture-in-picture resolves it by giving the game the space and the face the corner, and it is why gaming clips reframe more cleanly than people expect.

Webinars and screen shares

Cropping into a shared slide makes it unreadable. The screenshare layout keeps the shared content whole and places the presenter around it, which is the only version of this that stays legible on a phone.

Editors working at volume

Keyframing a crop by hand is a real skill and a slow one. Automating the routine ninety percent leaves the human time for the shots where framing is a creative decision rather than a mechanical one.

Marketing teams repurposing recordings

Conference talks and product demos were filmed landscape and need to appear vertically. Reframing is the step that makes that possible without a reshoot.

Video podcasters

Multi-camera podcast footage cuts between wide two-shots and single-camera close-ups, so the reframe has to change behaviour shot by shot. The podcast clip generator pairs this engine with episode-length moment detection.

Why a fixed centre crop keeps failing

The naive way to make a landscape video vertical is to take a column out of the middle and discard the rest. It is fast, it needs no analysis, and it works on exactly one kind of footage: a single person, centred, who does not move. Almost nothing is filmed that way.

Consider what the middle column actually contains in ordinary footage. In a two-person interview it is the space between the chairs, so you get a vertical video of a background while two people talk off screen. In a stage talk it is whatever the presenter happened to walk away from, since presenters pace. In a webcam recording the person is often framed slightly left or right by the camera itself, which puts them at the edge of the crop or outside it. In gameplay the middle is the game, and the streamer whose reaction is the entire point sits in a corner that has just been deleted.

The failure is not subtle when it happens, but it is easy to miss when you are exporting in bulk and checking thumbnails rather than watching. That is the argument for making the crop decision from the content: not that tracking is clever, but that the alternative is systematically wrong on the footage most people actually have.

The seven layouts and the footage each one is for

Fill crops into the frame so the subject occupies the full vertical rectangle, edge to edge. This is the default for a single speaker and the one that looks most like native vertical footage. It is also the most destructive, because everything outside the crop window is gone.

Fit does the opposite. The entire landscape frame is scaled down and kept whole, with the space above and below occupied rather than left black. Use it when the edges of the frame carry information — a whiteboard, an on-screen chart, a wide shot where the composition is the point. The subject ends up smaller, and that is the trade you are accepting in exchange for losing nothing.

Split stacks two panels vertically, one per speaker, each framed on its own face. This is the layout that rescues interviews and debates. Both participants stay visible continuously, so an interjection lands with the person who made it rather than arriving from off screen.

Three-up and four-up extend the same idea to panels and group calls, dividing the vertical frame into three bands or a two-by-two grid. Faces get small, so these work best on footage where people are already framed as head-and-shoulders rather than sitting across a wide room.

Screenshare treats the shared window as the priority. The shared content keeps its aspect ratio and enough of the frame to remain readable, with the presenter placed around it. Cropping into a slide is the single fastest way to make a webinar clip useless, and this layout exists specifically to avoid doing that.

Gameplay picture-in-picture gives the game the vertical frame and insets the webcam over it. The reaction and the thing being reacted to stay in the same shot, which is the whole grammar of a gaming clip.

A working routine: how to check a reframed batch in five minutes

Reviewing framing properly is a skill, and most people either skip it entirely or watch every clip end to end. Neither is necessary. Scrub each clip at three points — the first second, somewhere in the middle, and the last second — and you will catch nearly every framing error that matters, because framing failures are persistent rather than momentary.

Watch on a phone if you can, or at least at phone size in a browser window. A crop that looks slightly tight on a twenty-seven-inch display can be perfect on a handset, and a face that reads fine at desktop scale can be unrecognisably small in the four-up grid once it is on a phone. Judging vertical video at desktop scale is the most common reason people reject framing that was actually correct.

Two habits are worth building. Check the clips with more than one person in them first, since those are where automatic decisions are hardest and where an override is most likely to be needed. And when you do override, override toward fit rather than toward a different crop — if a tracked crop was wrong, it is usually because something outside the crop mattered, and fit is the layout that brings it back.

Finally, feed the tool your best source. A reframe made from a 4K master and a reframe made from a re-compressed download are the same decision executed on different amounts of information, and only one of them holds up on a close-up.

Troubleshooting a reframe you cannot fix by switching layout

Most disappointing reframes are solved by picking a different layout, and if you have not tried that yet, try it before anything else. The cases worth writing down are the ones where every layout is wrong, because those almost always mean the problem sits in the source rather than in the crop decision.

A subject who slides out of frame on a clip that is otherwise tracked well is usually a lighting problem. Detection needs the face to be separable from what is behind it, so a presenter lit from the back, a dim room, or a hard shadow cutting across one cheek gives the tracker only intermittent signal, and between hits the window eases back toward the middle. Fit is the dependable escape, because it asks nothing of the tracker at all.

Panes holding the wrong people in split or in one of the grids point at a face count that changed partway through. Someone leaning into shot for a few seconds, or a fifth participant arriving in a four-up, means the arrangement was decided against a different population than the one on screen for most of the clip. Trimming so the cast stays stable across the whole range fixes that far more reliably than rendering the same range again.

Framing that reads correctly in one preview and badly in another is a ratio mismatch rather than a fault. Every output ratio is framed independently, so a vertical preview tells you nothing dependable about the square version of the same moment. Check the ratio you are actually about to post.

If a clip is soft rather than mis-framed, no layout will rescue it, because softness is a resolution problem: the crop is spreading fewer source pixels across the same output rectangle. Re-run that one from the highest-resolution master you still have instead of adjusting the framing again.

When reframing is the wrong answer

Some footage should not go vertical, and a tool that pretends otherwise wastes your time. Dense screen recordings are the clearest case: a desktop capture with small type does not become readable by being scaled, and no crop can fix text that was already at the limit of legibility on a laptop. If your source is mostly code, spreadsheets or detailed slides, the honest path is to rebuild those visuals for a vertical frame rather than convert them.

Wide establishing shots and anything where the composition is the content are the second case. A landscape drone shot of a valley reframed to 9:16 is a shot of one hillside. Fit will preserve it at reduced size, which is sometimes the right call, but you are no longer showing what the shot was for.

And if there is no face and no obvious subject at all — product footage on a table, abstract B-roll, an animated explainer — tracking has nothing to lock onto and will behave like a centre crop, because that is the only sensible default left. Reframing helps when there is a subject worth following. It does not invent one.

What building the override path taught us about the engine

Observations from operating the pipeline in production — not general advice.

The uncropped original is stored before the first crop is made

Overriding a layout means rendering again from the full-width picture, so the original still has to exist at the moment you change your mind. It is copied into object storage under the project id before the pipeline begins cropping rather than afterwards, because by the time a run finishes the working copy has already been discarded. That ordering is the only reason the adjust-framing controls can hand you a square or landscape version of a clip that was first rendered vertical without asking you to upload the file a second time. Where that copy is absent on an older project, the editor quietly falls back to a face-biased preview rather than refusing to open.

The working copy is deleted mid-run, and the timing is deliberate

A source video is routinely somewhere between half a gigabyte and two gigabytes, and every minute it occupies disk is a minute another queued job cannot use that space. The pipeline therefore removes it the moment cutting and watermarking are done, before the phase that uploads finished clips even starts, rather than waiting for the job to terminate. That is because disk, not processing time, turned out to be the real ceiling on how many videos can run at once, which was not the constraint we expected to hit first. The extracted audio track is dropped at the same instant for the same reason.

Every re-render is counted, and the count is the useful part

When you override a layout and render again, the clip stores both the time of that render and a running total of how many times it has happened. One override is unremarkable and tells us nothing, since the automatic pick is a default rather than a verdict. A clip rendered four or five times is a completely different signal: it means none of the seven layouts suited that footage and you were fighting the tool instead of correcting it. Those clips are the ones worth reading, because they mark the footage types where our layout inference is actually wrong rather than merely arguable.

Compared with other ways to reframe

vs. keyframing the crop by hand

A skilled editor keyframing a position track will beat any automatic pass on a single important shot. The economics change immediately at volume: fifteen clips from one recording is a couple of hours of position keyframes, and the quality gap on routine talking-head footage is small enough that most of those hours buy very little.

vs. auto-reframe inside a traditional editor

Editors like Premiere ship reframing that tracks motion and does a competent job. What they do not do is decide between layouts — there is no notion that this particular clip is a two-person exchange that should be split, or a screen share that must not be cropped. You get one reframing strategy applied to everything.

vs. padding with blurred bars

Duplicating the frame, blurring it and using it as a background is the standard fallback, and it is not always wrong: on footage that must stay whole it is more honest than a bad crop. But it shrinks the subject in a format that is watched at arm's length on a phone, and a wall of blur reads as a repost rather than something made for the feed.

vs. reframing inside other AI clip tools

Most competent clipping tools include face-aware cropping, so the differences are in the corners: how many layout options exist, whether cuts in the source are handled distinctly from motion, and whether you can overrule the machine and re-render. Those corners are where footage that is not a single centred talking head lives.

Frequently asked questions

What does an AI reframe tool actually do?
It converts video from one aspect ratio to another while deciding which part of the original frame to keep. Instead of taking a fixed column out of the middle, it identifies the subject — usually the person speaking — and positions the crop window around them, changing the framing as the shot and the speaker change.
Does it track the person who is talking, or just any face?
It follows the active speaker. On footage with several people in shot that distinction matters constantly, because framing whoever happens to be nearest the middle produces a clip where the answer comes from off screen while a listener nods at the camera.
What aspect ratios can I reframe to?
9:16 vertical for Reels, Shorts and TikTok; 1:1 square for feed placements; and 16:9 if you need a landscape version. Each is framed independently, since the right crop for a square frame is not simply a wider version of the vertical one.
What are the seven layouts?
They are named fill, fit, split, three-up, four-up, screenshare and gameplay inset. Fill crops in tight, fit keeps the whole frame at reduced size, split and the grid options stack multiple speakers, screenshare protects shared content from being cropped, and the inset holds a webcam over gameplay.
Can I choose which layout a clip uses?
Yes. The engine picks one from what it sees in the footage, and you can override it and re-render if you disagree. People most often override toward fit, usually because something at the edge of the original frame matters more than the automatic pass assumed.
Why not just crop to the centre?
Because the centre of a landscape frame is frequently not where the subject is. In a two-person interview it is the gap between them. On a stage it is wherever the presenter is not standing at that moment. Centre cropping only survives on single, centred, stationary subjects.
Does the framing jitter or drift while watching?
The crop window is designed to hold still and move deliberately rather than react to every small movement, because constant micro-adjustment is more distracting than a slightly loose frame. If you spot drift on a particular clip, the fit layout removes camera motion from the equation entirely.
What happens when the source cuts between camera angles?
The cut is detected and the framing is re-established on the new shot instead of gliding across it. Sliding smoothly through a hard cut is one of the most obvious tells of automated reframing, so cuts and motion are treated as different events.
Will it work if there are no people in the video?
It will run, but with nothing to track it behaves close to a centred crop, which may or may not be what that footage needs. For product shots, animation or abstract B-roll, consider the fit layout so the composition survives intact.
Can it reframe a whole two-hour video into one long vertical file?
No, and it is better to say so plainly. This produces clip-length exports, not a full-runtime vertical version of your source. You can widen a clip's in and out points, but if your goal is an entire lecture as one vertical file, a conventional editor is the right tool.
Does reframing lose quality?
Cropping into a frame means fewer original pixels covering the same output size, so a 1080p landscape source has less real detail to work with than a 4K one does. It is still perfectly acceptable on a phone. If you have a 4K master, use it — the difference is visible on close-ups.
Can it reframe vertical footage to landscape?
Yes, 16:9 is an available output. Going that direction has the opposite problem: you have height to spare and not enough width, so expect the fit treatment more often, since there is simply no extra picture at the sides to reveal.
Does reframing work on screen recordings?
Screen content uses the screenshare layout, which protects the shared window from being cropped. That said, a dense desktop capture with small text is hard to read on a phone no matter how it is framed, and the realistic answer for that footage is to redesign the visuals rather than convert them.
Does it work on animation or motion graphics?
It runs, but the tracking has no face to follow, so you are effectively choosing between a centred crop and fit. Animation with a clear central character reframes fine; a busy motion-graphics sequence usually does not.
How does this handle a four-person podcast?
Typically with the four-up grid, which puts each participant in their own cell of a two-by-two arrangement. It works best when everyone is already framed head-and-shoulders. Four people across one wide room shot gives every cell too much empty space, and in that case split on the two active speakers reads better.
Are captions positioned around the reframe?
Captions are burned in after the frame is built, so they sit within the vertical picture rather than being cropped by it. There are eleven caption styles, and switching style is the quickest fix if the words land somewhere awkward on a particular clip.
Is reframing available on its own, without the clipping?
Reframing is part of the clip pipeline rather than a standalone converter. Every clip the tool produces goes through it. If you want the framing behaviour applied inside your own workflow, the API is the route.
Is there a free way to test the framing?
Yes. Run a video under thirty minutes through the free demo and inspect the results — framing is a visual judgement and no description substitutes for looking at your own footage. Demo clips carry a watermark; paid exports do not.
What does it cost after the demo?
A three-day trial is $1, and Pro is $29 per month afterwards. One click cancels it, and we send an email before the trial converts so nothing is charged quietly.
Does it handle ultrawide or unusual source ratios?
Yes, though the wider the source the more aggressive the discard has to be. An ultrawide frame going to 9:16 loses an enormous proportion of its width, so tracking matters more there than on ordinary 16:9 footage, not less.
Can I apply my own branding to reframed clips?
Yes. The Brand Kit holds your logo, colours and fonts and applies them at render, which keeps a batch of clips visually consistent even when they came out of different layouts.
Will it work on footage where someone walks around the stage?
That is a good case for tracking and a terrible one for a fixed crop. The window follows the presenter as they move. If they cross the frame very rapidly, the fit layout is the safer choice, because a fast horizontal traverse can look restless even when tracked accurately.
Does the reframe change the audio?
The framing pass does not touch audio, but the clip pipeline does remove filler words and dead air, which shortens the clip slightly. If you need audio untouched, that trimming is the part to be aware of.
Can I use this through an API?
Yes. The reframe engine sits behind the same API as the rest of the pipeline, and there is an MCP connector for driving it from Claude. Both are documented in the developer docs.
How does this differ from a vertical video converter?
A converter changes the container and the dimensions. The reframing decision — which part of the picture to keep — is the part that determines whether the result is watchable. Our vertical video converter page covers the format side of the same job.
Is my footage kept after processing?
Yes, and on this tool that is deliberate rather than incidental. The working copy on the machine that did the cutting is discarded as soon as rendering finishes, but a copy of your uncropped original is kept alongside the project, because overriding a layout and re-rendering needs the full-width picture back — without it, every framing change would mean uploading the file a second time. That stored original, the clips and the transcript all persist until you remove them: deleting a single clip drops its file and row straight away, deleting a project cascades through its clips, transcripts and thumbnails, and deleting the account purges the associated objects from storage within thirty days. None of it is published or shared by us at any stage. The privacy policy breaks the same thing down category by category.
What if I do not like any of the layouts on a clip?
Export at 16:9 and treat the clip as a normal landscape cut. Not every moment belongs in a vertical frame, and forcing one that does not is worse than posting it wide or skipping it.
When is it still worth cropping the shot by hand?
When one clip genuinely matters more than the rest — a launch video, a paid placement, the opening of a series. A person setting crop positions frame by frame will read intent that no automatic pass can infer, such as holding on a reaction the tracker would have moved away from. The moment you are working through a batch of fifteen, that advantage stops paying for the hours it costs.
What are the alternatives if vertical framing is not working at all?
Three, roughly. Post the moment wide and let it live as a landscape cut. Pad it with blurred bars, which keeps the frame whole and is honest enough for footage that cannot be cropped, though it reads as a repost. Or shoot the next one with the vertical crop already in mind, keeping the subject and any on-screen text inside the middle third of the frame, which removes the problem at the source instead of solving it afterwards.
What source resolution should I send, and what can I export?
Send the highest-resolution master you have. A vertical crop keeps only a tall column of the original, so a 4K source leaves real detail in that column where a re-compressed 1080p download leaves very little, and the gap shows on close-ups rather than on wide shots. Exports are available at 9:16, 1:1 and 16:9, each framed as its own decision rather than derived from the others.

Watch it reframe footage you have already tried to crop

Pick the recording you gave up on — the wide two-shot, the presenter who paced, the panel of four. Those are the clips that tell you whether the tracking is real.

⚡ Reframe My Video — $1 trial
3-day trial · just $1 · cancel anytime