Two-person footage breaks the two things automated clipping usually gets wrong. A centre crop frames the empty air between two chairs, and an answer lifted away from the question it responds to stops making sense. This cuts around the exchange, follows whichever person is speaking, and hands back vertical clips where the setup and the payoff are both still there.
🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
⏱ That video's a bit long for the demo
Keeps the question with the answer
Split-screen when both faces matter
Tracks whoever is talking
Word-by-word captions burned in
Every clip graded 0-100
One long video in, a week of posts out
Paste a link. Get the best moments.
Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.
The interesting work happens in step two, and none of it is yours.
1
Point it at the interview
Paste the link if the interview is already on YouTube, or upload the recording straight from your machine. Remote sessions exported from a call recorder work, and so does a two-camera edit that has already been cut together. Two hours is the ceiling on a single paid upload; the free demo takes anything shorter than thirty minutes.
2
It maps the conversation, not just the audio
The transcript is read as a dialogue rather than a wall of speech. It tracks who is asking and who is answering, which questions produced a real answer instead of a deflection, and where an exchange resolves. That structure is what lets a clip start on a question and finish on the beat after the point lands, rather than fifteen seconds into an answer nobody has context for.
3
Review the exchanges it picked
Each clip comes back framed, captioned and scored 0-100. Watch the top few, extend a boundary if you want the interviewer's reaction included, and publish. Send a copy to your guest as well — a good interview clip is the one piece of content a guest will reliably reshare, and their audience is the reason you booked them.
What it does that a general clipper does not
These are specific to two people talking. On a monologue most of them are irrelevant, which is precisely why generic tools do not bother with them.
🪞
Split-screen so nobody disappears
When the moment depends on both people — a challenge and a rebuttal, a claim and a raised eyebrow — the split layout stacks them so neither is cropped out. The alternative that other tools produce is a clip where you hear a question from someone who is not on screen, which reads as an editing mistake even to a viewer who cannot say why.
👀
Tracking that switches with the speaker
A fixed crop on a two-shot spends half the clip on whichever chair was closer to the middle. Speaker tracking moves the frame to whoever currently has the floor, so a fast back-and-forth stays legible in vertical instead of turning into a shot of the space between two people.
❓
Questions kept attached to their answers
An answer is often meaningless without the ten words that prompted it. Boundaries are chosen with the exchange in mind, so a clip can open on the question and close after the answer completes. That single behaviour is what separates an interview clip that works from a fragment that needs a caption to explain itself.
✂️
Cuts that land on sentence edges
Nobody in conversation speaks in tidy blocks, and the temptation for automated tools is to slice on silence. Silence in an interview is usually a person thinking, not a person finishing. Cutting on sentence boundaries respects the difference, so a clip does not begin halfway through a considered thought.
📊
A 0-100 score across the whole conversation
A long interview contains many decent exchanges and two or three genuinely strong ones, and after two hours in the room you are the worst-placed person to tell them apart. The grade sorts them for you on how the clip opens, paces and peaks, which is a more useful second opinion than asking a colleague to watch the whole thing.
🔤
Captions that carry accents and cross-talk
Interviews are the format most likely to be watched on mute in a feed, and the format where a strong regional accent or an unfamiliar name most needs to be on screen. Word-by-word captions in eleven styles are rendered into the video, timed to the speech, and they stay correct if you trim the clip afterwards.
🧽
Filler, crosstalk stumbles and dead air removed
Conversational footage carries more of this than any other format — the half-started question, the overlapping agreement, the pause while somebody remembers a name. Stripping it automatically usually buys back several seconds on a minute-long clip and makes both parties sound sharper than the raw recording did.
🎁
Guest-ready clips in every ratio
Export vertical for Shorts, Reels and TikTok, square for feeds, or landscape if the guest wants something for a newsletter. Handing a guest three finished clips within a day of recording is the cheapest way to make them promote the episode, and it costs you nothing extra once the clips exist.
🔴
Live interviews clipped while they air
If the conversation is being streamed to YouTube, Twitch or Kick, clips are produced during the broadcast rather than after. A remark that will be quoted everywhere tomorrow can be posted while people are still watching, which is not something a post-production workflow can offer.
Who is cutting interviews all week
🎙 Interview podcasters
The whole format is one guest and one host for ninety minutes. Clipping it is the growth channel, and doing it manually is why most shows stop after four episodes. See the podcast clip generator.
📺 YouTube interview channels
Long-form conversations get discovered through short excerpts of the sharpest exchange. Framing decides whether that excerpt is watchable on a phone.
🗞 Journalists and news teams
A recorded sit-down usually contains one quotable admission, and the value of getting it out fast decays by the hour rather than by the day.
🎬 Documentary and street interviewers
Vox pops produce dozens of short exchanges in a single shoot. Sorting them by score is faster than reviewing four hours of footage looking for the three good responses.
💼 Recruiters and employer brand teams
Employee interviews and culture pieces work far better as sixty-second answers than as a four-minute corporate film nobody finishes.
🏢 Customer marketing
A customer answering why they switched is the most persuasive footage a company owns. The webinar clip generator covers the B2B side of the same idea.
The framing problem nobody warns you about
Shoot an interview properly and you get a wide two-shot: both people, some room around them, a comfortable amount of air. It looks great on a television and it is close to unusable on a phone. Converting it to vertical removes about two thirds of the width, and if that crop is fixed to the centre it will often frame the coffee table, the microphone stand, or a nicely lit patch of wall.
Even a crop that finds a face has a second problem. Interviews alternate, so a crop locked on the host misses every reaction from the guest and vice versa. The viewer hears a voice with no source, which is subtly disorienting in a way most people register as the clip feeling cheap.
The fix is not clever, it is just work: identify who is talking, move the frame accordingly, and when both people matter simultaneously, stack them rather than choosing. Doing that by hand on every clip is a tedious afternoon. Doing it automatically is the reason this page exists.
Where an interview clip should start
Take a strong answer out of an interview and you often find it starts with a pronoun. "That happened to us twice." Twice what, and to whom? The line was perfectly clear in the room, because the question was still hanging in the air, and it is nonsense fifteen seconds later in a feed.
There are two ways to solve it. You can write an on-screen caption explaining the setup, which is what most people do and which costs you the first two seconds of attention. Or the clip can simply include the question. The second is almost always better, because a question is short, it creates tension, and it gives the viewer a reason to stay for the answer.
This does not always work. Some interviewers ask forty-second questions with three sub-clauses, and no reasonable clip can carry that as an opening. In those cases the clip starts inside the answer at the first line that stands alone, and you may want to add your own context. Knowing which situation you are in is a ten-second judgement once the clip is in front of you, and impossible before it exists.
Where interview clipping goes wrong
Heavy crosstalk is the hardest case. Two people talking over each other produces a transcript with uncertain speaker boundaries, and uncertain boundaries produce cuts in slightly wrong places. Panel discussions where three people interject are worse than a calm two-hander by a wide margin.
The second failure is a guest who cannot be brief. Some people answer in four-minute paragraphs where the point arrives at the end and depends on everything before it. There is no short clip inside that answer, and no tool can invent one. If your interviews consistently run that way, the fix is upstream — asking sharper questions, or interrupting more.
Third, audio recorded on a single room mic with both people some distance away transcribes poorly, and a poor transcript damages captions and cut points together. A cheap lapel mic on each person improves clip quality more than any software choice you will make this year.
A recording-day routine that pays off in clips
Almost everything that determines clip quality is decided before the camera stops. Four habits are worth building, and none of them cost anything.
Ask short questions. A twelve-word question can open a clip; a ninety-word preamble with three sub-clauses cannot, and it forces every clip from that answer to begin cold. If you catch yourself explaining the question, stop and ask it again in one sentence — you can always cut the first attempt.
Let a beat of silence sit after an answer finishes. Interviewers instinctively fill the gap with agreement noises, and those noises land exactly where the clip wants to end. A held pause gives the clip a clean edge and often prompts the guest to add the better line.
Encourage the guest to restate the subject rather than lean on a pronoun. "The migration took us eight months" survives being lifted out of context; "that took us eight months" does not. Say so before you start recording and most people adapt within ten minutes.
Mic both people separately and frame slightly tighter than feels right for a landscape edit. A generous wide two-shot looks handsome on a monitor and leaves the vertical crop working with a fraction of the pixels. Everything downstream — the tracking, the captions, the boundaries — is downstream of these choices.
Doing it by hand on a two-shot, and what that actually costs
The part people price is the cutting, and the cutting is nearly free. Everything either side of it is where the hours go. Say you have an hour-long conversation and you want eight clips out of it: that is the better part of an hour scrubbing to find the exchanges, eight separate decisions about where each one opens, and then the reframing — which on two-person footage is not one crop but a sequence of them, one per speaker change, keyframed across the length of the clip. That last part is why editors quietly resent vertical interview work.
A human still wins on two things and it is worth naming them precisely. They know the guest, so they know which line will get quoted approvingly and which will get screenshotted uncharitably. And they hear tone. A sarcastic answer lifted out of its exchange reads as sincere in a feed, and no transcript carries the difference.
What a person loses is every clip after roughly the sixth. Reframing is mechanical, repetitive and identical each time, and attention on repetitive work degrades in a way that surfaces as inconsistent framing across a week of posts. Keeping the judgement and handing over the mechanics is not a compromise between the two approaches — it is the division of labour that matches what each side is genuinely good at.
The honest test is throughput. If you publish two interview clips a month, cut them yourself and enjoy the craft of it. If the plan calls for a clip every weekday from a weekly episode, that plan does not survive contact with manual reframing, and most shows that attempt it quietly stop by the second month.
The alternatives, and the one most interview shows overlook
Platform clip buttons are the free option and they beat nothing at all. YouTube has its own clip feature and Twitch and Kick both ship one, and if a moment is genuinely good your audience will mark it without being asked. What comes back is a landscape excerpt with no captions and no vertical framing, so treat it as a signal about which moments landed rather than as a finished post.
A transcript editor is the second route. Deleting a paragraph of text deletes the matching stretch of video, which suits anyone already reading the transcript to write show notes. It removes the scrubbing but not the choosing — every clip is still your decision — and it does nothing whatsoever about framing a two-shot for a phone.
The overlooked one is the guest. Anyone with an audience of their own usually has somebody who edits their content, and a raw file plus three suggested timestamps frequently comes back as three finished clips at no cost to you, cut by a person motivated to make their principal look good. Most podcasters never ask, which is odd given the guest's editor is already on a retainer and the guest wants the episode to travel as much as you do.
Then there is the null option: publish the full interview with chapter timestamps and do nothing else. It is a real strategy and it fails for a specific reason. Timestamps are navigation for someone who already opened the video; clips are discovery for someone who never would. They are not substitutes, which is why a show that only does the first tends to grow strictly within the audience it already has.
What two-person footage has taught us in production
Observations from operating the pipeline in production — not general advice.
Who said what is not stored anywhere, and that shapes the boundaries
There is no speaker-diarisation step producing neatly labelled turns — the words arrive as one continuous transcript. Turn-taking is inferred from the shape of the language instead: a question, an answer that resolves it, a name being addressed directly. That holds up on a calm two-hander and breaks on sustained overlap, where two voices collapse into a single run-on sentence and the cut lands somewhere inside it. It is the mechanical reason panels need more review than a sit-down, and why separate microphones improve the transcript before they improve the sound.
An already-edited two-camera interview is the easy case, not the hard one
People assume raw wide footage is the safe thing to hand over and that a finished multi-camera edit will confuse the framing. It runs the other way round. A cut between an A-cam and a B-cam gives the tracker one face per shot with a hard boundary between them, so each stretch gets its own anchor and nothing has to arbitrate between two equally valid subjects. The genuinely difficult input is a single locked-off wide two-shot held for ninety minutes, because every frame contains two candidates and the choice has to be remade continuously.
Extending the in-point is the one edit interview clips actually need
Across every format we see, the adjustment made most often on interview output is dragging the opening back a beat — not to repair a bad boundary, but to catch the interviewer's reaction or the scrap of setup that makes an answer land. Boundary logic optimises for a clean start on a complete thought, and a reaction is not a thought, which is why it sits just outside the cut by design. So a clip that starts correctly and still feels abrupt usually wants two more seconds at the front rather than a different moment.
Compared with the alternatives
vs. clipping it yourself in an editor
Manually, the cutting is the easy part. The cost is rewatching ninety minutes to find the six exchanges worth keeping, then reframing each one so both people stay visible. Do that weekly and it becomes the reason the show stops publishing clips. Reviewing a scored shortlist takes twenty minutes and produces more posts than the diligent version ever did.
vs. a generic AI clipping tool
Most of them are tuned for one person addressing a camera, and they will confidently centre-crop your two-shot. On a solo video you would not notice the difference. On an interview it is the whole ballgame, and the tell is a clip where a disembodied voice asks something while the wrong person sits silently on screen.
vs. hiring a clip editor
A good editor will beat this on the pieces that matter most, because they understand your show, your guest and what your audience finds funny. They will not beat it on cost per clip or on turnaround, and turnaround is what determines whether a clip goes out while the episode is still current. Use a person for the trailer and the launch clip, and automate the other fifteen.
vs. posting the full interview and hoping
A long interview is discovered by people who already follow one of the two people in it. Clips reach everyone else, and they are also what convinces a guest's audience that the full conversation is worth ninety minutes. Both belong in the plan; only one of them is a distribution channel. Try it on your last interview and compare what the score picked against what you would have chosen.
Frequently asked questions
How does it decide where an interview clip should start?
It reads the conversation as a dialogue and places the opening at the start of a complete thought, which for interviews is frequently the question rather than the answer. Where the question is too long to include, the clip opens on the first line of the answer that stands up without context. You can always drag the boundary yourself and re-render.
Will both people stay on screen?
When the moment needs both of them, yes — the split layout stacks the two speakers so nothing is cropped away. When only one person is talking through a stretch, the frame follows that person, which reads more naturally than a static split for a long uninterrupted answer.
What happens with a wide two-shot?
That is the case a fixed centre crop handles worst, since the middle of a wide two-shot is usually empty space. Face tracking finds the people rather than the geometric centre of the frame, and the layout is chosen based on whether one or both need to be visible at that moment.
Does it work on remote interviews recorded over a call?
Yes. Grid recordings from call software are common source material, and the layout logic handles a two-tile grid well. If your recorder produces separate speaker files, export the combined version — a single composed video is what the tool expects.
Can I use a YouTube link instead of uploading?
Yes, and for anything already published that is the quicker route since the file never has to leave your machine. Uploading is for interviews that have not been posted publicly, including raw footage that will never be released in full.
Can I test it on an interview before paying?
A free demo handles videos under thirty minutes, which is enough for a short interview or a segment of a longer one. Beyond that it is a one-dollar trial for three days and then twenty-nine dollars a month. Cancel in one click; a reminder email arrives before the trial converts.
How long can the interview be?
A paid upload can run up to a full two hours, which covers most long-form conversations. Thirty minutes is the limit on the free demo. For a marathon session beyond two hours, split the file or clip from the live stream while it is happening.
Is there a badge on the exported clip?
Only on the free demo ones. Anything exported on a paid plan is clean, which matters here because interview clips get reshared by the guest and a third-party badge on their feed looks odd. Your own logo can be applied through the Brand Kit instead.
How many clips come out of a ninety-minute interview?
Often in the range of twelve to twenty-five, though the spread is wide. A crisp interviewer who keeps answers short produces many more usable clips than a rambling conversation of the same length. The score is there to help you pick the six you will actually publish.
Does it handle three or more people?
It does, with three-up and four-up grid layouts, but honestly a two-person conversation is the case it handles best. Panels involve interjections and crosstalk, which makes speaker boundaries less certain and means the clips need a closer look before publishing.
What about interruptions and people talking over each other?
Brief overlap is normal and handled. Sustained crosstalk is genuinely difficult, because the transcript cannot cleanly separate who said what, and cut placement depends on that separation. If a section is chaotic, expect the clips from it to need manual adjustment.
Can it clip a live interview as it is being streamed?
Yes. Connect a YouTube, Twitch or Kick channel and clips are produced mid-broadcast, so a notable answer can be posted during the conversation rather than after the VOD processes. This is uncommon among clipping tools and it changes what is possible for a live show.
Are captions accurate with strong accents?
Generally good and not perfect. Accents, technical jargon and proper nouns are where transcription is most likely to slip, and an interview contains a lot of names. Check the captions on any clip where a person's name or a company name is central before you publish it.
Can I edit the clip after it is generated?
Yes. Move the in and out points, switch among the eleven caption styles, change the layout, rewrite the title, and re-render. Extending the opening by a couple of seconds to catch a reaction shot is the most common adjustment people make.
Which formats can I export for a guest to share?
Vertical 9:16, square 1:1 and landscape 16:9. Interview clips are worth exporting in more than one ratio when a guest will share them, since their audience may live on a different platform than yours.
What happens to all the ums and false starts?
They come out automatically. Conversation is dense with them, and their removal is more noticeable on interview footage than anywhere else — a guest who said "you know" nine times sounds considerably more articulate once they have not.
Can I send clips straight to social platforms?
Yes. Connect TikTok, Instagram Reels and YouTube Shorts to publish now or schedule ahead. Spacing a single interview's clips across two weeks tends to outperform posting all of them the day the episode drops.
Will it pick the moment I would have picked?
Frequently, though not always, and that is the right expectation to hold. It is very good at surfacing candidates and ranking them sensibly. It does not know that your guest made the same point better on a previous show, or that a joke will read badly out of context. Treat it as a well-briefed assistant, not an oracle.
Can I wire this into our production pipeline?
Yes, there is an API, along with an MCP connector so Claude can run the clipping conversationally. Teams producing several interviews a week usually automate submission rather than uploading by hand. The developer documentation has the specifics.
Can I add my show branding?
Yes. The Brand Kit holds your logo, colours and typeface, and applies them across every clip so the whole set reads as one show rather than as a pile of separate uploads.
Does it dub interviews into other languages?
No. There is no dubbing, no voice cloning and no synthetic narration in production. Captions are transcribed from the original audio and burned in, which is what carries a clip in a silent feed, but the voices remain the voices that were recorded.
Does it generate any AI footage or B-roll?
No. Nothing is invented. Every frame in the finished clip came from the interview you supplied, which for journalism and customer testimonials is a requirement rather than a shortcoming.
What if the audio was recorded on one room microphone?
It will work, and the results will be weaker. Distance from a single mic flattens both voices and degrades the transcript, which harms caption accuracy and cut placement together. Individual lapel or desk mics improve clip quality more than any setting in the tool.
How quickly are the clips ready?
Usually a few minutes, depending on the length of the interview and the queue. Fast enough to clip an interview the same day it was recorded, which is when a guest is most likely to reshare it.
What happens to my footage?
It is processed to produce your clips and never published by us. Unreleased interview footage stays unreleased. The privacy policy sets out how it is handled and retained.
Can I clip an interview I did not conduct?
The tool will accept any video you give it. What you may publish depends on who owns the recording and on the rules of the platform you post to, and that call belongs to you rather than to us.
How do I stop the subscription if it is not for me?
One click from the account page ends it, and an email arrives beforehand so the conversion never comes as a surprise. Access carries on to the end of the period already paid for, and anything you exported before then stays yours to use.
Run it on your last conversation
You already know which exchange was the best one in that interview. Check whether the score agrees, and look at how the framing handled your two-shot.