HomeAI Video ToolsMulti-Speaker Split Screen
Multi-Speaker Split Screen

Four people, one vertical frame, nobody cropped out

Two people sitting across from each other end up at opposite edges of a 16:9 frame, and a vertical crop keeps neither of them. ClipSpeedAI stacks the speakers instead of choosing between them — a two-pane split, a 3-up or a 4-up grid — with tracking inside each pane so faces stay centred while people move.

🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
  • Split, 3-up and 4-up layouts
  • Tracking inside each pane
  • Layout chosen from what is on screen
  • Reactions stay visible, not cropped away
  • Works on live panels and co-streams
One long video in, a week of posts out

Paste a link. Get the best moments.

Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.

https://youtu.be/z_bX3runikk
Source video: Matthew McConaughey's Concerns Over AI Source video 15:27
Matthew McConaughey's Concerns Over AI
9 clips found  ↓  top 5
Matthew McConaughey speaking, captioned vertical clip
Score 92
9:16
Joe Rogan gesturing, captioned vertical clip
Score 90
9:16
Matthew McConaughey listening, captioned vertical clip
Score 89
9:16
Joe Rogan mid-sentence, captioned vertical clip
Score 89
9:16
Matthew McConaughey talking to camera, captioned vertical clip
Score 88
9:16

How multi-speaker clips get built

You do not draw masks, set keyframes, or nominate who goes on top.

  1. 1

    Give it the conversation

    Paste the YouTube URL of the episode or upload the recording — a Zoom export, a Riverside session, a conference panel, a co-stream VOD. What matters is that it is the finished combined video with everyone visible in the frame. Two hours per file on paid plans, thirty minutes on the free demo.

  2. 2

    Faces are located and the layout follows the count

    The engine detects how many people are actually in the picture across the moment it is cutting. Two faces get a split, three get a 3-up, four get a grid, and a stretch where only one person is in shot gets a single full-frame crop instead of an empty pane. The footage decides, not a setting you picked at the start.

  3. 3

    Review clips where everyone is legible

    Each cut arrives at 9:16 with the panes stacked, word-by-word captions burned in, and a 0-100 score. If a particular moment would play better as a single speaker filling the frame, switch the layout on that clip and render it again.

The layouts, and what each one is for

Seven layouts ship in total. These are the ones that exist because conversations do not fit in a vertical rectangle.

Split — two speakers stacked

The workhorse for interviews and two-host shows. Both faces sit in their own pane, one above the other, each cropped tight enough to read on a phone. The reason this beats cutting between speakers is that a listener’s reaction is often the better half of the moment, and a single-speaker cut throws it away.

3-up — a host and two guests

Three panes in a vertical column, which is roughly the practical limit for faces that stay readable at phone size. It suits panels where one person is steering and two are responding, since the eye can follow a three-way exchange without the frame having to cut at all.

4-up — the full grid

A two-by-two arrangement for four-person recordings and remote roundtables. Each tile is smaller, so this works best on close-framed webcam footage where heads already fill their source frames and badly on wide room shots where everyone is a distant figure at a table.

Tracking inside every pane

A pane is not a static crop of a fixed rectangle. Each one follows its speaker, so somebody leaning back, turning to a second camera or gesturing off to one side stays framed. Movement is smoothed rather than snapped frame to frame, because a crop that twitches on every small motion is more distracting than a slightly loose one.

The layout is chosen, not imposed

Forcing a split onto footage that only ever shows one person produces a black rectangle where the second speaker should be, which is the most common way automated vertical conversion fails visibly. Layout is selected from what is actually in the frame during that clip, and it is changeable per clip afterwards.

Screenshare when there is something to show

Panels that pull up a slide, a chart or a demo get the screenshare layout instead: the shared display takes the readable portion of the frame with the speaker kept visible below it. Losing the slide is as bad as losing the speaker when the whole point of the clip was the number on screen.

Picture-in-picture for co-streams

Gameplay with two or more people on cam is a different geometry problem, and the PiP layout keeps the action large with the reactions inset. This is the layout you want for a clip where what happened in the game and how people responded are equally the story.

Captions that share the space

Text and stacked panes are competing for the same vertical inches. Word-by-word captions are burned in across eleven styles, and on a 3-up or 4-up the plainer styles read better than the heavy ones simply because there is less room. Captions carry most of the comprehension on a muted feed, so they get priority over decoration.

Scored, cut on sentences, cleaned up

Multi-speaker clips get the same treatment as everything else: cut on sentence boundaries so a clip never opens mid-word, filler and dead air removed — which matters more in conversation because turn-taking generates gaps — and a 0-100 grade to rank the batch.

Live panels and multi-person streams

Connect a YouTube, Twitch or Kick channel and multi-speaker clips are produced while the broadcast is still running. A three-way argument on a live show is exactly the kind of moment that is worth far more posted within the hour than the next morning.

Formats this was built for

Interview podcasts

The guest says something and the host reacts, and both halves are the clip. See the podcast clip generator for the wider workflow around episode-length recordings.

Conference panels

Four chairs on a stage in one wide shot is nearly unusable vertically. Cropping each speaker into their own pane is what makes a panel postable at all.

Debate and reaction shows

Disagreement is the format, and disagreement is unreadable if you can only see whoever is talking. Both faces at once is the whole point.

Remote and Zoom recordings

Grid-view calls already put everyone in a box, and restacking those boxes vertically converts a call recording into something that works on a phone.

Co-streamers and duos

Two webcams over gameplay, or two friends reacting to the same thing. Both cams stay in shot instead of the clip picking a side.

Corporate roundtables and webinars

Multi-host webinars and internal panels produce long recordings with a handful of quotable exchanges buried in them, and those exchanges usually involve at least two people.

Why a centre crop is the wrong tool for a conversation

What a conversation loses to a vertical crop is the other person. Seat two people across from each other and they occupy opposite edges of a wide frame, so the strip that survives is the gap in between — a bookshelf, a mic arm, a doorway, a lamp. Neither face is in it. It is the single most recognisable failure of naive auto-cropping, and once you have seen it you cannot unsee it on other people’s clips. If you want the pixel-by-pixel version of how little of a wide frame comes through, it is worked out in full on the vertical video converter page; the part that matters here is simpler, which is that a two-hander gives you no version of the shot where cropping costs you nobody.

The next attempt is usually to track the active speaker and cut between them. That is better, and for a monologue it is correct. In a conversation it quietly removes the thing that made the moment work. Someone delivers a line and the other person’s face does something, and if the clip is on the speaker for all forty seconds the reaction never happened as far as the viewer is concerned.

Stacking is the answer because vertical has height to spare. A 9:16 frame is tall enough for two cropped faces at a size that reads fine on a phone, and the format is watched close to the eyes on a small screen where smaller faces still work. What is scarce in vertical is width, and a split screen spends the dimension you have on the problem you have.

What a split screen has to get right

The first requirement is that each pane keeps its person. If tiles reshuffle mid-clip, a viewer has to re-learn the arrangement every few seconds and the clip becomes exhausting to watch, no matter how good the exchange is. A stable arrangement is more important than an optimal one.

The second is tracking that does not fidget. Cropping tight to a face means small head movements translate into large frame movements, and a pane that re-centres on every micro-motion reads as a camera operated by somebody nervous. Following the subject with smoothing applied is the difference between a crop you notice and one you do not.

The third is knowing when a pane should not exist. Not every second of a four-person recording has four people visible — a director cuts to a close-up, someone leaves to get coffee, the camera pushes in on the speaker. A layout that insists on four tiles will render empty boxes. Choosing the layout from what is actually in the frame during that specific clip is the fix.

The fourth is leaving room for text. Captions are what most of the audience is reading, since the majority of feed viewing happens with the sound off. A layout that fills every pixel with faces and then drops caption text on top of somebody’s mouth has solved the framing problem by creating a legibility one.

A recording workflow that makes split screen clips better

Most of the quality of a multi-speaker clip is decided before you press record. If you already know panels are going to be clipped, a few habits at capture time are worth more than anything a tool can do afterwards.

Frame each person tighter than feels right for the landscape edit. A wide two-shot with plenty of headroom looks composed on a monitor and produces small, distant faces once it is cropped into a pane. Faces that already fill their portion of the source frame survive the crop with resolution to spare.

Keep everyone in shot rather than cutting to close-ups in the master edit. Every second where the director has punched in on one speaker is a second where no split screen is possible, because the other person genuinely is not in the pixels. If you publish both a directed cut and a raw multi-view recording, submit the raw one.

Light both sides of a remote conversation. It is common for a host in a studio and a guest on a laptop webcam to end up in the same grid, and the difference is jarring at pane size. Asking a guest to sit facing a window costs nothing and closes most of the gap.

Finally, leave a beat after a good line before jumping in. Overlapping speech is the hardest thing to clip cleanly, and the reaction you want in the second pane needs a moment of quiet to be visible in. Interviewers who learn this get noticeably better clips out of identical conversations.

Where this stops working, and when not to reach for panes at all

It composites from your finished video, not from separate camera files. If you recorded isolated per-speaker tracks in Riverside, Zoom or an equivalent, export the combined edit and submit that. There is no multitrack ingest that takes four ISO files and arranges them for you, and if that is what you need then a full NLE is the right tool.

Four panes is the ceiling. A six-person panel does not become a 6-up, because six faces on a phone screen are too small to carry any expression. What you get instead is the people who are actually speaking during that clip, which is usually the honest edit anyway.

It cannot show a person who is not in the picture. If your source is single-camera and the director cut to a close-up of the host for that whole exchange, the guest does not exist in those pixels and no layout can retrieve them. Multi-speaker layouts work on multi-speaker footage, which sounds obvious but is the most common mismatch we see.

And wide room shots are a genuinely hard case. If four people are sitting at a distant conference table, cropping each of them into a small tile gives you four low-resolution figures rather than four faces. Close-framed webcams and proper podcast setups produce far better results than a single camera at the back of a room.

What an editor does by hand that a layout engine does not

It is worth being precise about what you are giving up, because the honest answer is not nothing. A person cutting a panel by hand breaks the split deliberately: they hold both faces through the setup, then punch in to full frame on the exact syllable where the reaction lands, then pull back out. That rhythm — two panes, one pane, two panes — is the thing a good conversation editor is actually doing, and the layout here is chosen once per clip rather than changed inside it.

A human also sizes panes unevenly. Sixty per cent of the height to whoever is carrying the exchange and forty to the listener reads more naturally than an even split, and an editor adjusts that per moment. Panes here are equal by construction. That is a real trade and not a subtle one on a two-hander where one person talks for most of the clip.

What swings the other way is that the engine does the same thing on clip fifteen as it did on clip one. Say you run a weekly three-person show and pull eight clips from every episode: that is four hundred clips a year, each needing masks, positions and keyframes whenever somebody shifts in their chair. Nobody sustains hand-built splits at that rate, and the version that actually happens is not the careful edit — it is three careful clips and five that never got made.

The biggest mistake here is deciding this in the abstract. Cut the batch, find the one clip that genuinely deserves an hour, and rebuild only that one in a timeline. Everything you learn doing it makes the next batch easier to review, and the other seven clips went out on Tuesday instead of never.

Alternatives when panes are the wrong answer

Panes are a solution to a specific problem — several people, one wide frame, a vertical output — and there are several moments where the alternatives beat them outright.

A single tracked crop of whoever is speaking is the strongest option more often than people expect. Any moment that is really a monologue with an audience, any stretch where one answer runs a full thirty seconds, any clip where the other person is politely nodding: a full-frame face reads better than a half-frame face plus a half-frame of nodding. It turns out this is also the right call on wide room shots, because one legible face beats four unreadable tiles.

Keeping the original 16:9 is an alternative nobody considers because vertical has become the default. For YouTube, LinkedIn and an embed on your own site the landscape frame already holds everyone at full size and needs no layout decision at all. A 90-minute conference panel destined for a company blog does not need to be vertical, and forcing it there costs resolution for no gain.

For gameplay and co-streams, picture-in-picture usually beats an even split, since the action and the reaction are not equally important and PiP says so in the composition. For a moment built around a slide or a chart, the screenshare layout is the alternative: the thing being shown is the content and the faces are context.

And if you have isolated per-speaker camera files, a full editing suite with a multicam workflow will beat any automatic layout, because it is cutting between full-resolution sources instead of cropping regions out of one composite. That is a different amount of work, but if the deliverable is one flagship clip a month rather than eight a week, it is the better tool and we would rather say so.

Troubleshooting a panel clip where the panes came back wrong

Most disappointing multi-speaker clips are not broken, they are mismatched — the layout is doing something reasonable for footage other than yours. Four symptoms account for nearly all of it, and three have a fix that takes one re-render.

The split picked the wrong pair. On a five or six-person panel the two panes go to the people carrying that stretch of talk, so somebody who lands one sharp interjection can lose their pane to whoever simply spoke more seconds. If the interjection is the reason the clip exists, override that clip to a 3-up and render it again — three panes will hold all the participants who matter, and the person you wanted is back in the frame.

The clip opens on the listener. Conversations overlap at the edges, and a cut placed on a clean sentence boundary can still begin with the tail of the previous turn while the camera is already on the person about to answer. Nudge the in point a beat later. It is a two-second adjustment and it is the single most common reason a good exchange feels like it starts in the middle of something.

The framing is fine but the clip drags. Filler and dead air have already been taken out by this point, so what is left is real conversational time — thinking, nodding, a pause before an answer. A split doubles how much of the screen is visibly doing nothing during those beats. Either tighten the in and out points around the exchange itself, or drop to a single crop, which hides the waiting instead of framing it twice.

One pane looks worse than the other. That is a studio host next to a guest on a laptop webcam, and no layout corrects exposure or lens quality. What does help is spending fewer panes: a two-pane split gives the weaker source considerably more width than a 4-up tile does, so if the choice is available on that clip, take the split.

What we have learned running this engine

Observations from operating the pipeline in production — not general advice.

The layout has to be read across the clip, not sampled at one frame

Face count is not a constant inside a thirty-second exchange. A directed edit will punch in on the host, hold there four seconds, then cut back out. If the layout decision is taken from a single frame, it can land inside that punch-in and render the whole clip as a solo crop with the guest absent from their own answer. Reading detections across the entire span being cut is what avoids that, and the reason it matters far more on panels than on a lone presenter is that a multi-camera conversation changes its composition several times inside a window a talking head would fill with one unbroken shot.

A tracker that has stopped working looks completely fine on one face

We once shipped a long stretch of clips where face detection had quietly stopped reporting and every crop fell back to the centre of the frame. On a single presenter that is nearly invisible: somebody who stands still is framed almost identically whether the tracker is running or not, which is precisely why it survived as long as it did before anyone caught it. On a two-person split the same failure is obvious inside a second, because the panes stop following anybody and one of them settles on the wall over a shoulder. If you ever suspect tracking in your own output, test it on multi-speaker footage first — it breaks loudly there and quietly everywhere else.

Four panes is a resolution decision before it is a design one

A 4-up on a 1080 × 1920 output hands each tile 540 pixels of width, and that number governs the result more than the choice of layout does. Close-framed webcam footage, where a head already fills its source frame, scales down into 540 and stays readable. A single camera at the back of a room gives each face a small patch of the original frame and then asks for that patch to be enlarged up to 540, and enlargement does not restore detail that was never recorded in the first place. Measuring how tall a head actually is in your source predicts the quality of a grid better than anything in the settings.

Compared with the alternatives

vs. building the split by hand in Premiere or Resolve

Manually this means duplicating the clip, masking each speaker, positioning both, and keyframing the crops whenever someone shifts in their chair. For a single hero clip that is a reasonable afternoon. For every clip from a weekly show it is a job, and it is the kind of job that gets skipped first when a deadline moves.

vs. multicam editing with separate camera tracks

A proper multicam edit with ISO recordings gives you the most control there is, including full-resolution close-ups of each speaker. It also requires the recordings, the sync, and someone who knows the workflow. This works from the single video you already published, which is why it fits people who do not have an editor.

vs. plain auto-reframe or a centre crop

Auto-reframe features in general editors follow motion or a single subject, which handles a solo talking head perfectly well. Point one at a panel and it either sits between people or oscillates between them. The layouts here exist specifically because a one-subject assumption breaks on a conversation.

vs. the split screen in other clipping tools

Split screen is fairly widely offered now; 3-up and 4-up are less so, and the quality gap is usually visible in whether the layout adapts when the number of visible faces changes mid-clip. Run a panel episode through and look for empty panes and reshuffling tiles — those are the tells.

Frequently asked questions

How do I make a split screen clip from a podcast?
Submit the episode as a YouTube link or a file. Faces are detected across the moment being cut, and a two-person exchange comes back as a vertical split with each speaker in a pane. You do not select the layout in advance or mark up who is where — that is read from the footage.
How many speakers can be on screen at once?
Up to four, in a two-by-two grid. Two people get a stacked split and three get a vertical 3-up. Beyond four the tiles become too small for faces to read on a phone, so larger panels are handled by showing whoever is actually speaking during that clip.
Can I pick the pane arrangement myself?
The layout is selected automatically from what is visible, and you can override it per clip and re-render. Overriding is worth doing when a moment is really about one person — a single full-frame crop often hits harder than a split, even in a two-person show.
What if only one person is on screen for part of the clip?
That stretch is treated as single-speaker rather than being forced into a half-empty split. Rendering a black pane where nobody is visible is the most obvious sign a tool is applying a layout blindly, and it undermines an otherwise good clip.
I recorded ISO tracks in Riverside — which export should I actually submit?
The grid or gallery export, not the speaker-switching one. Most recording platforms offer both: a composed view holding everyone at once, and an active-speaker view that cuts to whichever person is talking. The second reads better as a recording somebody watches end to end, and it is the weaker input here, because it has already made the choose-one-face decision permanently, in the file, on exactly the footage where you wanted both. If the switching view is all your platform will give you, it still works — you get single-speaker clips instead of splits. Stitching raw ISO files into a grid yourself is an editing-suite job, and doing it once at export costs far less than confronting it on every clip.
What does it cost to run a panel through?
Anything under thirty minutes goes through the free demo, which covers one panel segment or a short two-hander. Full access then starts at $1 for three days and settles at $29 a month for Pro. One click ends it, and a reminder email precedes the conversion.
Will the panes stay in the same position?
They are meant to. A stable arrangement matters more than a theoretically optimal one, because tiles that reshuffle mid-clip force a viewer to re-orient every few seconds. Where a person is in the frame should be something the audience learns once.
Does each pane follow the speaker or is it a fixed crop?
It follows. Each pane tracks its person so leaning, turning and gesturing stay in shot, with smoothing applied so small movements do not make the crop jitter. A twitchy pane is more distracting than a slightly loose frame.
Does this work on Zoom recordings?
Yes, and grid-view calls are among the easier cases because each participant already occupies a defined box in the source. Restacking those into a vertical arrangement is what converts a call recording into something watchable on a phone.
Our panel was filmed from the back of the room — is there anything that helps?
Two things, and neither of them is a setting. If a 4K master of that shot exists anywhere, submit it rather than the 1080p export: each tile is cut from whatever real pixels are there, and a 4K frame simply has more of them to hand to every face. Failing that, override the grid on the clips you actually care about and take a single tracked crop of whoever is speaking. One legible face beats four unreadable ones, and it is the same trade a human editor makes on the same footage. If neither option is open, the recording is telling you it wanted a closer camera, which is a fix for the next panel rather than this one.
Are captions included on split screen clips?
Yes, burned in and synced word by word, in your choice of eleven styles. On a 3-up or 4-up the simpler caption styles work better because vertical space is already committed to panes, and legibility beats decoration when most viewing is muted.
Can it do split screen on a live stream?
It can. Connect YouTube, Twitch or Kick and multi-person clips are cut during the broadcast. Live debate and co-stream moments lose most of their value overnight, so producing them mid-stream is the difference between timely and archival.
What if the shared screen is a spreadsheet or a terminal rather than a slide?
Slides survive the shrink because their text was already sized to be read from the back of a room. A forty-row spreadsheet, an IDE or a terminal at its normal font size does not. The layout keeps the shared window whole instead of cropping into it, but whole at phone width is still small, and nothing downstream can enlarge eight-point text into something legible. For those moments the stronger clip is usually the person explaining the thing, with the captions carrying the detail the pixels cannot. Two follow-ups come up a lot here. The presenter inset is part of a fixed composition, so you switch between layouts rather than resize or drag the box. And the layout is settled for the clip as a whole, from whatever dominates its span — a share running most of the clip stays in screenshare to the end, while one that only covers the opening is treated as a conversation instead. When a share drops midway and the result looks wrong, either trim the out point to the moment it drops or move that clip to a split and re-render.
How does it know who is speaking?
Speaker activity is inferred from the audio and the visible faces together, which is what drives both the layout choice and how the panes are emphasised. On footage where nobody is visible during a line, the layout falls back to what is actually in shot.
Can I export in something other than 9:16?
Yes. Vertical is the default because that is what Shorts, Reels and TikTok want, but 1:1 square and 16:9 landscape are both available. A split reads differently in square, so it is worth previewing before committing a batch.
Do exported clips have a watermark?
Only the free demo watermarks. Paid exports carry nothing of ours in the frame, and the Brand Kit will place your own logo consistently across a batch if you want the clips visibly yours.
Does it remove filler and pauses in a conversation?
Yes, and multi-person recordings need it more than solo ones because turn-taking creates gaps of its own. There is a detailed breakdown on the silence remover page if pacing is your main concern.
How long can my source recording be?
Two hours per upload on a paid plan, thirty minutes on the demo. Long panel recordings can be split across submissions, or you can connect a channel and let live clipping handle the runtime instead.
How many clips will a panel episode produce?
It depends entirely on the conversation rather than a fixed quota. A ninety-minute panel with a lot of agreement might yield five worthwhile moments; a heated one produces considerably more. Only moments worth posting are returned, which is why the count varies.
Can I edit the framing after the fact?
You can change the layout, adjust the in and out points, switch caption style and edit the title, then re-render. There is no per-pane manual positioning, which is the trade being made for not having to build any of it yourself.
Does it work if people are wearing headphones or masks?
Headphones are fine and extremely common in this footage. Anything that obscures most of a face makes detection harder, and in that situation a wider single-speaker crop is usually the better choice for that clip.
Is there an API for multi-speaker clips?
Yes. The same engine including layout selection is reachable over the developer API, and through an MCP connector that lets Claude submit an episode and hand back finished clips. Both are documented at developers.
Which languages does it support?
Layout and face tracking are language-independent, but captions and moment selection depend on transcription, which is strongest in English. Run a real episode through the free demo if you work in another language and want to judge the captions rather than the framing.
How long does a panel episode take to process?
Usually a few minutes, longer for a two-hour recording or a busy queue. You can close the tab while it works — the clips will be in your library when you come back to it.
What happens to my recording?
It is processed to produce your clips and nothing is published by us anywhere. Rights to the output stay with you. The privacy policy covers handling and retention in detail.
Is there a lock-in after the three-day trial?
None. One click in your account stops the subscription and access runs to the end of whatever period you have already paid for. Before the trial rolls into a subscription we send an email, so the decision is always yours to make in advance.

Put the whole conversation in one vertical frame

Take an episode where the reaction shot matters as much as the answer. Watching the split version next to a centre-cropped one settles the argument in about ten seconds.

⚡ Clip My Conversation — $1 trial
3-day trial · just $1 · cancel anytime