Feed it a long recording — a match, a tournament stream, a conference day, a four-hour broadcast — and the AI marks the moments that stand on their own, cuts each one out, captions it, and grades it 0-100 so the review is a ranked list instead of a scrub through the whole thing.
🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
⏱ That video's a bit long for the demo
Built for multi-hour source footage
Each highlight scored 0-100
Cuts on sentence boundaries
Vertical, square or landscape
Live events clipped as they run
One long video in, a week of posts out
Paste a link. Get the best moments.
Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.
The slow part of highlight work is the watching. That is the part being removed.
1
Hand over the recording
Upload the file or paste a URL if the footage is already published. Paid plans handle up to two hours in a single go; the free demo takes anything under thirty minutes, which is enough to sanity-check the output on one period, one session or one talk before you commit a whole event to it.
2
It watches everything, including the boring parts
The model goes through the full runtime rather than sampling it, which is the only way to find the thing that happened at minute 97 while everyone had stopped paying attention. Segments that hold up out of context get marked, and each is cut around a complete thought rather than a fixed number of seconds either side of a timestamp.
3
Skim the ranked list and take the top ones
Highlights come back captioned and framed, each with a 0-100 grade. Sorting by score turns thirty candidates into a five-minute review. Take the strongest few, adjust the in and out points on anything that needs tightening, and download or schedule them from the same screen.
What the highlight generator handles
Long footage has specific problems. These are the ones being solved.
🔎
Full-runtime review, not sampling
Highlight tools that skim keyframes miss anything that is interesting for what was said rather than for how it looked. Reading the whole transcript and the audio is slower to run and much better at finding the moment nobody thought to note down at the time.
📊
Ranked, not just extracted
Getting forty candidates out of a six-hour event is only half a result, because forty candidates is still forty things to watch. The 0-100 grade decides the order you open them in, so the review becomes a top-down skim you can stop the moment you have enough for the week. What feeds that number is defined in full on the viral score checker; on this page it is simply the thing that stops a long event producing a second pile of work.
✂️
Highlights that start where the thought starts
The tell of an automatic highlight is a clip that opens two words into a sentence. Cutting on sentence boundaries means each highlight begins on a complete thought and lands on the end of one, which is what makes it feel chosen rather than sliced at a timestamp.
🔤
Legible when the room audio is not
Event and match audio is rarely clean — crowd noise, a distant mic, someone shouting off camera. Word-level captions synced to the speech keep the highlight legible when the sound is not, and they carry the clip for the large share of viewers watching muted.
🎯
Framing that follows the person talking
Wide event footage is mostly empty space, so a fixed vertical crop tends to land on a table edge or the gap between two chairs. Face tracking moves the crop to whoever is speaking, which is the difference between a usable vertical highlight and one you throw away.
🗂
Seven layouts for different footage
A panel wants a split or a grid, a screen-share session wants the slides readable, gameplay wants the camera in picture-in-picture. Choosing the layout from what is actually on screen prevents the standard failure of jamming every kind of footage into the same vertical template.
🤫
Dead air and filler cut out
Long recordings are full of pauses, restarts and throat-clearing. Removing them tightens every highlight by several seconds, which matters most in this format — a highlight that takes eight seconds to get going is not a highlight.
🔴
Live events, clipped during the event
If the source is a broadcast rather than a file, clips can be produced while it is still running. For a tournament day or a conference livestream that means highlights go out during the thing, not in a recap post nobody clicks. See the livestream clip generator.
🤖
Drive it from Claude or your own stack
The MCP connector lets you ask Claude to pull highlights from a recording and hand them back in the chat. For batch work — a season of matches, a year of webinars — the API is the sane route.
Where long footage piles up
🎙 Conference and event teams
A two-day event produces dozens of hours nobody will rewatch. Highlights are what the marketing team actually needed, and they are needed the same week.
🎮 Esports and tournament organisers
Bracket days generate more usable material than any human can cut in time. Casting audio gives the model plenty to work with, which makes commentated esports one of the strongest cases here.
📚 Course creators and trainers
A long teaching session contains several self-contained explanations. Pulled out individually they become the search-friendly answers people find you with.
💼 Sales and enablement
Recorded demos and QBRs hold the objection-handling moments worth circulating internally. Highlights make a two-hour call reviewable in six minutes.
⛪ Services and community broadcasts
One long broadcast reliably contains a handful of passages that travel on their own. Captions do the heavy lifting, since most of that viewing happens with the sound off.
📺 Talk shows and panels
Panel discussion is dense in exactly the way detection likes. The split and grid layouts keep more than one speaker on screen so an exchange still reads vertically.
The real problem is the ratio
Highlight work is not hard, it is long. A six-hour recording might hold twelve minutes worth keeping, which means the labour is almost entirely in watching the other five hours and forty-eight minutes to find them. Nobody can do that at a reasonable cost, so in practice most long footage is never highlighted at all — it sits in a drive as an asset that was expensive to capture and produced nothing.
Automating the cut is only marginally useful. Automating the watching is the thing. Once a machine has read the entire runtime and produced a ranked shortlist, a person spends their attention on judgement — is this the moment we want to be known for — rather than on transport controls.
That is also why the score matters more here than on shorter sources. With a twenty-minute video you can review every candidate. With a full event you cannot, so an ordering that is roughly right is worth more than a perfect extraction with no ordering at all.
When this is not the right tool: footage with nobody talking
Detection reads what is said. That makes commentated footage — an esports cast, a broadcast match with a commentary track, a livestream where someone is reacting to the play — a strong case, because the commentator reliably gets loud and specific exactly when something worth clipping happens. The transcript effectively labels the highlights for us.
Footage with no commentary is the weak case, and it would be dishonest to pretend otherwise. A youth game shot from a tripod with only crowd ambience gives the model very little to distinguish a goal from a throw-in. We do not run object tracking on a ball or detect a scoreboard change, so if that is your footage, this is not the right tool and you should not spend a trial finding out.
The line is simple enough to apply before you upload: if a person is describing what is happening, the highlights are findable. If the only signal is on the pitch, they are not.
Individual highlights, not a stitched montage
This produces separate highlight clips, each standing alone, ready to post. It does not assemble them into one montage with a music bed and transitions. That is a deliberate scope choice rather than a missing feature — short-form platforms reward individual moments, and a five-minute compilation performs poorly in a feed built for one idea at a time.
If you do want a traditional highlight reel for a website or an opening slot at an event, the practical route is to take the top-scoring clips out of here and assemble them in whatever editor you already use. The tedious part, finding and cutting them, is the part that has been removed.
The same applies to music. Clips come out with the original audio and burned-in captions, not with a licensed track laid under them. Adding a bed is thirty seconds of work in any editor and keeps the licensing decision where it belongs, with you.
Automatic selection versus cutting the event by hand
Somebody who sat in the room will beat this on any individual highlight, and it is worth being specific about why. They know which speaker the client is paying attention to, they heard the line that got the laugh, and they remember that the best answer of the day came after a question the microphone never picked up. A model reading a transcript afterwards has none of that context and cannot acquire it.
What the manual route cannot do is keep up with runtime. Reviewing a nine-hour bracket day honestly costs real time plus a second pass, so a hand-cut event is two working days before a single file is exported. Most teams never spend those two days. They cut the keynote, promise the breakouts, and the breakouts quietly never happen — which is why the realistic comparison is not "machine clips versus better human clips" but "machine clips versus no clips".
Most people expect the division of labour to fall the other way round: the machine does the mechanical cutting and the human finds the moments. It inverts. Finding is the part that is mechanical at this scale, and choosing which three of fifteen strong candidates represent the event is the part that needs someone who was standing at the back of the room.
So the arrangement that holds up is sequential rather than either-or. Let the full-runtime pass produce the ranked shortlist, then put your editor on the top fifteen candidates instead of on six hours of transport controls. Trimming a second off the front of a good clip is a thirty-second job; locating that clip was the two days.
Alternatives, and when one of them is the better answer
If the footage is one talk under twenty minutes and you already know the moment you want, use the trimmer you already own. Running a whole selection pipeline to extract a clip you could scrub to in ninety seconds is not a saving, and we would rather say that than take the upload.
If the event is being streamed and you mainly want reaction moments, the platform clip button is free and immediate. Twitch and Kick both ship one, and a moderator with a hotkey produces perfectly serviceable markers while the show is running. What that route cannot give you is captions, a crop that follows the speaker, or any ordering across a nine-hour day — so in practice the two get used together, with the hotkey covering what a human noticed live and the full-runtime pass covering the eight hours nobody was watching.
For a flagship event with a real budget, a human clipper or an agency is a legitimate answer and often the right one. The economics only turn when the volume does: one keynote a year is a freelancer, twelve sessions a quarter is a pipeline, and the awkward middle is where teams overspend on both.
And you can build a version of this yourself. An open transcription model plus a script that cuts on timestamps gets you a rough shortlist in a weekend. The parts that take much longer than a weekend are the ones nobody budgets for: word-level caption timing, speaker-following crops, and a scoring pass that is consistent enough to sort by. Worth knowing before you start, because the transcript step is the easy 20% that makes the project look finished.
Where an automatic highlight pass goes wrong
The most expensive mistake is feeding it bad audio. Highlight selection is downstream of the transcript, so a camera mic picking up a hall from thirty feet away produces a mangled transcript and, inevitably, mangled choices. If there is a desk feed, a board mix or a lapel recording available, use that file instead — it changes the output more than any setting.
The second is uploading raw footage with the pre-roll still attached. Twenty minutes of an empty stage while people find their seats is twenty minutes of nothing to score, and on a demo capped at thirty it is most of your test. Trim the top before you submit.
The third is assuming the top-scoring clip is the one to lead with. Ranking is reliable at separating the strong from the weak and much less reliable at separating first place from third. Watch the top few and choose with your own judgement about which one represents the event.
The fourth is running silent footage through it and concluding the tool is broken. It is not broken, it is doing what a speech-driven system does with no speech, and no amount of retrying changes that. Check whether someone is narrating before you upload.
A practical workflow for an event with hours of footage
Split the day the way the schedule already split it. One file per session, per match or per talk gives cleaner results than one enormous recording, because each piece gets judged against itself instead of against the whole day, and a strong moment in a quiet session is not buried under a louder one from the keynote.
Run the first session on its own and review it properly before submitting the rest. Ten minutes spent checking whether the framing holds up on your setup — your stage, your camera position, your two-shot — will tell you whether to change layouts before you have processed twelve files that all need the same correction.
Then queue the remaining sessions, come back once, and sort everything by score across the whole event. Pull the top clip from each session rather than the top ten overall, which keeps your coverage broad instead of six clips from the one speaker who happened to be the most quotable.
Schedule the output across the days after the event rather than the hour it ends. Interest in an event lasts longer than the event does, and a clip released on the Thursday reaches the people who were too busy on the Tuesday.
Export settings that actually matter on event footage
Choose the aspect ratio from how far away the camera was, not from where you intend to post. A lectern shot from the third row crops to vertical perfectly well. The same talk captured from the back of a ballroom does not, and the honest export there is 1:1, which keeps enough of the stage for the speaker to still read as a person, or 16:9 for a recap page or a screen in the lobby. Any clip can be re-rendered into another ratio later without regenerating it, so the ratio is a reversible decision in a way the camera position is not.
Pick the layout from what is on screen rather than applying one choice to the whole event. A two-person exchange wants split-screen, a four-person panel wants the grid, a workshop that lived in its slides wants the screen-share layout so the deck stays legible in a tall frame, and a casted tournament wants picture-in-picture so the play and the reaction are both visible. One event will legitimately use three of these, and the batch that gets a single layout applied across everything is the batch with an unusable third.
Set the caption style before you render the event, not after. Room audio is the norm here and the burned-in text is what carries the clip on a muted phone, so the style that survives contact with a busy hall is the high-contrast one rather than the one that matched your slide deck. If you want a mark on the output, put the event or team logo in the Brand Kit — paid exports come out clean by default, and your own badge is the only one worth carrying on a conference clip.
Render vertical first and treat every other ratio as a second pass. Most event clips only ever need 9:16, and the two or three destined for the recap page or a sponsor reel can be re-rendered once you know which ones they are. Doing it the other way round, every ratio for every clip on the way out, is how a twelve-session event becomes ninety files in which nobody can find anything.
What multi-hour footage has taught us about this pipeline
Observations from operating the pipeline in production — not general advice.
Pre-roll is not dead air, it is reading budget already spent
Selection does not read your runtime, it reads a transcript that is truncated before it ever reaches the model, so everything in the file competes for a fixed number of characters. Event recordings waste them more than any other source we handle: twenty minutes of an empty stage, a mic check, a compere reading housekeeping off a card. All of that transcribes, all of it occupies the window, and none of it can score. In practice, trimming the top of the file before you submit changes the shortlist more than any option on this page, because the characters you free up are spent reading further into the session instead of into the queue for coffee. The character limit itself is set out on the best moments finder.
The upload ceiling is a property of files, and live has no files
Two hours per submission is a limit on an uploaded file, and a tournament day is not a file. A live session works differently in kind: it follows a public broadcast on a window that keeps being extended while the stream runs, and it commits to each clip inside the part of the show that is currently airing rather than from one read of a finished transcript. That is the reason a nine-hour bracket is easier to clip as it happens than it is to clip the next morning, and why we tell event teams to connect the channel before the doors open. The cost is that it is not retrospective: a session you start at 4pm cannot go back for the thing that happened at 11am.
A speaker who is small in the frame is the case tracking cannot rescue
Face tracking moves the crop to whoever is talking, which assumes there is something in the frame worth moving to. A hall camera thirty metres back breaks that assumption: a 9:16 window is roughly a third of the width of a 1080p frame, so when the speaker occupies a tenth of the original shot, every candidate crop contains the same small figure and the choice between them is cosmetic. With nothing it can hold confidently, the framing stays put, because a still crop that is slightly off is far easier to watch than one drifting across a stage hunting for a face. The reason we push event teams to run one session before the other eleven is that this is settled by your camera position rather than by a setting.
How this compares
vs. the timecode log somebody kept during the day
Every event team has a version of this: a marker key on the switcher, a runner with a notepad, a channel where people drop timestamps as things happen. It is real signal and worth keeping. Its structural hole is that it only records the minutes when somebody was in that room, paying attention and not currently rescuing a microphone — so it is dense on the keynote and empty on the third breakout, which is where the unexpected thing usually happened. A full-runtime pass covers the sessions nobody was assigned to. The productive arrangement is to use the log as a checklist against the ranked shortlist afterwards: anything your team flagged that the shortlist missed is worth a look, and anything the shortlist found that nobody flagged is the reason you ran it.
vs. a loudness or excitement detector
Volume-based highlight tools find the moments where something got loud, which correlates with interest and is not the same thing. A quiet, devastating answer registers as nothing. Reading the language finds the substance as well as the noise, and drops the loud moments that turn out to be an alarm going off.
vs. publishing the session archive and calling it coverage
Uploading every session in full is worth doing and most event teams already do it. What the archive cannot do is reach anybody who was not already looking for it: chaptered two-hour recordings serve the delegate who attended and wants the slide again, and nothing about them travels. Highlights are the part aimed at the people who were not in the building, who will never scrub a timeline for the good ten minutes, and who are the audience the event was supposed to grow. The two outputs answer different questions, and only one of them has any reach.
vs. general AI clipping tools
Most are tuned around twenty-to-sixty-minute talking-head video. The differences that show up on long footage are full-runtime review rather than sampling, ranking that survives a large candidate set, and the ability to work on a live event instead of waiting for the recording. Run one event through it and judge on the shortlist.
Frequently asked questions
How long can the source footage be?
Paid plans take up to two hours per upload. The free demo is limited to thirty minutes. For an event that runs longer than two hours, either split the recording into sessions, which usually matches how it was scheduled anyway, or connect the live broadcast and clip it during the event.
Does it make one highlight reel or several clips?
Several separate clips, each cut to stand on its own and ready to post. It does not stitch them into a single montage with music and transitions. If you need a traditional reel, export the top-scoring clips and assemble them in your usual editor.
Will it work on a sports match with no commentary?
Poorly, and we would rather say so up front. Moment detection is driven by speech, so with only crowd ambience there is very little for the model to work from. Commentated or reacted-to footage works well; silent tripod footage does not.
What about esports with a casting team?
That is one of the better cases. Casters get loud and specific precisely when something notable happens, so the transcript effectively marks the highlights. Add the picture-in-picture layout and you get vertical clips where both the play and the reaction are visible.
How many highlights will I get from a long recording?
It reflects what is actually in the footage rather than filling a quota. A dense panel session can produce a couple of dozen candidates; a slow-moving three-hour broadcast might produce five. Padding a thin recording to a round number would only waste your review time.
What does the score do for me on a six-hour recording?
It sets your reading order, and that is the whole of its job here. Thirty candidates come back from an event day and the grade tells you which four to open before lunch and which twenty-six can wait until something is already scheduled. What the number is built from is spelled out on the viral score checker, which is the one page on this site that defines it rather than paraphrasing it. Rank inside a session rather than across the whole event.
Can I get highlights while an event is still live?
Yes, if the source is a broadcast on YouTube Live, Twitch or Kick. Clips are produced during the stream so you can post from a tournament or conference while it is happening, which is when the audience for it exists.
Can I try it on one session before paying?
There is a free demo for anything under thirty minutes, so you can test it on a single session before deciding. After that a 3-day trial costs $1 and Pro is $29 a month. Cancellation is one click and we email before the trial rolls over.
Do the highlights carry a watermark?
Exports on a paid plan are clean. Demo clips carry a small watermark, which is the trade for trying it without paying. You can add your own event or team logo instead through the Brand Kit.
Which aspect ratios do highlights come out in?
Vertical 9:16, square 1:1 and landscape 16:9. Event teams often want both: vertical for social and a 16:9 version for a recap page or a screen in the venue. You can re-render the same clip into another ratio without starting over.
Will it keep two speakers on screen during an exchange?
That is what the split-screen and grid layouts are for. The failure mode with panels is a vertical crop that centres between two people and shows neither properly, and choosing a multi-speaker layout is how that is avoided.
Can it read slides from a screen-share recording?
There is a screen-share layout that keeps the shared content legible in a vertical frame instead of cropping the slide to nothing. It does not summarise the slide contents — the clip is chosen from what is being said about it.
Does it remove pauses and filler?
Automatically, on every clip. Long-form footage is unusually full of dead air, and stripping it typically recovers several seconds per highlight. Tight pacing matters more in this format than almost anything else you could adjust.
Can I adjust a highlight after it is generated?
Yes. In and out points, caption style, title, layout and aspect ratio are all editable, then you re-render. Most highlights need no change, and the ones that do usually need two seconds trimmed off the front.
Can I process a whole season or back catalogue?
Through the API, yes — that is what it is there for. Batch submission and retrieval are documented at developers, and it is a far better path than uploading a hundred files by hand through the dashboard.
Does Claude connect to this?
Yes, through the MCP connector. You can ask Claude to pull highlights from a recording, check on progress and get the finished clips back in the conversation, which is convenient when the footage is already a URL you have on hand.
What languages does it handle?
Transcription spans major languages, with English the most reliable. Because highlight selection depends on understanding the speech, output quality follows transcription quality closely — worth checking on a single session in your language before running an event through it.
Is the audio quality of my recording going to matter?
It matters a great deal. A clean lapel or board feed gives the model an accurate transcript and good selection. A camera mic at the back of a hall gives it a noisy one, and highlight choices degrade accordingly. If you can capture a direct audio feed, do.
Can I use it for interviews and long conversations?
Yes, and conversation is the format detection is strongest on. If the source is a recorded show rather than an event, the podcast clip generator is the page tuned for that workflow.
Do you keep my footage?
It is processed to produce your clips and nothing is published anywhere by us. The clips belong to you. Retention and handling are described in the privacy policy.
How long does a long recording take to process?
Usually a matter of minutes, scaling with runtime and current queue. You do not need to keep the tab open — the clips are in your library when you come back, which is the sensible way to run a two-hour file.
Can I schedule the highlights to post automatically?
Yes. Connect TikTok, Instagram and YouTube and queue them from the review screen. For an event, the useful pattern is to space the top highlights across the following week rather than dumping all of them the same afternoon.
Will it find a moment if nobody reacted to it?
Sometimes, because it reads the substance of what was said rather than the volume of the response. A precise, quiet statement can score well. But it has no knowledge of your audience or your history, so a moment that is only meaningful to insiders will not register.
Does it add music or transitions?
No. Clips come out with the original audio and burned-in captions. Keeping music out means you decide what is licensed for your use, which for events and brands is not a decision worth automating.
What is the fastest way to know if this suits my footage?
Take the single most typical hour you have, cut a thirty-minute piece of it, and run it through the free demo. You will learn more from one honest test on real footage than from any feature list, including this one.
What happens if I only need this for one event?
Then take the trial, process the event, and cancel in one click from your account — access continues through the period you already paid for. A reminder email goes out before the trial converts, so a one-off project stays a one-off cost.
Give it your longest recording
The best test is the file you have been avoiding. Run it once and see what the shortlist looks like against what you would have found yourself.