Point it at a recording and it works through a long stretch of the transcript, marking the segments that stand on their own and ranking them 0-100. You get a shortlist with timestamps, already cut and captioned, instead of a scrub bar and an evening. Reading is deepest across the first half-hour, so a two-hour file is best sent in parts.
🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
⏱ That video's a bit long for the demo
Reads ~30 minutes of speech at a time
Each moment ranked 0-100
Timestamps you can jump to
Cut on complete thoughts
Finds moments during a live stream
One long video in, a week of posts out
Paste a link. Get the best moments.
Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.
You hand over the recording. The searching, ranking and cutting all happen without you in the loop.
1
Give it the whole recording
Paste a YouTube link or upload a file — a two-hour interview, a workshop recording, a stream VOD, a Zoom export. There is no project to configure and nothing to pre-mark. Paid uploads run up to two hours each; the free demo handles anything under thirty minutes so you can test the detection before spending anything.
2
The transcript gets read, not skimmed
A person hunting through a long recording pokes at it: beginning, a couple of stabs at the middle, the end. The model reads instead. It works through the transcript, looking for where a thought completes, where a claim gets specific, and where an answer actually answers something. There is a ceiling on how much transcript goes into one pass — roughly the first half-hour of speech — so a very long source is read most closely at the front, and splitting a two-hour file into parts is how you get even attention across all of it.
3
You get a ranked shortlist
Results come back as a short, ordered list — not a dump of every timestamp with a keyword in it. Each entry carries a score out of 100, a title, and a finished vertical cut with captions burned in, so you can judge the moment as a viewer would rather than reading a transcript excerpt and guessing.
4
Keep, discard, or re-cut
Nothing is locked. Nudge the in and out points if you want an extra beat before the punchline, swap the caption style, rewrite the title, or throw the whole clip away. The finder is doing triage; the final call on what ships stays with you.
What the finder looks for
Detection is the whole product. Here is what it is actually measuring, and what it does with what it finds.
🧠
Semantic detection, not keyword matching
Searching a transcript for "the most important thing" finds throat-clearing. The model evaluates meaning: whether a segment resolves something, introduces a specific number, contradicts an expectation, or delivers a payoff the preceding sentences set up. That is why it surfaces moments that contain none of the words you would have thought to search for.
📈
A 0-100 rank on every find
A finder that returns forty equally-weighted timestamps has moved the problem, not solved it. Each moment is graded on how it opens, how quickly it gets to the point, and where its peak sits — so the output is an order of operations, not another list to triage.
🧩
Moments that survive being removed
A great line that depends on ten minutes of setup is not a moment, it is a callback. Detection favours segments that hold up cold, for a viewer who has never heard the rest of the recording — which is the only condition a clip is ever watched under.
📍
Timestamps you can act on
Every find comes with its position in the source. Useful even when you plan to edit elsewhere: you can jump to 01:14:22 in your own project file and cut it yourself, using this purely as the search step and doing the craft work in the editor you already like.
✂️
The moment starts on the moment
Detection is worth nothing if the cut lands four seconds early on someone clearing their throat. Clips open and close on sentence boundaries, and filler words and dead air inside the segment are stripped, so the first second of the clip is the first second worth hearing.
🔴
Finds moments while a stream is running
The hardest recording to search is one that has not finished. Connect a YouTube, Twitch or Kick channel and detection runs during the broadcast, surfacing moments minutes after they happen rather than after you download a six-hour VOD the next morning.
🖼
Each find comes back postable
Finding is only half the job if the result is a note in a spreadsheet. Every moment is reframed to 9:16 with the speaker tracked, captioned word-by-word, and laid out to match the footage — so the shortlist is a set of finished clips you can post, not a research document.
🗃
Everything stays in one place
Moments from every recording you run live in the same library with their scores and titles attached. That turns a scattered folder of exports into something searchable later, when you want the one clip where you explained pricing well and cannot remember which session it was in.
🤖
Detection from Claude or your own code
Run the finder conversationally through the MCP connector — ask Claude for the best moments in a URL and get scored clips back in the chat — or call the same detection engine directly from your stack. The developer docs cover the endpoints.
Who needs a moment finder
🎙 Interview and podcast hosts
You were in the conversation, which makes you the worst judge of it. Hosts consistently remember the parts they enjoyed rather than the parts that travel. A podcast clip generator view of the same episode is often uncomfortable and usually right.
🎮 Streamers with unwatchable VOD lengths
Six hours is not searchable by hand and you already forgot half of it. Detection is the only realistic path from a full stream to the four moments in it.
🏫 Course creators and workshop leaders
A three-hour workshop contains maybe six self-contained explanations. Those six are the marketing for the other two hours and fifty minutes.
📞 Revenue teams reviewing calls
Recorded calls hold the exact objection language your prospects use. Finding where a rep handled one well is a training asset most teams never extract because nobody wants to re-listen to forty calls.
🎬 Editors working from raw footage
Use it as an assistant rather than a replacement. Let the model produce the candidate list overnight, then apply your own taste to a shortlist instead of a timeline.
📣 Event and conference organisers
A day of talks is a year of clips nobody ever cuts. Running each session through detection turns the recording budget into something with an afterlife.
What actually counts as a moment
Most people describe the thing they are hunting for as "the good part", which is not a definition you can build against. In practice the segments that perform share a shape. Something is set up and then resolved inside the clip. There is a turn — an expectation gets broken, a vague topic becomes a specific number, a polite disagreement becomes a real one. And crucially, none of it requires the viewer to know what came before.
Laughter is the classic false positive. Audio energy spikes are trivially easy to detect and they are why so many automatic highlight tools return a reel of people chuckling at nothing. A laugh marks that something happened; it does not tell you whether the thing that happened means anything to a stranger scrolling past. Detection has to read the sentence that caused the laugh, not the laugh.
The opposite failure is quieter and costs more. The single best ninety seconds of a technical interview is often calm, low-energy, and completely flat on a waveform — one person explaining the thing they understand better than anyone. Nothing about the audio marks it. Only the content does.
Why finding is harder than cutting
Cutting is a mechanical task with a known cost. Twenty seconds of trimming, a caption pass, an export. It is annoying but it is bounded, and you can hand it to anyone. Finding is unbounded: to be sure you have the best moment in a two-hour recording you have to have heard all two hours, and then hold every candidate in your head at once to rank them.
That is the real reason so much footage never gets clipped. Not the editing — the review. A working week does not contain enough uninterrupted attention to re-watch everything you record, and the half-measure people fall back on is clipping the bit they happen to remember. Memory is heavily weighted toward whatever happened in the first ten minutes and whatever happened last.
There is also a self-assessment problem that no amount of discipline fixes. You cannot hear your own explanation the way a stranger hears it, because you already know what you meant. The moment you think is your sharpest is frequently your most familiar, which is a different thing. An outside pass — even an imperfect machine one — corrects for that bias simply by not having it.
A workflow for turning finds into published clips
Run the search first and do nothing else with it that day. The instinct is to open the top result and start polishing, which puts you back in the timeline you were trying to escape. Read the shortlist as a list, mark the ones you would defend to someone whose opinion you respect, and delete the rest before you get attached to any of them.
Then work in one batch. Fixing openings, rewriting titles and choosing caption styles are all the same kind of attention, and doing them together for six clips takes less than half the time of doing them one at a time across six sittings. This is also when to catch the clip that technically scored well but says something you would not want quoted back to you.
Publish across days rather than in one dump. A single source can fill a fortnight, and spacing the posts gives each one a fair chance to be seen instead of having them compete with each other in the same feed. Queue them, then go and record the next thing rather than watching the first one perform.
Every few weeks, compare what actually did well against how it was scored. You are calibrating your own reading of the shortlist, and after two or three cycles you will know whether your instinct runs ahead of the ranking or behind it. That is the point at which the tool starts saving judgement as well as hours.
Troubleshooting a shortlist that missed the obvious moment
The most common complaint is that the list came back without the one segment you already knew was the best thing in the recording. Before concluding the ranking is broken, check where that segment sits in the runtime. One pass reads roughly the first half-hour of speech, so a moment landing at the ninety-minute mark of a long upload may never have been read at all. Re-send the source in shorter parts and look again before you judge the detection.
If the moment was inside that window and still did not surface, the usual cause is dependency. Segments that need the preceding twenty minutes to make sense are marked down on purpose, so a payoff whose setup lives somewhere else in the recording drops. Check the bottom half of the list rather than assuming it is absent — a near-miss usually sits at rank eight, and one you can rescue by dragging its start point earlier was never really a miss.
A shortlist that is weak all the way down is normally a transcript problem rather than a ranking problem. Crosstalk, a distant mic and heavy room noise all corrupt the words the model is reading, and nothing downstream can recover meaning that never made it into the text. Play thirty seconds of the source back and ask whether you could follow it with your eyes shut. If you could not, fix the capture before spending another pass on it.
And if the timecodes sit a second or two wider than where you would have cut, that is intentional rather than sloppy. Starts are placed slightly ahead of the first word so the clip never opens on a syllable already underway. Drag the handles to taste and re-render — the export is built from your points, not from the original suggestion.
Reading a transcript by machine versus watching the recording yourself
The two methods do not have access to the same evidence, and almost every disagreement between them traces back to that. You watching a recording get tone, timing, the pause before an answer, the face someone pulls while another person is still talking, and everything you already know about who is listening. The picker gets words and their positions in the runtime. It is a narrower input than most people assume, and it explains both what it is unusually good at and where it is blind.
Narrow turns out to be an advantage for the specific job of triage. Text has no charisma, so a segment that sounded compelling because of how it was delivered has to survive on what was actually said — which is the same test it faces from a stranger reading captions with the sound off. A person re-watching their own recording cannot run that test, because they hear the delivery again whether they want to or not.
The blind spots are equally structural. Anything whose meaning lived in the room — a gesture, a slide, a look exchanged between two people, a demonstration nobody narrated — leaves no trace in the words and cannot be ranked. So does sarcasm delivered deadpan, which reads on the page as the opposite of what was meant. Those are the segments a human reviewer should be scanning for, and they are worth a deliberate pass rather than a hope that the ranking catches them.
The arrangement that works is not a contest. Let the machine read all of it and hand back an ordered list, then spend your own attention on the two jobs it cannot do: vetoing anything that misrepresents you, and rescuing the moment whose substance was visual. That is fifteen minutes of your judgement applied to a shortlist rather than two hours of it applied to a scrub bar.
Where the finder gets it wrong, and when not to run one at all
Start with the cases where the search is not worth running. A recording under about fifteen minutes does not need one — you can watch the whole thing in the time it takes to read a shortlist, and your own memory of it is still intact. Neither does a session you already know cold, where you could name the timecode of the good part from memory. The finder earns its place on material you have not reviewed and were realistically never going to.
It has no idea what your audience is already tired of. If you have explained the same framework in nine videos, the tenth explanation will still score well, because the model is grading the segment and not your posting history. Only you know that one is played out.
It will miss inside jokes, running bits, and community references. A moment that is enormous to two thousand regulars and meaningless in isolation is exactly the kind of thing detection undervalues, and it is worth scanning the low-scoring finds for those before you discard them.
And it cannot find a moment that is not in the recording. A meandering ninety minutes with no clear thought in it produces a thin shortlist, and that is the honest result rather than a failure to pad. If the finder consistently returns two moments from your hour-long episodes, the useful conclusion is about the episodes, not about the tool.
What running the moment finder in production has taught us
Observations from operating the pipeline in production — not general advice.
Half an hour of speech is the real reading window
The picker is handed the transcript as plain text, and that text is capped at 28,000 characters before it reaches the model. Once the inline timecodes are counted, a spoken minute costs somewhere around nine hundred characters, which puts the practical ceiling near the first thirty minutes of talking. That is the reason we tell people to send a two-hour recording in parts instead of trusting one pass to weigh the final hour as carefully as the opening one. A tool that truncates quietly and then claims to have read everything is a thing you only find out about when a moment you know exists never appears, so we would rather publish the number than have you discover it.
The model reads timecodes, never a waveform
Before the transcript reaches the picker we rewrite it with a position marker roughly every ten seconds, so what the model reads looks like a script with running timecodes rather than an undifferentiated wall of words. That is the reason returned start and end points land where they do — the model is matching a moment to the nearest marker it can see, not measuring anything in the sound. Loudness, silence and laughter never reach it on an uploaded file. Energy analysis is real and it is in the stack, but it belongs to the live path, which has to commit to a moment during a broadcast before a finished transcript exists to read.
Widening a clip to the nearest safe point was a mistake
An early version of the boundary logic fixed endings that landed mid-sentence by pushing each clip outward to the next safe stopping point. It worked in the narrow sense that endings stopped being severed, and it also added more than 30 seconds to the average clip. Half a minute of drift is not a rounding error on a forty-second cut, it is a different cut with a different pace. What replaced it runs the other way round: the segment is chosen along sentence structure to begin with, so the ending is already correct before any optimiser is allowed near it and nothing needs padding to rescue it.
Compared with the other ways to find moments
vs. hunting through the timeline by hand
You will find good moments this way — you will just find them slowly, and only in recordings you can face re-watching. The cost is not the hour it takes on one video. It is the eleven videos you never open because you know what opening them involves.
vs. auto-generated chapters
Chapters segment a video by topic, which is a genuinely different job. A topic boundary tells you where the conversation moved on; it says nothing about whether anything in that stretch was worth watching. Chapters help a viewer navigate. They do not rank.
vs. searching the transcript
Transcript search is precise and blind. It will find every instance of a phrase you already know to look for, which means it can only surface moments you had already half-remembered. The ones you completely forgot — usually the best ones — are invisible to a keyword.
vs. reading your comments and chat replays
Audience reaction is real signal and worth mining. The catch is that it only exists for content people already watched, and it clusters around drama rather than usefulness. Treat it as a second opinion on the shortlist, not as the shortlist.
vs. rival AI clipping products
Plenty of tools cut clips. Fewer treat the ranking as the deliverable, and almost none run detection against a stream that is still live. If detection quality is what you are shopping for, the fair test is to run a recording whose best moment you already know and see whether it comes back near the top — do that here.
Frequently asked questions
How does the AI decide what a "best moment" is?
It reads the timestamped transcript and looks for segments with internal structure: a setup that resolves, a specific claim, a change in direction, or an answer that lands. Then it checks whether the segment still makes sense with everything around it removed, because that is the condition a clip is actually watched under. Segments that fail that test rank low even when they sounded lively in context.
Does it just look for loud parts or laughter?
No, and on an uploaded recording it could not if it wanted to. Audio-energy detection is cheap to build and it is why so many highlight tools return montages of people laughing at nothing. The picker for an uploaded file receives the transcript and nothing else, so loudness is not an input at that stage at all. Energy analysis does exist in our stack, but it runs on the live path, where a moment has to be chosen mid-broadcast before any finished transcript exists. Plenty of the highest-scoring finds are completely flat on a waveform.
How many moments will it find in an hour of footage?
It varies by how dense the recording is, and the number is not padded to hit a target. A focused hour-long interview usually yields somewhere in the high single digits to the mid teens. A rambling hour might yield three. If the material is thin, a short list is the accurate answer rather than a failure.
Do I get timestamps, or just the finished clips?
Both. Every moment shows its position in the source video alongside the rendered clip. If you would rather do the craft work in your own editor, you can ignore the exports entirely and use this purely as a search layer, jumping straight to the timecodes it returned.
Is there a free way to test the detection?
Yes. There is a free demo for videos under thirty minutes, which exists precisely so you can check the moment selection against your own judgement before paying anything. After that, full access is a 3-day trial for $1 and then $29 a month, cancellable in one click.
How is the 0-100 rank calculated?
It grades how strongly the clip opens, how efficiently it reaches its point, and where its peak falls within the runtime. It is calibrated for ranking clips against each other rather than predicting view counts. Used as a sort order it is reliable; used as a forecast it is not, and nobody honest would sell it as one.
Can it find moments in a video I did not make?
The tool accepts any URL it can reach, so mechanically yes. What you are allowed to publish afterwards depends on copyright and the rules of the platform you post to, and that judgement belongs to you rather than to us.
Does it work on a live stream?
Yes, and this is the part most alternatives cannot do. Connect a YouTube, Twitch or Kick channel and detection runs while the broadcast is happening, so moments surface during the stream instead of after you have downloaded and processed a huge VOD.
Is there a limit on how long the recording can be?
Two hours per upload on a paid plan, and the free demo caps at thirty minutes. There is a second limit worth knowing about on the long end: one detection pass reads roughly the first half-hour of speech, so a full two-hour file is read most closely at the front. Sending it as three shorter uploads costs you nothing extra and buys even coverage. For a marathon stream, connecting the channel and letting live detection run is a better route than trying to process the whole VOD in one piece.
Will it find the moment I already have in mind?
Usually, and that is the test worth running first. Take a recording where you know exactly which segment is the good one, run it, and see where that segment lands in the ranking. If it comes back near the top, the detection matches your taste well enough to trust on recordings you have not reviewed.
What if it misses something obvious?
Two things usually cause it. Either the moment depends on context outside the clip, which detection intentionally penalises, or it is a community reference the model has no way to weight. Scan the lower-scoring finds before discarding them, since near-misses often sit at rank eight rather than being absent entirely.
Can I change the length of the moments it returns?
The cut length is chosen to fit the thought rather than a fixed duration, so a complete idea is not chopped to hit forty seconds. You can adjust the in and out points on any clip afterwards and re-render if you want it tighter or want an extra beat of reaction at the end.
Does it remove pauses and filler inside a moment?
Yes, automatically. Ums, false starts and silences are cut out of the segment, which usually recovers several seconds and noticeably tightens the pacing. A good moment with three seconds of dead air in the middle of it stops being a good moment.
Are the found moments captioned?
Every one, with word-by-word animated captions burned into the video in your choice of eleven styles. Since most short-form viewing happens with the sound off, an uncaptioned moment is effectively an undiscovered one.
Can I use this on Zoom or Teams recordings?
Yes, if you export the recording as a video file and upload it. Multi-person grid recordings are handled with the split and grid layouts so participants stay legible after the crop to vertical.
Does the finder work for non-English recordings?
Transcription covers major languages and detection runs on the transcript, so it functions beyond English. Quality is strongest in English and does vary with language and audio conditions. Running one representative recording through the free demo will tell you more than any claim on this page.
Can I search my past moments later?
Yes. Everything you process stays in your library with its score, title and source attached, so the clip you cut three months ago is findable when it suddenly becomes relevant again.
What aspect ratios do the moments export in?
Vertical 9:16 by default for Shorts, TikTok and Reels, plus 1:1 for square feeds and 16:9 if you want the moment as a landscape cut for a newsletter or a site embed.
Do the exported moments carry a watermark?
Clips produced by the free demo carry a small badge. Paid exports are clean, and you can apply your own logo through the Brand Kit if you want your mark on everything you publish.
Can I post the moments straight to social?
Yes. Connect TikTok, Instagram and YouTube and either publish immediately or queue moments on a schedule, which is generally the better use of a batch of finds than dumping all of them out on one day.
Is there an API for moment detection?
Yes, the same detection engine is exposed through a developer API, and there is an MCP connector so Claude can run a search conversationally. Both are documented in the developer docs if you want detection inside your own pipeline.
How long does a search take?
Usually a handful of minutes, stretching out for a two-hour source or when a lot of jobs are running at once. You are not required to watch it run — close the tab and the shortlist is there when you come back.
Will it work on a video with poor audio?
Detection depends entirely on the transcript, so bad audio degrades it in proportion. Heavy background noise, overlapping crosstalk and very quiet mics all cost accuracy. If the audio is bad enough that you struggle to follow it yourself, expect the shortlist to reflect that.
Is a moment finder useful if I already have an editor?
Often more useful, not less. Editors are expensive to point at raw footage because reviewing is the slow part of their day. Handing an editor a scored shortlist with timecodes moves their hours from searching to actually making the clips good.
What do you do with my footage once the search finishes?
It is used to produce your results and is not published anywhere by us. The clips and the moment list belong to you. The privacy policy sets out how the files and data are handled.
Can I stop paying once the trial ends?
Yes, in a single click from your account, and we send an email before the trial converts so it never turns into a surprise charge. Whatever you have already paid for stays usable right up to the date it was due to renew.
What are the alternatives to an automated moment finder?
Scrubbing the timeline yourself, which is accurate and does not scale past the recordings you can face re-watching. Searching the transcript for phrases you remember, which only ever returns moments you had already half-recalled. Reading comments and chat replays, which is real audience signal but exists only for content people already watched. Or paying an editor or assistant to review the footage, which works well and turns into a per-hour cost that grows with every recording you make. Each is a reasonable answer depending on how much footage is piling up.
How is this different from a general clipping tool?
A clipping tool is judged on its output — the caption style, the crop, the export. A moment finder is judged on its choices. If you want to see the same engine framed around finished vertical output instead, the YouTube Shorts maker is the same detection with a publishing emphasis.
Point it at the recording you have been avoiding
Pick the long one you never got around to reviewing. One pass will tell you whether there were four moments in it or none.