HomeAI Video ToolsAI Video Summarizer
AI Video Summarizer

An hour of video, reduced to the parts that carry it

This summarises by extraction rather than paraphrase. The model works down the transcript and pulls out the passages that hold the argument — each one a short, self-contained clip with a title and a rank — so you can absorb a long recording in the time it takes to drink a coffee, in the speaker's own words.

🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
  • Extracts watchable moments, not a wall of text
  • The speaker's own words, never paraphrased
  • Each one titled and ranked
  • Cut on complete thoughts
  • Works on a live broadcast too
One long video in, a week of posts out

Paste a link. Get the best moments.

Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.

https://youtu.be/z_bX3runikk
Source video: Matthew McConaughey's Concerns Over AI Source video 15:27
Matthew McConaughey's Concerns Over AI
9 clips found  ↓  top 5
Matthew McConaughey speaking, captioned vertical clip
Score 92
9:16
Joe Rogan gesturing, captioned vertical clip
Score 90
9:16
Matthew McConaughey listening, captioned vertical clip
Score 89
9:16
Joe Rogan mid-sentence, captioned vertical clip
Score 89
9:16
Matthew McConaughey talking to camera, captioned vertical clip
Score 88
9:16

Summarising a long recording

You supply the link. Everything after that is the tool deciding what mattered.

  1. 1

    Give it the recording

    Paste a YouTube URL or upload a file — a conference keynote, a two-hour interview, a lecture, a recorded meeting, a product deep dive. Paid plans accept up to two hours per upload. The free demo covers sources under thirty minutes, which is enough to see whether its judgement lines up with yours.

  2. 2

    It reads before it decides

    The audio is transcribed, and a selection pass then works down that transcript marking the passages that carry weight: the claim with evidence behind it, the answer that resolves a question, the concession, the number, the point the speaker returned to twice. Nothing is chosen for sitting near a marker or a timestamp somebody added afterwards. On a recording longer than about half an hour, submit it in halves — the reason is spelled out below, and it is the one thing worth knowing before you rely on the output.

  3. 3

    Skim the moments instead of the video

    What comes back is a short, ordered set of clips with titles and 0-100 ranks. Read the titles to see the shape of the recording, watch the three that matter to you, and ignore the rest. If a moment was cut tighter than you needed, widen its start or end point and it re-renders with the surrounding context.

What the summariser gives you

The point is not to produce fewer words. It is to spend your attention on the eight minutes that were worth an hour.

Continuous speech, not chapter hopping

People skim video by dragging the scrubber, which surfaces whatever happens to sit beside a chapter marker and quietly discards everything between them. The selection pass reads continuous transcript instead, so a passage earns its place on what it actually says rather than on whether somebody put a label near it. Chapter titles are written to divide a recording, not to grade it, and those are separate jobs.

Moments you can watch, not a paraphrase

The summary is made of the actual footage. You hear the pause before the answer, the hedge in the phrasing, the emphasis on one particular word. Those are things a written précis flattens, and they are frequently the reason you were watching a video rather than reading a document.

Each moment is a complete thought

Cuts land on sentence boundaries, so a clip opens where an idea starts and closes where it finishes. An extract that begins mid-clause is useless as a summary because you have to go back to the source to work out what was being referred to.

Titles that work as an index

Every moment arrives with a short suggested title, editable if it misses. Read down the list and you have the argument of the recording laid out in order without playing anything, which is often all you needed.

Ranked so you know where to start

A 0-100 grade on each moment orders the list by how strongly it lands. When a keynote yields fourteen extracts, the ranking is what turns them from an undifferentiated pile into a sequence you can work down and abandon halfway.

The waffle is already removed

Filler words, false starts and long pauses are stripped from every extract. A summary should be dense by definition, and several seconds of throat-clearing at the top of each moment defeats the purpose of summarising at all.

Readable with the sound off

Word-by-word captions are burned into every extract, which means you can consume a summary in an open-plan office, on a train, or at double speed with the audio off. For a lot of people this is the difference between reviewing a recording and never getting to it.

A summary other people will actually open

Sending a colleague a two-hour link is a request they will decline. Sending three ninety-second clips with titles is something they will watch in the lift. The output is designed to be forwarded, which is where most video summaries are actually consumed.

Summarise a broadcast while it is still running

Connect a YouTube, Twitch or Kick channel and the key moments are extracted during the stream rather than after it. For a long live event, that means a running digest exists before the broadcast has even finished.

Let Claude write the prose part

ClipSpeedAI runs as an MCP connector, so you can ask Claude to pull the key moments from a link and then have Claude write the narrative summary around what came back. That combination — extraction from us, prose from the model — is the closest thing to a complete summary of a video that currently exists. See the developer docs.

Who summarises video, and why

Students and researchers

Recorded lectures and conference sessions pile up faster than anyone watches them. Extracting the defensible claims lets you decide in five minutes whether the full ninety are worth your evening.

Sales and customer teams

Recorded calls hold objections, competitor mentions and the exact phrasing a buyer used. Reviewing them at volume is impossible; reviewing the moments that mattered from each one is not.

Journalists and analysts

A long interview or an earnings call has three quotable passages inside it. Finding them by scrubbing is the slowest part of the job, and having them as clips means they are ready to publish rather than ready to be edited.

Anyone with a meeting recording backlog

The recording exists, nobody rewatches it, and the decisions inside it get relitigated a month later. A short set of extracted moments is a record people will actually consult.

Podcast listeners and producers

Deciding whether a two-hour episode deserves your commute is a real question, and the moment list answers it. Producers use the same output as show notes material — see the podcast clip generator.

Trainers with long lesson recordings

A long lesson usually contains three self-contained explanations. Pulling them out serves two purposes at once: a summary for existing students and a free sample for prospective ones.

Two kinds of summary, and which one this is

Summarisation splits into two approaches. Abstractive summarisation reads the source and writes something new about it, in different words — that is what you get when you paste a transcript into a chat model and ask for the gist. Extractive summarisation selects the most important parts of the original and hands them back unchanged. This tool is firmly the second kind, and the distinction is not academic, because the two fail in opposite directions.

A written abstract is compact and searchable, and it is also a report from a system that may have quietly smoothed over a hedge, dropped a caveat, or stated something with more confidence than the speaker did. If you are going to act on what was said — quote it, forward it, base a decision on it — you generally end up going back to the source to check, which means the summary saved you less than it appeared to.

An extractive summary cannot misrepresent the source in that way, because it is the source. What it cannot do is compress an argument that was made slowly across twenty minutes into three sentences. So the useful rule is about what you need next. If you need to know what was said and be sure of it, extraction. If you need something short to paste into a document, paraphrase.

A workflow for getting the written takeaways as well

We are not going to pretend this outputs a research memo. There is no essay at the end of the process and no transcript file to download — captions are burned into the clips rather than delivered as a sidecar document, which is right for posting and wrong for note-taking. Stating that plainly is more useful than implying a feature we do not ship.

The written layer you do get is the titles. Each moment carries a short editable headline, and read together as a list they function as an outline of the recording: the claims made, in the order they were made. For most review tasks that outline plus three watched clips is the entire job.

When you need real prose, the MCP connector is the answer and it is genuinely good. Ask Claude for the key moments from a video, and it can call ClipSpeedAI, receive the extracted passages, and then write the summary, the bullet points, the action items or the email around them. The extraction stays grounded in what was actually said and the writing is done by a model built for writing, which is a better division of labour than one system attempting both.

A checklist for summarising a recording over an hour long

Split it before you submit it. The pass that chooses moments starts at the beginning of the transcript and has a working limit on how much it takes in, so on a feature-length source the closing stretch can sit outside what it ever considers. Two submissions of an hour each are each handled properly, and you end up with two moment lists to read in sequence rather than one list that quietly thins out toward the end.

Decide what you are looking for before you open the list. "What was decided" and "what is quotable" pull toward different passages, and skimming a ranked set with no question in mind is how people watch six extracts and retain none of them. The titles are there so you can settle that in about thirty seconds.

Widen anything that reads as an assertion with nothing under it. Extraction errs toward brevity, so the sentence that set a claim up is the one most likely to have been left just outside the cut. Pulling the start point back a few seconds and re-rendering usually recovers it, which is cheaper than reopening the original to find out what the speaker was responding to.

Myths about AI video summaries that cost people time

The first is that summarising is a compression ratio — that two hours ought to reduce to a fixed ten minutes whatever was in them. Density belongs to the recording, not to the tool. A disciplined half-hour talk can be nearly all signal and reduce badly, while a loose two-hour conversation might hold four minutes worth keeping. Anything that returns the same proportion every time is padding or discarding to hit a number.

The second is that the top-ranked moment is the one you came for. The rank measures how strongly a passage stands on its own, which overlaps with what you need without being the same thing. The quiet sentence where somebody concedes a point almost never outranks the confident claim, and it is often the more informative of the pair. Read the titles before you trust the order.

The third is that a summary can stand in for the recording whenever detail matters. It cannot, and no amount of extraction changes that. What it does do is tell you cheaply whether the detail matters enough to go back for, which is a smaller question and a much faster one to answer.

Troubleshooting a moment list that missed what you were after

A summary disappoints in four fairly distinct ways, and telling them apart is quicker than re-running the job. Three of the four are fixable from the results screen in under a minute. The fourth is a property of the recording and no setting will move it.

The usual complaint is that a passage you distinctly remember is simply absent. Check where it fell in the runtime before anything else. If it sat in the closing third of a feature-length source, the likelihood is that it was never considered rather than considered and passed over, and resubmitting that stretch as its own job settles the question in a couple of minutes. If it sat early and still did not appear, it lost to other material, and widening the neighbouring extract usually brings it back into frame.

The second is an extract that lands as a firmer claim than the speaker actually made. Selection cuts to the sentence that carries, and the qualifying clause frequently lives in the sentence before it. Drag the start point back by one sentence and re-render before you forward that clip to anyone. A hedge that was present in the room and missing from the extract is the single editing error in this workflow with consequences outside your own head.

The third is captions containing the wrong proper nouns. Names, product names and domain jargon are where transcription degrades first, and burned-in text is not something a viewer can mentally skip past. Caption text is editable before a clip renders, so correcting the three words that matter costs seconds — noticing after export costs a re-render instead.

The fourth is a short list from a long recording, and that is usually the recording being honest rather than the tool being lazy. A two-hour conversation that circles the same three ideas contains three moments, not twenty. Worth knowing that this is the case you cannot configure your way out of, because padding the count would hand you weaker material carrying exactly the same confident rank as the good material.

Skimming a recording by hand versus letting a pass read all of it

Manual skimming is not reading a recording, it is sampling one. You drag the scrubber, land somewhere arbitrary, hear half a sentence, drag again. The method systematically over-weights the opening — everybody watches the first five minutes properly — and it finds material at chapter markers and topic changes because those are the places a waveform or a thumbnail strip gives you something to aim at. Whether a passage was any good is invisible to all of those cues.

An automatic pass has the opposite shape. It reads every sentence with equal attention, which is exactly what a human skimmer cannot do, and it has no idea what you personally came for, which is exactly what a human skimmer knows perfectly. Say you sit down with a 45-minute internal review already knowing you need the part where pricing came up: searching a transcript for the word "pricing" beats any ranked list, and it always will. Known target, manual wins.

The moment you do not know what you are looking for, the comparison inverts. An unfamiliar two-hour keynote gives you nothing to search for, and sampling it by scrubber means judging a talk on the four random seconds you happened to land in. That is where reading a ranked set of titles first is not a shortcut but a genuinely better method, because it puts the shape of the whole recording in front of you before you spend attention on any part of it.

The workflow most people settle into after a week uses both. Run the pass, read the titles as an index, watch two or three extracts, and if a specific detail is still missing, go back to the original with a timestamp you now actually have. Nothing about the automatic pass prevents you from opening the source; it just means you open it knowing where to land.

Alternatives worth considering before you automate this

A plain transcript with a search box is the strongest alternative and it is frequently the correct one. If your question is factual and you can name the word that would appear in the answer, a transcription service costs less, returns a searchable document, and answers you faster than any selection pass. It fails on the question people more often have, which is not "where was X mentioned" but "was any of this worth an hour".

Meeting note-takers occupy the neighbouring problem and solve it well. They produce written minutes, decisions and action items in text, filed and searchable, which is what an internal record needs to be. Use one for the record and something extractive for anything you intend to quote, forward or verify — the two outputs are not competing so much as answering different questions about the same call.

Show notes, chapter lists and a colleague who already watched it are all free and all underrated. If somebody on your team sat through the webinar, ask them which ten minutes mattered; a human who was present carries context no transcript contains. The reason automation is worth reaching for is volume, not superiority — nobody has a colleague available for the fortieth recorded call of the quarter.

Then there are other clipping and summarising services, most of which now cover uploaded files competently. Two things to test them on if you are comparing: what they do with the back half of a long recording, since that is where selection quietly thins out, and whether they can work from a broadcast still in progress. If you regularly need a digest before an event has finished, the livestream side of the engine is the axis that separates tools rather than the summary formatting.

Export settings, clip length, and where the files go next

Every moment comes back as a standard MP4 with the words burned into the picture. The default frame is 1080 × 1920, and 1:1 and 16:9 are available on the same job when the extract is destined for a document, a slide or an internal wiki rather than a feed. For a summary you plan to forward to colleagues, 16:9 usually reads better on a laptop than a tall frame does — the vertical default exists because most extracts end up as social clips, not because it is right for every use.

Clip length is not a setting and that is deliberate. Duration follows where the thought closes, so extracts from the same recording vary, and a set that all came back at a suspiciously round number would be a sign of a tool trimming to a target rather than to a sentence. If a particular extract is longer than you want to send, move its end point back to an earlier sentence break rather than cutting to a duration.

Caption style is worth setting differently for review work than for publishing. The bold animated styles that hold attention in a feed are noise in a clip somebody is watching to extract a fact, and a plain, high-contrast style is the better default for internal circulation. There are eleven of them, and the choice applies per clip, so one recording can produce internal extracts and public ones from the same pass.

For anything internal, turn the Brand Kit off. A logo and brand colours on a clip you are sending to two colleagues is furniture nobody asked for, and the same footage can be re-rendered branded later if it turns out to be worth publishing. The files themselves are unremarkable video — they attach to an email, drop into a channel, sit in a shared drive, and play without anyone installing anything.

Where this is not the right tool, and the mistakes that follow

If you need a verbatim record — legal review, compliance, accessibility, or study notes you will revise from — you want a transcription service, not this. Burned-in captions on a set of clips are not a document, and using extracts where a transcript is required will cost you more time than it saves.

If the video is under about ten minutes, just watch it. The overhead of processing and reviewing a moment list is not worth it against pressing play, and any tool claiming otherwise is selling you a workflow rather than a result.

And if the value of the recording is visual rather than spoken — a design walkthrough, a dense slide deck, a screen recording where the explanation is happening on the screen instead of in the audio — the extraction has weaker signal to work with, because it is reasoning primarily about what was said. It will still find the moments where someone explains something out loud. It will not summarise a diagram nobody described.

What we have learned running a summariser in production

Observations from operating the pipeline in production — not general advice.

The transcriber reaches further than the selector does

Every second of audio gets transcribed, but the pass that decides which passages matter is handed a bounded slice of that transcript — about 28,000 characters, which works out somewhere near the half-hour mark depending on how quickly the speaker talks. Past that point the tail of a long recording is not losing to better material, it is absent from the decision entirely. We would rather write that down than let a reader infer coverage we do not deliver, and the workaround is dull but effective: send a feature-length recording in two halves and read the two moment lists together.

A summary is a spread, not the top scores

The obvious way to build an extractive summariser is to grade every window and return the highest numbers, which in practice hands you five takes on the same strong five minutes and nothing else. Our selection step forbids exactly that: extracts may not overlap, each has to begin after the previous one ends, and each is required to cover a different moment. What comes back is therefore shaped like the recording rather than clustered around its loudest passage, and it is the reason the moment ranked fourth is frequently the one you needed.

Ending on a complete thought is harder than starting on one

Openings are recoverable because there is always an earlier sentence break to fall back to. Endings are not. A thought that runs past the length cap has no lawful place to stop, and an extract with no speech in it has no sentence to end on at all. Both cases are real, and we let them fail open rather than chop mid-word. It is also why boundary work runs after the moment has been picked: choosing what to include and choosing where the sentence closes are separate problems, and fusing them is how extracts end up opening mid-clause.

Compared with the other ways to get through a long video

vs. watching at double speed

Speeding up a recording still costs you half its runtime, and comprehension drops on anything dense enough to have been worth summarising. It also does not solve the actual problem, which is that most of a long video is transition, restatement and setup rather than substance.

vs. an AI text summary of the transcript

Text summaries are faster to read and easier to file, and they are the right output for a lot of tasks. Their weakness is verification: you are trusting a paraphrase of something you have not heard. Pairing the two covers both — the extracted clips are the evidence, the written version is the convenience.

vs. auto-generated chapters and platform summaries

Chapters divide a recording by topic, which tells you where things are but not which of them was worth your time. A chapter marked "Q&A" could contain the most interesting ninety seconds of the whole session or twenty minutes of housekeeping, and nothing about the marker distinguishes those cases.

vs. skimming the transcript yourself

Reading a transcript is genuinely effective and genuinely slow, and an hour of speech is a great deal of text with no formatting to guide your eye. It is also hard to tell from a page of prose which line landed in the room, which is exactly the signal an extractive pass is trying to capture.

Frequently asked questions

How does the AI summarise a video?
The recording is transcribed, and a selection pass then reads down that transcript picking the passages that carry the most weight — claims with support behind them, resolved questions, points the speaker emphasised or repeated. Each one becomes a short clip with a title and a 0-100 rank, cut so it opens and closes on complete sentences.
Does it write a text summary or paragraph?
No. Each moment gets a short editable title, and those titles read as an outline of the recording, but there is no written abstract at the end. If you want narrative prose, use the MCP connector and let Claude write it from the extracted moments — that combination works well and keeps the writing anchored to what was actually said.
Can I download the transcript?
Not as a separate file. Captions are burned into the clips rather than exported as a subtitle document, which is the right choice for posting and the wrong one for note-taking. If a transcript file is your actual requirement, a dedicated transcription service will serve you better than this will.
How long is the summary?
It depends on the source, because the output is however many moments genuinely earned a place rather than a fixed count. A tight thirty-minute talk might reduce to four extracts running six or seven minutes in total. A rambling two-hour interview can produce more, and the ranking tells you where to stop.
Will it find the important part if it happens near the end?
On a source of half an hour or less, yes — nothing about being late in the recording counts against a passage, which already beats manual skimming because that reliably over-samples the first ten minutes. On a much longer recording the honest answer is that the selection pass may not reach the end at all, so split the file and submit the closing stretch as its own job.
Can I summarise a YouTube video by pasting the link?
Yes, and that is the fastest path since nothing needs downloading to your machine. Direct file upload is there for recordings that were never published — internal meetings, client calls, footage you would rather not put on a platform.
How long a recording can it handle?
Paid uploads run to two hours each, and the free demo covers sources shorter than half an hour. Beyond two hours you would split the recording, or connect a channel and let the live mode extract moments while the broadcast is still running.
Can I test it without paying?
The demo is free for videos under thirty minutes, which is deliberately enough to test its judgement on a recording you already know well. Beyond that the trial is $1 for three days, then $29 a month for Pro, stoppable in a single click with an email warning beforehand.
Does it work on recorded meetings?
Yes, and meeting recordings are one of the better fits, since the ratio of substance to scheduling talk is usually poor and the substance is what you need. Upload the file from your meeting tool. Multiple speakers are handled, and the split layout keeps two participants visible if you plan to share the clips.
Can it summarise a lecture?
It works well for lectures where the teaching happens verbally. For a maths lecture where the content is being worked out on a board with minimal narration, the extraction has much less to reason about, and you would be better served by the recording itself plus your own notes.
How accurate is the moment selection?
Good enough that most people find the shortlist overlaps heavily with what they would have chosen, and not infallible — it occasionally over-values a punchy line with little behind it. The honest way to evaluate it is to run a video whose content you already know and see how the list compares to your own memory of it.
Does it handle multiple speakers?
Yes. Speaker changes are tracked and factored into where clips begin and end, so an extract usually contains a full exchange rather than half an answer. For interview footage the split layout keeps both faces on screen through the moment.
What languages does it work in?
Transcription covers major languages with English the strongest by a clear margin. Selection quality follows transcription quality, so noisy audio or heavy crosstalk degrades both. Running one representative video through the free demo will tell you more about your specific case than any general claim would.
Can I make a moment longer if it cut too tight?
Yes. Adjust the start and end points and re-render to pull in the surrounding context. This is worth knowing because summarisation errs toward brevity, and occasionally the setup sentence before a claim is the part that makes it make sense.
Does it summarise the visual content or just the speech?
The selection reasons mainly about what is said. Visual context informs how the clip is framed — who is speaking, whether a screen is being shared — but a chart nobody talks about will not be surfaced as a key moment. Videos where the substance is on screen and unspoken are the weak case.
Is this the same thing as a note-taking assistant?
No. Meeting note-takers produce written minutes and action items in text. This produces the moments themselves as watchable clips. They serve different needs, and a fair number of people use one for internal records and this for anything they intend to share or verify.
Can I share the summary with someone else?
Yes, and the shareability is a large part of the point. Clips are ordinary video files, so they go into a message, a document, an email or a channel and get watched, which is more than can be said for a link to the two-hour original.
Do the extracted moments carry a watermark?
Only on the free demo, and it is unobtrusive. Paid exports have nothing added, so an extract you forward to a client or a colleague looks like a plain video file.
Can it summarise a live stream while it is happening?
Yes. Connect a YouTube, Twitch or Kick channel and moments are extracted during the broadcast, which is useful for long live events where a digest is more valuable during the event than after it.
How quickly does a summary come back?
A few minutes for most recordings, scaling with length and queue depth. Processing runs on our servers, so closing the tab is fine and the results are waiting when you return.
Can summaries be produced automatically in a pipeline?
Yes. The extraction engine is exposed through the developer API, which makes it practical to summarise recordings as they land in storage rather than submitting each one by hand through the browser.
Is my video kept or used for anything else?
It is processed to produce your output and nothing further. Nothing is published by us and the results are yours. For recordings containing anything confidential, read the privacy policy first and make your own call.
Does the AI generate any content that was not in the video?
No footage and no voice. Everything you watch came out of your source recording. The only generated text is the suggested title on each moment, and you can rewrite any of them.
Can I use the extracted moments as social clips?
That is what they are built as, since a good extract and a good short clip turn out to be nearly the same object. They come back vertical, captioned and ready to publish, and the Instagram Reel maker covers that side of the workflow.
What if I disagree with what it picked?
Delete the ones you do not want, extend the ones that were cut short, and reorder the rest. Treat the output as a first pass by a fast assistant rather than a verdict, which is the correct expectation for any automated selection.
Can I read the summary on a phone?
Yes. Submitting a link and reading through the moment list both work on a handset, which suits the common case of wanting to know what was in a recording before you are back at a desk.
What is the cancellation process?
One click from your account, and an email lands before the trial rolls into a subscription. Stopping it leaves your access in place until the period you already paid for runs out.

Summarise something you have already watched

Run a recording you know well and compare its shortlist to your own memory of what mattered. That is the only test of a summariser that means anything.

⚡ Summarise My Video — $1 trial
3-day trial · just $1 · cancel anytime