HomeAI Video ToolsAI Video Editor
AI Video Editor

An AI video editor that makes the decisions

Ordinary editors hand you a timeline and wait for instructions. This one watches the footage first, decides which stretches deserve to survive, trims them on clean sentence boundaries, keeps the speaker centred when it crops to vertical, burns in the captions, and then grades every result 0-100 so even the posting order is settled before you look at it.

🎬 Free demo: use a video under 30 minutes · paid plans up to 2 hours
  • No timeline to learn
  • Every cut graded 0-100
  • Captions burned in, word by word
  • Speaker-aware vertical reframe
  • Edits a live stream while it airs
One long video in, a week of posts out

Paste a link. Get the best moments.

Real ClipSpeedAI output — every clip below was scored, cut, reframed to 9:16 and captioned automatically from the video above it.

https://youtu.be/z_bX3runikk
Source video: Matthew McConaughey's Concerns Over AI Source video 15:27
Matthew McConaughey's Concerns Over AI
9 clips found  ↓  top 5
Matthew McConaughey speaking, captioned vertical clip
Score 92
9:16
Joe Rogan gesturing, captioned vertical clip
Score 90
9:16
Matthew McConaughey listening, captioned vertical clip
Score 89
9:16
Joe Rogan mid-sentence, captioned vertical clip
Score 89
9:16
Matthew McConaughey talking to camera, captioned vertical clip
Score 88
9:16

How the automatic editing pass runs

You provide footage at the start and taste at the end. Everything between is handled.

  1. 1

    Hand it the footage

    Paste a YouTube link or send a file up from disk. There is no project to create, no bin to organise, no proxy pass, and no sequence-settings dialog to get wrong. Uploads on a paid plan may run a full two hours, while the no-cost demo stops at the thirty-minute mark.

  2. 2

    It looks before it cuts

    The audio is transcribed, the transcript is read for structure, and the picture is examined to work out how many people are present and whether a screen, a game or a slide is in shot. Only after that does anything get trimmed. Doing it in that order is why the boundaries land where a thought opens instead of at a tidy round number of seconds.

  3. 3

    Approve, adjust, or simply post

    Finished cuts arrive vertical, captioned and graded. Drag the handles a second either way, swap the caption style, change the layout, rewrite the title, re-render. Or take the three highest scores and publish them. Most sessions end up being a review rather than an edit, which is the whole intent.

The editing jobs it takes off you

Every item here is something a human editor repeats identically on clip after clip. That repetition is the part worth handing to software. Judgement is not, which is why the last call stays yours.

Moment selection, which is the real bottleneck

Dragging clips around is not what makes short-form expensive. Deciding is. The model reads the full transcript alongside the audio and marks passages that survive without setup — the ones a stranger can drop into cold. You end up judging a shortlist instead of hunting through raw footage with a scrub bar.

A 0-100 grade on every cut

Each finished clip carries a number reflecting how hard its first seconds work, how evenly it moves, and where its strongest beat falls relative to its length. When twelve come back at once, that number is the difference between a folder of files and an actual posting plan for the week.

Boundaries that respect sentences

A cut that begins three syllables into a word identifies itself as machine output before the content has a chance. Trims are placed where a sentence starts and where one resolves, so the clip opens on a complete thought and lands on one.

Reframing that follows whoever is talking

Taking a wide 16:9 master down to 9:16 throws away roughly two thirds of the frame width, and a fixed centre crop will happily settle on the lamp between two chairs. Tracking the active face keeps the subject where an eye expects to find them, and that alone is why a vertical looks edited rather than merely cropped.

Seven layouts, picked from what is on screen

Fill, fit, split, three-up, four-up, screenshare, and a gameplay picture-in-picture. A slide deck wants the slide legible with the presenter inset; a four-person panel wants a grid; a piece to camera wants neither. Choosing per clip is itself an editing decision, and it is made from the frame rather than from a preference you set once and forget.

Captions timed to the word, painted into the pixels

Words illuminate as they are spoken, in eleven styles. They are rendered into the video rather than delivered alongside it, so nothing needs re-syncing after a trim and nothing falls off when the file gets reposted somewhere that ignores sidecar subtitle tracks.

Dead air and verbal tics removed unprompted

Pauses, restarts and the small noises people make while thinking are stripped without being asked. Over a minute-long cut this commonly reclaims several seconds, and reclaimed seconds are the entire distance between a clip that feels tight and one that feels merely acceptable.

It can edit while the source is still recording

Attach a YouTube, Twitch or Kick channel and cuts are produced mid-broadcast. A timeline editor structurally cannot do this, because until the stream stops there is no file to open. Being first to a moment is worth more than being polished about it two hours later.

Drivable from Claude or from your own code

One engine sits behind both an MCP connector and a REST endpoint, so an edit can be requested inside a conversation or fired by a cron job in your stack. The API reference lists what is exposed.

Who swaps which part of their workflow

Podcast teams

The weekly episode is already produced. What never happens is the twelve vertical cuts, because that is a second job nobody was hired for. The podcast clip generator covers that case.

B2B marketing

Recorded demos and customer sessions hold the most persuasive footage a company owns, and virtually none of it is ever cut down. See the webinar clip generator.

Solo creators

One person shooting, editing, thumbnailing and posting runs out of evenings long before running out of ideas. Automating the mechanical third makes a publishing schedule survivable.

Newsrooms and media desks

A press conference or a long sit-down needs a quotable forty seconds out while the story is still moving. Turnaround beats polish for that job, every time.

Course and curriculum creators

A lesson usually contains two or three explanations that stand alone perfectly well. Those are the pieces that pull strangers back toward the paid material.

Faith and community teams

Volunteer media crews edit on borrowed weeknights with whatever laptop is available. The sermon clip generator is written for exactly that constraint.

Tooling was never the shortage

Video editing software got extremely good and extremely cheap somewhere around a decade ago. Anyone can install something capable of broadcast work this afternoon for nothing. Despite that, most people with a long recording still publish zero clips from it. The missing ingredient was obviously never the software.

What is actually missing is a sequence of small judgements repeated dozens of times: where does this thought begin, does it stand up without the ten minutes before it, is the payoff early enough, does the crop hold the right person, do the captions read at arm’s length on a phone. None of those are hard individually. All of them together, forty times over, is a day of work that competes with the thing you are actually good at.

So the automation worth having is not a faster razor tool. It is a system that arrives at defensible answers to those judgements on its own and hands you a shortlist to overrule. That is the difference between an editor with an AI feature and an editor whose entire premise is the decision.

What this deliberately is not

It is not a general-purpose non-linear editor. There are no keyframes, no colour wheels, no audio bus, no motion graphics, no multicam sync, no titles you can animate by hand. If your job this afternoon is grading a short film or building a lower-third package, install a real NLE and use it. Nothing here competes for that work.

It also does not fabricate footage. There is no text-to-video, no synthetic presenter, no invented B-roll, no cloned voiceover. Everything that comes out was in what you put in. Some people arrive at an AI editor expecting generation and leave disappointed, so it is fairer to say it plainly at the top than to let the discovery happen after payment.

And it will do poorly on some material. Footage with no speech has nothing for a language model to reason about. Heavy background music over quiet dialogue degrades the transcript, and a bad transcript produces bad boundaries and worse captions. Highly visual content whose meaning lives in the picture rather than the words is not the strong case. Clear talking, in any setting, is.

Where automated editing goes wrong

The first way it goes wrong is trusting the score as an oracle rather than a sort order. A high grade means the clip is well constructed; it does not know your audience saw the same joke last month, or that the point contradicts something you published in March. People who post the top-scored clip without watching it eventually publish something they regret, and then blame the model.

The second is feeding it material with no speech worth transcribing. A montage set to music, a silent process shot, a gameplay session with the mic muted — none of these give a language model anything to reason about, and the output reflects that. If the meaning of your video lives in the picture rather than the words, this is the wrong category of tool and no settings change that.

The third is exporting everything. A dozen clips arrive and the instinct is to post all twelve, which trains an audience to scroll past your name. Publishing four and deleting eight is almost always the stronger week, and it costs nothing because you did not make the eight.

The fourth is skipping the check on names and numbers. Transcription is strong on ordinary speech and weakest on proper nouns, so the one clip built around a person's name or a specific figure is the one worth reading the captions on before it leaves.

A weekly workflow that survives a busy month

A realistic session looks like this. You submit a recording, walk away, and return to finished vertical clips with scores attached. You watch the top four at double speed. Two are obviously right and go out. One needs its opening pulled back two seconds because the setup line was trimmed a beat too tight. One gets discarded because the model liked a rhetorical flourish you know landed badly in the room.

That fourth clip matters. Anything claiming to be right every time would be lying, and you would learn not to trust it within a week. Eight decent candidates out of twelve, with the strong ones near the top, is a far more useful thing to own than an imaginary perfect single answer.

The economics hold because failures are free. Throwing away four clips you did not make costs nothing. Throwing away four clips that took twenty minutes each is precisely why people stop clipping after a fortnight.

The cadence that works for most people is one processing run per recording, one twenty-minute review, and clips spaced across the following week rather than dumped in an afternoon. Batch the review with something else dull — do it while a render finishes elsewhere — and the whole habit costs less attention than deciding what to have for lunch.

Automatic editing versus doing it by hand

Hand editing has one advantage nothing automated can match: you were in the room, or you at least sat through the whole recording, so you know which line actually landed and which one merely sounded clever. None of that context reaches a model reading a transcript. Where hand editing loses is throughput and consistency — the fortieth cut of the month is reliably worse than the first, and every editor knows it.

Say you have a ninety-minute recording and you want eight clips out of it. By hand that means scrubbing the full runtime, making eight in-and-out decisions, reframing eight times to vertical, running eight caption passes and queueing eight exports. Call it most of a working day if you are quick and fluent in the software. The automatic pass returns graded candidates in minutes, and the day collapses into a twenty-minute review plus whatever finishing the best two deserve.

The split that holds up is this: hand work wins on anything carrying weight, automation wins on volume. Cut the launch film, the sponsor read and the client deliverable yourself, because those live or die on taste and context. Eight conversational cuts from a Tuesday recording are a different category of work, and treating them as craft is precisely why they never get made at all.

The biggest mistake is framing this as a choice between the two. In practice the arrangement that works is automatic first, hand-finishing second — let the pass settle the shortlist, the boundaries and the framing, then spend your attention on the two clips that earn it instead of the ten that do not.

The jobs where this is not the right instrument

The clearest bad fit is an argument that only works cumulatively. Some talks are a single forty-minute proof where each step depends on the one before it; lift ninety seconds out of the middle and you have removed the reasoning that made it persuasive. No boundary logic repairs that, because the problem is the shape of the content and not the position of the cut.

Recordings that must not be trimmed are the second case. Depositions, compliance training, regulated financial commentary and anything where removing a sentence changes what was said on the record belong published whole or not at all. Software whose entire premise is deciding what to leave out is the wrong instrument for material where leaving things out is itself the risk.

Length is a limitation at both ends. A source of a couple of minutes has nothing to choose between, so you get one clip that is really just the original with captions painted on. Past two hours a single upload is out of scope, and if the source is a broadcast you are better connecting the channel and letting it cut live than waiting for a file that then has to be split.

Worth knowing before you plan around it: a recording in a language your audience does not read is not solved here either. Captions are transcribed from the audio you supply and burned in exactly as spoken. There is no translation step and no dubbing, so a Spanish source yields Spanish captions and the reach problem stays yours.

What running this pipeline actually taught us

Observations from operating the pipeline in production — not general advice.

The better cutting path existed for months with the switch turned off

Audio-verified bad boundaries sat at 71 per cent until cutting moved onto sentence and thought boundaries, which took the figure down to 17. The uncomfortable detail is that the better path was already written — the fix was configuration, not code, and the correct setting was simply off. In practice that generalises well past our own codebase: when an automatic editor produces sloppy cuts, the cause is far more often a default nobody revisited than a capability nobody built.

Stripping filler moves every timestamp behind it

Ums, restarts and dead gaps are removed from inside the clip, which means every word after a removal now sits at a different time than the transcript claims. Burn captions in before that pass and they drift further out of step with each cut you make, which is why the order is fixed rather than configurable: choose the boundary, strip the dead audio, re-align the word timings, then paint the captions into the pixels.

There is no project file, so a re-render is the smallest unit of change

Nothing sits in an editable timeline between runs. A clip is a set of decisions — in-point, out-point, layout, caption style, title — executed into a new file. Nudge a handle by one second and the whole thing is produced again from the source. That sounds wasteful and turns out to be the cheaper model, because it deletes an entire class of failure where a project points at media that has since been renamed, moved or thrown away.

Where it sits against the usual suspects

Not all of these are competitors. Two of them are things you should probably keep using.

vs. Premiere Pro, Final Cut or DaVinci Resolve

Those are craft instruments and this is not trying to replace them. They will execute anything you can specify, at any level of precision, forever. What they will never do is tell you which ninety seconds of a two-hour recording is worth cutting, and that unanswered question is why the footage sits untouched on a drive. Keep the NLE for flagship work; the volume work is a different task.

vs. CapCut

CapCut is a genuinely good mobile-first editor with strong auto-caption and template features, and it is free. The gap is upstream of it: you still have to arrive already knowing which moment you are editing. If you have the clip and want to dress it, CapCut is fine. If you have two hours of raw recording and no idea where to start, the dressing tools are not the problem.

vs. AI features inside a traditional editor

Transcript-based trimming, auto-reframe and speech-to-text all exist inside major editors now, and they work. They are, however, features you invoke after opening the file, choosing the range and setting the parameters. The distinction here is that no human is in the loop until the output already exists, which changes how many clips actually get produced in a week.

vs. the other automatic clipping products

Most of the category converges on the same feature list for uploaded video, and on cut quality you should judge with your own footage rather than a demo reel. Two things are genuinely uncommon: clips produced from a live broadcast in real time, and the whole engine being reachable from Claude or a script. Run one of your own videos and compare boundaries — that is the only test that settles it.

Frequently asked questions

Is this a full video editor like Premiere Pro?
No, and it is not trying to be. There is no multi-track timeline, no colour grading, no motion graphics and no audio mixing desk. It is an automatic editor for one specific job: turning long recordings into short, captioned, correctly framed clips without you opening a timeline at all. For anything requiring frame-level craft, keep your existing editor.
What does the AI actually decide?
Four things, mainly. Which segments of the recording are worth cutting, where each one should start and stop so it opens and closes on a whole thought, how the frame should be composed once it becomes vertical, and how strong the finished result is compared with the others. Captions, filler removal and layout selection follow from those decisions.
Can I change a clip after the AI produces it?
Yes. In and out points can be nudged, the caption style can be switched among eleven options, the layout can be changed, the title can be rewritten, and the clip re-renders with your changes. Treat the output as a strong first pass rather than something locked.
Is the AI video editor free to try?
A free demo runs on any video shorter than thirty minutes, so you can test the editing decisions on footage you already know well. Full access begins with a three-day trial priced at one dollar, after which it is twenty-nine dollars a month for Pro. Cancel in one click. We email you before the trial converts, so nothing is charged quietly.
Do exported clips carry a watermark?
Clips produced in the free demo do carry a small watermark, because the demo exists to show you the quality rather than to supply finished inventory. Paid exports come out clean. If you would rather have your own mark on every clip, the Brand Kit will place your logo instead.
How long a video can I put in?
A single paid upload may be up to two hours long. The demo is limited to thirty minutes. If your source runs past two hours, either split it into parts or, when it is a broadcast, connect the channel and let live mode cut it as it happens.
Do I need to download the video first?
Not for anything already published to YouTube. Paste the link and it is fetched server-side, so nothing lands on your machine and nothing eats your upload bandwidth. For footage that was never posted anywhere, upload the file directly.
Does it do colour correction or motion graphics?
No. Grading, keyframed effects, animated titles beyond the caption styles, and compositing are all outside what this does. Those are timeline tasks and a timeline tool does them better. What arrives here is a clean cut with captions and correct framing, which is what short-form actually needs.
Can it generate video, B-roll or AI footage?
No. There is no text-to-video, no synthetic actors and no generated B-roll of any kind. Every frame in your output came from the footage you supplied. If a tool promises to invent shots for you, that is a different product category and we are not in it.
Does it create a synthetic voiceover or dub other languages?
No voiceover generation and no dubbing. The voices in your clips are the voices in your recording. Captions are transcribed from that same audio and burned in, which covers the mute-viewing problem, but the audio itself is never synthesised or replaced.
What exactly is the 0-100 score?
It is a grade for how the clip is built: how much work the opening seconds do, how consistently it holds pace, and how its strongest beat is positioned. Read it as a ranking device for the batch in front of you rather than a prediction of views. Ordering twelve clips correctly is a solvable problem; forecasting an algorithm is not.
How many usable clips come out of a single file?
It depends entirely on how much of the source stands on its own. A dense twenty-minute talk might give three or four genuinely usable pieces. A long, conversational recording often gives ten to twenty. We would rather return six good candidates than pad the count to a headline number with weak filler.
Which aspect ratios does it export?
Vertical 9:16 for Shorts, Reels and TikTok, square 1:1 for feed placements, and 16:9 if you want a landscape version for a site or an email. Vertical is the default because that is where short-form distribution lives.
Will it frame a two-person conversation correctly?
Yes, and this is one of the clearer wins over naive cropping. Speaker tracking follows whoever is talking, and for a two-hander the split layout is usually chosen so both faces stay visible through the exchange. The interview clip generator goes into how that framing is decided.
Does it remove filler words and silence?
Automatically, on every clip, without a setting to enable. Ums, throat-clearing, restarts and long gaps come out. The effect is most obvious on conversational footage, where the recovered seconds usually amount to a noticeably tighter piece.
Can I control the length of the clips?
Clip length is chosen to fit the thought rather than a fixed target, because a forty-second idea forced into fifteen seconds loses its point. You can shorten or extend any individual clip afterwards using the handles, and re-render.
Does it support a brand kit for consistent styling?
Yes. The Brand Kit stores your logo, colours and typeface and applies them consistently, so a month of clips looks like it came from one channel rather than from whatever caption style you happened to click that day.
Can it post the clips for me?
Yes. Connect TikTok, Instagram Reels and YouTube Shorts, then publish immediately or drop clips into a schedule. Doing the clipping and the scheduling in one sitting is what turns a single recording into several weeks of posts.
Can it edit a live stream while it is running?
It can, and this is the capability most editors cannot match on architecture alone. Connect a YouTube, Twitch or Kick channel and clips are produced during the broadcast rather than after the VOD is available, so a moment can be posted while people are still talking about it.
Is there an API or a way to script it?
Yes. There is a developer API for submitting jobs and retrieving finished clips, plus an MCP connector so Claude can run the whole flow conversationally. Endpoints and authentication are documented at the developer docs.
Do I have to install anything?
No. It runs in the browser and the processing happens on our machines, so an ageing laptop is not a constraint. You can start a job and close the tab; the clips are waiting when you come back.
What footage does it handle badly?
Anything without clear speech. Music videos, silent B-roll, ambient footage and heavily scored montages give the model nothing to reason about, and the boundary logic depends on sentences existing. Loud background music over quiet dialogue also degrades transcription, which then degrades both the captions and the cut points.
How long do I wait before the clips exist?
A few minutes is typical, and the figure moves with the runtime of the source and how loaded the queue happens to be. Nothing needs watching while it runs. Submit the job, go and do something else, and the finished clips will be sitting there when you come back.
What happens to my footage afterwards?
Your upload is used to produce your clips and nothing else, and we do not publish it anywhere. Ownership of the output stays with you. The specifics of storage and retention are laid out in the privacy policy.
Is the AI going to replace my editor?
For volume short-form, it removes work an editor probably resented anyway. For anything carrying your brand — the flagship video, the launch film, the client deliverable — a human is still obviously better, because those depend on taste and context that no scoring model has. The sensible arrangement uses both, aimed at different work.
Which alternatives should I look at before settling on this?
Three groups are worth an afternoon. A traditional editor if you need frame-level craft control. A mobile editor such as CapCut if you already know which moment you are cutting and only want it dressed. And the other automatic clippers, which largely converge on the same feature list for uploaded files — judge that group on where their boundaries land using your own footage rather than their demo reel, and check whether any of them touch a live source at all.
How is this different from an auto-cut button in another app?
An auto-cut button acts on a range you already chose inside a file you already opened. Here, no human touches the material before output exists. That sounds like a small distinction and it is not: it is the difference between editing faster and editing without sitting down.
Is cancelling the one-dollar trial straightforward?
It is a single click from the account page, and a reminder email lands before the trial turns into a subscription. Cancelling leaves your access intact through the period you already paid for and stops anything further being charged.

Give it footage you already know well

The fastest way to evaluate an automatic editor is to run it on a recording where you already know exactly which moment is the best one, and see whether it agrees.

⚡ Get A.I Clips — $1 trial
3-day trial · just $1 · cancel anytime