If you generate an AI clip in 16:9 and crop it to 9:16 for Shorts, Reels or TikTok, you keep about a third of the frame's width and lose the rest. The part you lose is often what the shot was about: the hands, the product, the second person in the conversation.
Generating in 9:16 from the start avoids that, but the setting alone isn't enough. The app draws buttons and text over your video, and captions need a clear area to sit in. This guide covers where the ratio setting lives, how to compose for a tall frame, safe zones and caption space, when cropping still makes sense, and what each resolution costs. Model facts and prices are as of October 2026 and link to their sources.
ClipSpeedAI finds the strongest moments in your streams, podcasts and YouTube uploads, cuts them to vertical 9:16, burns in captions from what was said and gives each clip a viral score.
Try ClipSpeedAI →Orientation is a generation setting: a dropdown in an app or a parameter in an API. Don't count on "9:16" or "vertical" in the prompt to change the frame the model renders: if the setting is left at 16:9, expect a widescreen clip whatever the text says. Our AI video prompt guide covers what belongs in the prompt and what belongs in settings.
Native vertical support differs by model and by host. Here is what we could and couldn't verify:
| Model | Native 9:16 | Resolution tiers |
|---|---|---|
| Veo 3.1 (Google) | Yes, since the January 13, 2026 update | 720p, 1080p and 4K. 1080p and 4K need an 8-second clip, and the Lite tier has no 4K |
| Seedance 2.5 (ByteDance) | Not confirmed in our sources; check your host | 480p and 720p at launch, per fal's comparison; fal now lists 1080p. 4K is unconfirmed |
| Kling 3.0 Turbo (Kuaishou) | Not confirmed in our sources; check your host | 720p on Standard, 1080p on Pro; no 4K tier listed |
| Gemini Omni 1.1 Flash (Google) | Not confirmed in our sources | 360p, 720p, 1080p and 4K, with 1080p and 4K made by upscaling |
| MiniMax H3 and fal's H3 Max | Not confirmed in our sources | Up to 2K via the API; the open weights default to 768p. H3 Max tops out at 768p |
The same model can expose different options on different hosts, so check the ratio list where you actually generate. Then check the file you get back. This command prints a clip's width, height and frame rate:
ffprobe -v error -select_streams v:0 -show_entries stream=width,height,r_frame_rate -of csv=p=0 clip.mp4
A vertical clip prints the smaller number first, such as 720,1280 followed by the frame rate. If the width is bigger than the height, the setting didn't take.
A 9:16 window cut from a 16:9 frame at full height keeps 81/256 of the width, about 32%. Here is what that means in pixels, measured against 1080×1920, the standard full-HD vertical frame:
| Source | What you keep | To fill 1080×1920 |
|---|---|---|
| 1280×720 landscape (720p), cropped | 405×720 | Scale up about 2.7× |
| 1920×1080 landscape (1080p), cropped | 608×1080 | Scale up about 1.8× |
| 3840×2160 landscape (4K), cropped | 1215×2160 | Scale down |
| 720×1280 native vertical (720p) | The full frame | Scale up 1.5× |
| 1080×1920 native vertical (1080p) | The full frame | None |
The row to notice is the second one. A crop from a 1080p landscape clip is 608 pixels wide, narrower than a native 720p vertical clip at 720. You pay for the higher tier and end up with fewer pixels across.
Composition is the bigger problem. A model composing for 16:9 tends to spread the action across the width: two people at opposite sides, a camera tracking sideways. A centered crop cuts those apart, and a moving subject drifts out of a fixed crop unless someone keyframes the reframe.
Otherwise, generate twice. On Veo 3.1 Fast at $0.10 per second at 720p, an 8-second clip costs $0.80, so a 16:9 version and a 9:16 version cost $1.60 together. They won't be identical shots, which is usually fine for a hook or b-roll. One 8-second 4K Fast clip to crop from costs $2.40, and Google describes Veo's 1080p and 4K as upscaling, so you'd be cropping an upscaled image rather than a sharper original render.
Write these as shot directions in the prompt, not as ratios. For example:
Medium close-up of a woman at a kitchen counter, centered,
head and shoulders in the upper half of the frame. She lifts
a ceramic mug toward the camera. Shallow depth of field, soft
window light, plain cream wall behind her. Slow push-in.
For image-to-video, make the start frame 9:16 as well. If the still is landscape and the video is vertical, the model or host has to crop it or invent the missing edges. Our image-to-video workflow covers making start frames that match the output.
Every short-form app draws its interface over your video. On Shorts, Reels and TikTok that generally means the account name, post caption and sound line across the bottom, a column of like, comment and share buttons down the right edge, and tabs or a search bar across the top. Product tags, links and ad labels can add more.
We're not printing pixel figures: layouts change with app updates and vary by device, so an old template's numbers can quietly be wrong. Measure your own:
Plan for one more overlay if the shot looks realistic. YouTube requires creators to disclose realistic altered or synthetic content and shows the label as an overlay on Shorts, and TikTok's Community Guidelines require a label on AI-generated content showing realistic scenes or people. Our disclosure rules guide covers when each platform expects one.
Pick the caption position before you generate, because it decides where the subject can't be. A face sitting under a block of captions is hard to fix in the edit.
Don't ask the model to render captions or other text into the video: you can't move it, restyle it or fix a typo later. Add text in the edit. And if the clip has generated dialogue, caption what the model actually said, not the line you wrote, because the two can differ. Google itself calls natural speech in short segments an area of active development for Veo.
When generated shots sit between pieces of real footage, keep the caption position fixed and generate b-roll with its subject in the same zone as your talking head, so nothing important lands under the text after a cut. More on that in AI b-roll for talking-head videos, and on caption styling in how to add captions to Shorts.
Resolution labels usually refer to the short side of the frame, so a 720p vertical clip is typically 720 pixels wide and 1280 tall, and 1080p vertical is 1080×1920. Run the ffprobe check above to confirm what you got.
Higher tiers cost more per second. Here is what one 8-second clip with audio costs at list per-second rates, as of October 2026, for four options that offer both 720p and 1080p:
| Model | 8 s at 720p | 8 s at 1080p |
|---|---|---|
| Veo 3.1 Lite | $0.40 | $0.64 |
| Veo 3.1 Fast | $0.80 | $0.96 |
| Kling 3.0 Turbo (Standard / Pro) | $0.90 | $1.12 |
| Seedance 2.5 on fal | $3.78 | $9.31 |
Veo needs the 8-second length for 1080p anyway. Kling 3.0 Turbo's 3 to 15 second range comes from a third-party guide, and our sources don't confirm 9:16 for Kling or for Seedance 2.5 on fal, so check length and ratio options before paying for a vertical run.
720p vertical is a sensible default for hooks and b-roll. Softness shows first in fine texture, wide scenes and small faces; close-ups hide it. If a 720p shot looks soft on your phone, try framing closer before paying for 1080p.
For upload settings on each platform, see our resolution guide for Shorts, Reels and TikTok.
ffprobe.ClipSpeed's AI Creator puts five video models side by side: Seedance 2.5, Kling 3.0 Turbo, MiniMax H3 Max, Veo 3.1 and Gemini Omni 1.1 Flash, plus Nano Banana Pro and GPT Image 2.5 Flare and Sunburst for images. For vertical shots, start with Veo 3.1, the one our sources confirm for native 9:16. Use Kling 3.0 Turbo for talking-head clips and volume, Gemini Omni 1.1 Flash to edit existing clips, and an image model for the start frame. Generations use ClipSpeed creation credits (pricing).
Next to it sits AI clipping. Paste a video link and ClipSpeed finds the strongest moments, cuts them to 9:16 or 16:9, burns in captions from the spoken words and gives each clip a viral score. It also clips Twitch, Kick and YouTube live streams in real time. The ClipSpeed MCP server runs the same clipping from Claude. For vertical work, clip your real footage first, then generate only the hook or b-roll shots it's missing.
Set 9:16 in the generator's controls, not the prompt, and check that the file came back vertical. Native vertical beats cropping on both pixels and composition: a crop from 1080p landscape is narrower than a native 720p vertical clip, and a widescreen composition rarely survives the cut.
Compose for the tall frame with one subject, depth and vertical motion, keep faces and action clear of the app's bottom band and right-hand column, and decide where captions go before you generate. Draft cheap, finish at 720p or 1080p, and spend the difference on retries rather than on resolution you'd crop away.
Can I just type "9:16" in my AI video prompt?
Don't rely on it. Aspect ratio is a setting, either a dropdown in the app or a parameter in the API, and typing "9:16" or "vertical" in the prompt isn't a dependable way to change the frame the model renders. Set the ratio in the controls, then use the prompt for framing: where the subject sits, how close the camera is and which way things move.
Is it better to generate vertical or crop a 16:9 AI clip to 9:16?
Generate vertical when your model and host support it. A full-height 9:16 crop keeps about 32% of a 16:9 frame's width, so a 1080p landscape clip crops to roughly 608×1080, narrower than a native 720p vertical clip at 720×1280. Cropping also cuts apart compositions built for a wide frame. Crop only when you need both orientations from one specific take or the model you need has no 9:16 option.
Which AI video models generate native 9:16?
Veo 3.1 has generated native vertical video since a January 13, 2026 update. Our sources don't confirm native 9:16 for Seedance 2.5, Kling 3.0 Turbo, Gemini Omni 1.1 Flash, MiniMax H3 or fal's H3 Max, so check the ratio list on the host you use.
What resolution should vertical AI video be?
A 720p vertical clip is typically 720×1280 and 1080p is 1080×1920, the common full-HD vertical export. 720p is a sensible default for hooks and b-roll; 1080p costs more per second. As of October 2026, Seedance 2.5 on fal lists $0.473 per second at 720p and $1.164 at 1080p. Check the actual file dimensions with ffprobe.
Where should captions go on AI-generated vertical video?
Decide before you generate, because the caption position decides where the subject can't be. A common choice is lower-middle, just above the app's bottom interface band, with the face prompted into the upper half of the frame. Add captions in the edit rather than asking the model to render text, and caption what the generated dialogue actually says.
What are safe zones on Shorts, Reels and TikTok?
They are the parts of the frame the app leaves clear of its interface. The account name, post caption and sound line cover the bottom, a column of buttons covers the right edge, and tabs or search cover the top. Layouts change with app updates and differ by device, so build an overlay from screenshots of your own feed instead of trusting old pixel templates.
Can ClipSpeedAI generate vertical AI video?
ClipSpeed's AI Creator runs Seedance 2.5, Kling 3.0 Turbo, MiniMax H3 Max, Veo 3.1 and Gemini Omni 1.1 Flash for video, plus Nano Banana Pro and GPT Image 2.5 Flare and Sunburst for images such as start frames. For vertical shots, start with Veo 3.1, which Google updated for native 9:16, and check the ratio options on the others before you generate. Generations use ClipSpeed creation credits. Separately, ClipSpeed's AI clipping cuts your existing videos into vertical, captioned shorts. You can start clipping here.
Published by ClipSpeedAI · AI video generation and AI clipping in one place — create with Seedance, Veo, Kling and Nano Banana, then cut it into captioned shorts.