How do I generate a video with AI? A step-by-step guide
Generate a video with AI in 6 steps: pick the format, write a prompt the model can follow, choose a model, generate, fix what breaks, then add voice and captions.
By the Revid teamUpdated 16 min readBeginner
Quick answer
To generate a video with AI, write a prompt that describes one shot (the subject, what it does, where, in what style and how the camera moves) or upload an image to use as the first frame. Then pick a video model, a clip length of 5 to 10 seconds and an aspect ratio (9:16 for TikTok, Reels and Shorts), and generate. On Revid, the fast tier returns a clip in about 23 seconds (median); a premium model takes 4 to 10 minutes. For a finished video, chain several clips and add a voice-over, captions and music, or let a prompt-to-video tool do all of it from one sentence.
- 1Decide what you are making: one clip, or a finished video with a voice-over.
- 2Pick the format first: 9:16 for TikTok, Reels and Shorts, 16:9 for YouTube.
- 3Write a prompt: subject, action, setting, style and camera.
- 4Choose a model, a length and a quality tier. Draft on the cheapest one.
- 5Generate, review, and change one thing at a time.
- 6Add voice, captions and music, then export in 1080p and publish.
- Median time for one AI clip on Revid's fast tier
- 23 s
- Median time for one AI clip on Revid's fast tier
- Sep 2026, 19,225 clips
- Median for a premium model (Seedance 2.5)
- 6.7 min
- Median for a premium model (Seedance 2.5)
- Aug to Sep 2026
- Faster when you start from an image instead of text
- 2.2×
- Faster when you start from an image instead of text
- 18 s vs 39 s median
- Credits per 5 seconds of video, depending on the model
- 15 to 130
- Credits per 5 seconds of video, depending on the model
- Revid pricing, Oct 2026
Two ways to generate a video with AI
"Generating a video with AI" means two different jobs, and most confusion comes from mixing them up. A clip generator turns a prompt or an image into one shot that lasts a few seconds. A video generator turns an idea or a script into a finished video: several shots, a voice, captions and music, cut together.
The underlying models are the same either way. A video generator calls a clip model once per scene, then does the work around it: writing the script, splitting it into scenes, generating a voice, timing captions and editing. If you want to learn how prompting works, start with single clips: they are cheap and fast. Move to a full video when you know what you want to say.
Step 1: Pick the format before you write anything
Choose the aspect ratio and length for the platform first, because the model composes every shot for the frame you ask for.
A video model frames the shot for the aspect ratio you give it. Cropping a 16:9 clip into a 9:16 vertical keeps only 32% of the width, so subjects get cut in half and the composition falls apart. Decide where the video will be posted, then set the ratio before the first generation.
Then decide the length. A single AI clip lasts between 4 and 30 seconds depending on the model, so a 30-second post is usually 6 to 10 clips of 3 to 5 seconds each, one per sentence of the script. Short clips also fail less: the longer a clip runs, the more frames the model has to invent and the more the subject drifts.
Step 2: Write a prompt the model can follow
Describe one shot with five parts: subject, action, setting, style and camera, in plain visual language.
Video models respond to what a camera could see. They do not understand intentions ("she misses her brother"), but they understand faces, objects, motion, light and lenses. A reliable prompt answers five questions in two to four sentences.
Too vague
a dog at the beach
Specific
A golden retriever puppy sprints along wet sand at sunrise, ears flapping, spray kicking up behind its paws. Low tracking shot at the dog's eye level, shallow depth of field, warm golden light, cinematic realism.
The vague prompt leaves every decision to the model, so you get an average shot: a dog standing still, centered, flat light. The specific one fixes the action, the speed, the camera height and the light, which are the things that make a clip look intentional.
Rules that save the most retries
- One action per clip. A 5-second clip holds one clear action. "He opens the door, walks in and sits down" becomes three clips.
- Name the camera move once. "Slow push-in", "tracking shot", "static wide shot", "aerial drone shot". Two moves in one prompt usually means neither happens.
- Describe the end of the shot for longer clips. For 10 seconds and more, say where the action ends ("...and stops at the edge of the water"). Premium models also accept time ranges such as "0-5 seconds: ... 5-10 seconds: ...".
- Keep words out of the frame. Models draw letters as shapes, so signs and labels come out as gibberish. Add text later as captions or overlays.
- Do not name real people, brands or characters. Describe the look instead ("a tall man in a grey suit", "a red sports car"). Named people and protected characters are the main reason providers refuse a generation (see step 5).
- Write the style once, at the end. Mixing "photorealistic" and "anime" in the same prompt produces neither.
Copy a prompt that already works
These four prompts follow the formula. Each one opens in Revid with the prompt and settings filled in, on the cheapest tier, so you can test them before you spend more.
Product shot
A frosted glass perfume bottle stands on wet black stone. Water droplets slide down the glass as a slow ring of mist curls around it. The camera orbits slowly from left to right at bottle height. Soft studio light from above, deep shadows, luxury commercial look.
Cinematic realism
A golden retriever puppy sprints along wet sand at sunrise, ears flapping, spray kicking up behind its paws. Low tracking shot at the dog's eye level, shallow depth of field, warm golden light, cinematic realism.
3D animation
A small copper robot with round glowing eyes waters a single flower on a windowsill, then turns to the camera and waves. Slow push-in. Pixar-style 3D animation, warm morning light, soft colors.
Story moment
An elderly man sits alone at a window table in a small Paris café, two cups of coffee in front of him. He slides the second cup toward the empty chair and looks out at the rain. Static medium shot, then a slow push-in on his face. Muted film colors, soft window light.
Build your own prompt
Fill in the five parts and the builder writes the prompt in the order models read best. Copy it, or open it in Revid.
Your prompt
A golden retriever puppy with a red collar sprints along the shoreline, spray kicking up behind its paws, on an empty beach at sunrise. Low tracking shot. Golden hour light, cinematic realism.
Step 3: Choose a model, a length and a quality tier
Draft on a fast, cheap tier, then regenerate the shots you keep on a premium model.
Revid's view is that the prompt matters more than the model: the leading models are sold through APIs to every video tool, so two tools that send the same prompt to the same model get comparable clips. Models still differ in speed, length, price and what they are best at. These are the ones available on Revid in October 2026.
How long should each clip be?
Pick the shortest length that holds the action. Cost is billed per second, and waiting grows with length. For social video, 5 seconds per shot is the default that works: long enough for one action, short enough to keep the pace up. Use 8 to 15 seconds for slow, ambient shots (a landscape, a product orbit) and Seedance 2.5's 30 seconds only when you need one continuous take.
What it costs, with real numbers
- A 30-second video made only of Pro AI clips uses 300 credits (6 clips × 50).
- The same 30 seconds on the Low tier uses 90 credits.
- A 30-second video illustrated with Pro AI images and a voice-over costs about 40 to 80 credits. Switching only the 2 or 3 key shots to AI video keeps most of that saving.
- Revid's Growth plan includes 2,000 credits a month: about 200 seconds of Pro AI video, or 660 seconds on the Low tier.
Step 4: Generate, review, and change one thing at a time
Check identity, motion, camera and text in every clip, and change one variable per retry.
Watch each clip twice, once for the whole and once frame by frame on the parts that usually break. Then change one thing in the prompt and generate again. If you change three things at once, you will not know which one fixed it.
- 1Identity: does the subject keep the same face, clothes and colors from the first frame to the last?
- 2Hands and bodies: count the fingers, check the joints, check that limbs do not merge.
- 3Motion: is it physically plausible, at the speed you asked for?
- 4Camera: did it do the move you named, and only that one?
- 5Text and logos: anything written in the frame is probably wrong.
- 6The last frame: does it end in a state that cuts well to the next shot?
Fix what breaks
85%
of failed Seedance calls on Revid were content refusals, not technical errors
Jun 29 to Sep 26, 2026
55%
of all failed Seedance calls were the real-person privacy filter alone
Same window
In other words, when a generation fails, the content of the prompt or the image is usually the problem, not the service. Rewording is faster than retrying the same request. The full breakdown by model is in Revid's AI video model index.
Step 5: Turn clips into a finished video
Write the script, generate one visual per sentence, then add a voice-over, synced captions and music.
A clip is a shot, not a post. People keep watching a short video because of what it says and how fast it gets there: a hook in the first two seconds, one idea per sentence, and a reason to stay until the end. This is the assembly order that works:
- 1Write the script as narration. One sentence per shot. The first sentence is the hook: a question, a surprising number, or a promise.
- 2Split it into scenes. One scene per sentence, 2 to 5 seconds each.
- 3Generate one visual per scene, reusing the same style sentence so the video looks like one piece.
- 4Add the voice-over. An AI voice, in one of 70+ languages on Revid, or your own cloned voice.
- 5Add captions synced to the voice. Many people watch with the sound off, and captions carry the words.
- 6Add music under the voice, low enough that every word stays clear.
- 7Export in 1080p. Use 30 fps for narrated videos and 60 fps for fast action. Revid exports in 720p, 1080p or 4K.
Doing this by hand in an editor takes an afternoon per video. Prompt to Video does all of it from one sentence: it writes the script, picks a visual per scene, generates the voice and captions, and gives you a timeline to edit. These three videos were made that way, each from the one-line prompt shown under it.
The Last Voicemail
A 53-second narrated short film about an old man in a café, with the same character across every shot and word-by-word captions.
The whole promptCreate a 2-3 minute cinematic short film. An elderly man sits alone in a small café every Friday at exactly 4 PM. He always orders two cups of coffee.
The Little Inventor and Pip
A 78-second animated story: a boy builds a copper robot in his workshop. The boy and the robot stay consistent from scene to scene.
The whole promptIn the colorful village of Sunberry Hollow, a curious 10-year-old boy named Oliver dreams of becoming the world's greatest inventor.
Who Owns These Seas
A 19-second cinematic opening in landscape: a captain on a warship at dawn, with his crew behind him.
The whole promptA weathered medieval naval captain stands at the bow of a colossal wooden warship just before dawn.
Notice what the prompts do not contain: no camera directions, no scene list, no style keywords. A video generator writes those for you. When you generate single clips yourself (step 2), you are the one writing them.
Generate your first video
Type one sentence, choose a style and a voice, and Revid builds the script, the scenes, the voice-over and the captions.
Step 6: Export, label and publish
Export in the platform's size, label realistic AI content where the platform asks for it, and publish.
Export at the size from step 1. A 4K export of clips generated at a lower resolution makes the file heavier without adding much detail, so pick 4K only when the clips were generated in 4K. Then check two things before you post:
- The first two seconds. Watch them with the sound off. If they do not make you want the next two, change the hook, not the visuals.
- The AI label. TikTok and YouTube ask creators to disclose realistic AI-generated content, such as a realistic person saying or doing something they did not. Use the platform's label when it applies. Obvious animation and fantasy usually does not need it, but the rules are each platform's to set.
Revid publishes directly to TikTok, Instagram and YouTube from the same place, or on a schedule, and shows the views and likes of each post afterwards, so you can see which prompts and styles actually work for your audience.
Can I generate a video with AI for free?
You can try it for free, but no serious AI video generator is free without limits. Every clip runs a large model on a GPU for seconds to minutes, so each one has a real cost. In practice, "free" means a small allowance of credits or clips, often with a watermark, a lower resolution or a slower queue.
On Revid, signing up is free, and new accounts usually get free credits to try video generation. A free account can also search the library of 2.77 million viral videos and posts and run up to 60 free deep analyses a day. Regular creation uses the monthly credits of a paid plan, from $39 a month ($32 billed yearly) for 2,000 credits.
How to get the most out of free credits
- Generate on the cheapest tier (15 credits per 5 seconds) until the prompt is right.
- Generate 5-second clips, not 10: half the cost, and they fail less.
- Start from an image when you can: on Revid it is 2.2 times faster on the fast tier, and the first frame is exactly what you want.
- Check the credit estimate before you click. Revid shows the cost of every generation and every video before it starts.
Frequently asked questions
How long does it take to generate a video with AI?
On Revid, one 5-second clip takes about 23 seconds on the fast tier and 31 seconds on Pro (medians, September 2026). Premium models are slower: Seedance 2.5 has a median of about 7 minutes. A full narrated video takes a few minutes, because every scene is generated.
How much does it cost to generate a video with AI?
On Revid, AI video costs 15 credits per 5 seconds on the Low tier, 50 on Pro, 100 on Seedance 2.5 and Veo 3.1, and 130 on Sora 2. Plans start at $39 a month for 2,000 credits. The exact cost of a video is shown before you generate it.
Can I generate a video with AI for free?
You can try it for free: signing up to Revid is free and new accounts usually get free credits. Unlimited free generation does not exist anywhere, because every clip costs GPU time. See the section on free options.
How long can an AI-generated video be?
A single AI clip lasts 4 to 30 seconds, depending on the model (up to 30 seconds with Seedance 2.5). Longer videos are made of several clips. Revid assembles them into videos of several minutes with a voice-over and captions.
What is the best AI video generator?
It depends on the job. For one cinematic shot, compare models (Seedance 2.5, Veo 3.1, Sora 2, MiniMax Hailuo). For a finished short video with a voice and captions, use a video generator that calls those models for you. Revid's AI video model index compares speed and failure rates across 713,649 generations, and Revid vs other tools compares the products.
Can AI generate a video with sound or with my voice?
Yes. Some models, such as Veo 3.1 and Seedance 2.5, generate sound with the picture. Revid also adds AI voice-overs in 70+ languages and can clone your own voice on the Elite plan and above.
Why was my AI video generation refused?
Usually because of the content: on Revid, 85% of failed Seedance calls were content refusals. The most common reasons are a realistic real person, a named brand or character, and sensitive subjects. Describe the look instead of naming people or brands.
Do I need video editing skills to make an AI video?
No. A clip generator needs only a prompt, and a video generator like Prompt to Video writes the script, edits the scenes and adds the voice and captions for you. Editing skills help for the final polish, in Revid's timeline editor or any other editor.
Sources and method
- Generation times, refusal reasons and tiers: Revid AI video model index, 713,649 successful AI video generations on Revid between June 29 and September 26, 2026.
- Credit prices and plans: Revid, October 2026. The live price list is on revid.ai/pricing.
- Clip lengths and model capabilities: Revid's model settings as of October 2026.
Keep going
Turn a photo into a short video with AI
Image-to-video, step by step, with motion prompts for 9 kinds of photo.
Video prompt generator
Turn a rough idea into a detailed, model-ready video prompt.
Prompt to Video
One sentence in, a narrated and captioned video out.
AI video generator
Every way to make an AI video on Revid, in one place.
Text to video
Turn a written text into a short video with visuals and voice.
AI video model index
Speed and failure rates of every model, from 713,649 generations.