Solopreneurship.eu
Build & Vibecoding

DIY: make a promo video without a camera — AI stills, a few seconds of motion, a voice and word-timed captions, for under a dollar

The exact production line behind my video channels, applied to one 20–30-second promo: generate the frames, animate only the seconds that matter, record or synthesise the voice, time the captions to the words, assemble with free tools. Real per-clip costs, the hook rule, and the four places it breaks.

EU-focused
Konstantin Filatov

Solo operator · one-person venture studio in Europe (SEO · affiliate · micro-SaaS) · 21 September 2026 · updated 21 September 2026 · 4 min read

DIY: make a promo video without a camera — AI stills, a few seconds of motion, a voice and word-timed captions, for under a dollar

You need a 20–30-second video: for the product page, the app listing, the launch post, the ad. No camera, no actor, no studio, no editor on retainer. This is the production line behind my seven channels, which turns out to be exactly a promo-video line — one episode a day per channel, under a dollar each.

You end up with
A finished 20–30-second vertical or horizontal video: your visuals, one animated moment, a voice, captions timed to the words, your track underneath
What it costs
Under a dollar in generation: ~$0.03 per still, ~$0.25 per five-second animated clip, voice and assembly free. Plus a track from the music recipe
How long
Two to three hours the first time; under an hour with the template
Runs on
An image model · an image-to-video model (five-second clips) · your phone or a free TTS engine for the voice · a free speech recogniser for word timing · ffmpeg or any free editor

Step 1 — the script is eight words, then the rest

Write the first line first, and make it do the work. The rule the channels run on: the first three seconds are a spoken hook of eight words or fewer that gives the viewer a reason to stay — a promise, a stake, a question with an open loop. Never “Welcome”, never “In this video”, never “Did you know”. Then four to six short sentences that pay the promise off, and a last line that closes the loop.

Read it aloud with a timer. A 25-second promo is about 55–65 spoken words. Cut until it fits — the model won’t, the platform will.

Step 2 — the frames

One still per sentence, from the creatives recipe: your style sentence, a subject, a seed per frame. Generate at the aspect ratio of the destination — 9:16 for Shorts, Reels and stories, 16:9 for a product page. If the product has a screen, use the grey plate and composite the real screenshot; if the product is physical, generate the scene and composite the real photo.

Four to six frames is the right number. More than that and the video becomes a slideshow.

Step 3 — animate one moment, not the whole thing

Pick the one frame where motion matters — the product turning, the door opening, the hand reaching — and send only that frame to an image-to-video model for a five-second clip. Prompt for minimal motion (“slow, subtle, camera nearly still”) and add a negative prompt against text appearing on screen. The model rewrites small details after about three seconds, so plan to use the first two or three.

The other frames get a slow push-in or drift in the editor — a 3–5 % scale over the clip’s duration. That is enough motion for a still to feel alive and it costs nothing.

Step 4 — the voice

Two options, both free:

  • Your own. Phone, quiet room, a blanket over your head if the room echoes, read it three times, keep the take where you weren’t trying. Normalise loudness in a free editor.
  • A synthetic narrator. Open-source and free text-to-speech engines are good enough for a promo; pick a voice once and use the same one on every video so it becomes part of the brand. Say the brand name the way the engine says it, or respell it phonetically in the script.

Either way, export a clean WAV before touching anything else.

Step 5 — captions timed to the words

Sixty percent of viewers watch with the sound off. Captions are not accessibility garnish; they are the video for most people. Run the voice track through a free speech recogniser with word-level timestamps (the small open-source models do this well and run on a laptop), then align the output to your script so the spelling is yours, not the recogniser’s. Show one or two words at a time as they’re spoken, in the brand font, in the safe zone — top 15 % and bottom 20 % of a vertical frame are covered by platform UI.

Step 6 — assemble

Timeline: hook frame at 0:00, the animated insert where the script points at it, the rest of the frames in order, each on screen for exactly its sentence. Voice on top, captions on top of that, the track underneath at −18 to −22 dB so it never fights the voice. End on the last caption, hold half a second, cut. No outro, no logo sting longer than a second — the last frame is the one the platform shows on loop.

Save the project as a template. The second promo takes an hour because Steps 1–6 are now a checklist.

Where it breaks

  • Nobody stays past three seconds. The hook was a greeting or a description. Rewrite the first line as a promise or a stake; it is the only line that matters.
  • The animated clip drifts. You used more than three seconds of it, or the prompt asked for a lot of motion. Use less, ask for less.
  • The captions are out of sync. You typed the timing by hand or trusted the recogniser’s spelling. Word-level alignment against your own script fixes both.
  • It reads as AI. Everything is animated, everything is “cinematic”, and a face keeps changing. Stills with push-ins, a medium instead of an adjective, no faces or faces composed away.

Two doors from here

What you just did has a name: you solved a business problem without hiring an intermediary. People who keep doing that are called solopreneurs. If that sounds like you, start here — or check whether you are ready. All recipes: Do it yourself.

Related: how to make a film alone with AI · the faceless channel playbook · all recipes in Do it yourself.

Frequently asked questions

Do I need to show my face or use my voice?
Neither. The recipe works with no person on screen — that is how my channels run — and the voice can be your own recorded on a phone, or a synthetic narrator from a free text-to-speech engine. If you use a synthetic voice, check the platform's disclosure rules: YouTube's altered-content label covers specific cases such as a real person appearing to say something they did not, not 'made with AI' in general. Read the current rule rather than ticking the box by reflex.
What does one 25-second promo cost?
My episodes run between about 25 cents and a dollar in generation: roughly three cents per still, about a quarter per five-second image-to-video clip, voice free (recorded or open-source TTS), assembly free (ffmpeg or any free editor). A promo with four stills, one animated insert and a voice lands under a dollar. Your time is the real cost: two to three hours the first time, under an hour once the template exists.
Why not just animate everything?
Because image-to-video is the expensive step and the one that drifts. Five seconds of motion is enough for one insert; beyond three seconds the model starts rewriting details, and a fully animated 30 seconds costs more, looks worse and reads as generated. Still frames with a slow push-in, one animated moment, a voice and captions hold attention better and cost a tenth.
Was this useful?

Keep reading

Everything here is free. If something saved you time, you can support the author — no product, no signup.