People often say editing takes 80% of a video's production time. True, but the other 20% is expensive too: filming yourself, botching the take, starting over, handling sound and lighting.
In this article, I show you the full workflow to create a talking-head AI short video without ever turning on the camera, in a single run. Be forgiving: I share the result further down, with no retouching.
Before getting practical, a word on what has changed lately. LLM performance has improved enormously (Opus 5.5, Astra), and with it the video creation and editing tools (avatar, voice, video).
Here is an example of how video has evolved over the last few years:
Here are the 6 steps to create a short:

Set up your avatar and cloned voice once
There is a one-time setup: create the avatar, clone the voice, install the tools, store the API keys. It takes some time at first, but once it is done you can focus on the rest: you give a topic, you approve four decisions, you get an MP4. A few minutes of your time.
The AI avatar on HeyGen, filmed once
You film yourself in one go, 30 seconds minimum, 2 minutes recommended, on HeyGen. I kept a free plan for this demo.
Watch your camera quality: I shot with my Mac's camera, at rather low quality. Think about the target format too. The source format is landscape, and for a 9:16 short HeyGen crops the center. You lose sharpness as a result (which is exactly what happened to us).
I covered AI avatars in detail in this comparison of AI avatar tools, so I won't go over it again.
Cloning your voice with ElevenLabs
HeyGen's built-in voice clone is average. So we go through ElevenLabs, then import the audio.
Two ways to clone:
- Instant clone: 1 to 2 min of audio, ready in a few minutes. Decent likeness, no more. Good for testing.
- Professional clone: 30 min of audio minimum (1 to 2 h recommended), several hours of training, and a voice check where you read a sentence aloud. That is the one we kept for the short.
Producing the short: the cloned voice sent to HeyGen
Once the cloned voice is ready, Claude writes the script, then goes through ElevenLabs to get an MP3. We send that MP3 straight to HeyGen, which lip-syncs the avatar to it.
On cost, you pay per minute produced: 48 credits/min with Avatar V (about $1.45 on the Creator plan for our short).

Video editing with HyperFrames
Here we use HyperFrames, HeyGen's open source framework. The idea: you write HTML, you get a video. Everything runs locally, for free, included in your Claude Code subscription.
I covered it in detail in my article on video editing with Claude Code.
The interesting part: the agent does not jump straight into editing. It transcribes the audio word by word, then writes a dressing plan in a file, with every element of the frame at each moment:

It presents the decisions to make, each with its recommendation. Nothing is built and nothing is paid for until we approve.
The verification loop with /goal
We set an end condition with the /goal command. The agent edits, exports, takes screenshots of its own render, looks at them, spots what is wrong and starts again. On its own.

On our short, it noticed the title did not show on the first frame and relaunched the render by itself. It ran for about an hour while we did something else.
Why Claude Code rather than the chat to create a short?
Because the chat does not accept video attachments, and Claude cannot read a video.
And that is the whole problem. A short is not a text, it is a tree of files talking to each other. A voice-over MP3, an avatar MP4, a word-level timestamped transcript, a dressing plan in Markdown, three B-roll shots, sound effects, and the edit file that calls all of them.

Each track above is an HTML layer, and each step leaves a file that the next one reads. Schematically:

Keep in mind that the real value is no longer in the model. It is in the workflow, the skills and the verification loop.
The raw AI short video, with no retouching
Here is the raw result, launched from Claude Code, with no retouching at all. A 37-second vertical short, with animated subtitles, three generated B-roll shots, music and 37 sound effects. My face and my voice from start to finish, without ever turning on the camera.
So yes, it is a first draft, with limits and things that do not work well. But honestly, on screen it holds up. The visual thread is consistent and the pacing is good. If you did not know it was generated, you might not notice.
How much an AI short video costs, and what went wrong
The real cost of the short: $0.37 in total.
- B-roll on fal.ai: $0.36, three shots kept on the first try
- ElevenLabs voice: $0.01
- HeyGen avatar: free in our case, the Free plan includes 3 videos (otherwise about $1.45)
- Music and sound effects: €0, HeyGen and HyperFrames libraries
- Editing, transcription, background removal: €0, all local
What went wrong was the avatar quality. HeyGen's free plan refuses 1080p, so we rendered in 720p with a watermark before upscaling locally. My avatars were filmed horizontally, hence a loss of sharpness when cropping to vertical.
There is also room for improvement on the voice. We did not use the latest ElevenLabs model, v4, which was not available for voice cloning at the time of the test.
The limits, today:
- Video quality: avatars break down on long shots, and generated B-roll still looks like AI.
- Taste: without art direction or a visual reference, you quickly get slop. The viral one-shots you see all rely on a skill prepared upfront and an idea that works.
- Rights: to clone a voice or a face, you need a precise written agreement (medium, purpose, duration). And fully generated content cannot be copyrighted in the United States, according to the Copyright Office report of January 2025.
- Platforms: YouTube, TikTok and Instagram require labeling AI-generated content, and the AI Act sets metadata and disclosure obligations for deepfakes.
Should you create your short videos with AI?
What struck me during this test was not the technology. It was that the rare skill has moved.
In recent years, knowing how to edit a video was a barrier. It is falling. What remains is knowing what you want to tell, and describing it precisely enough for a machine to execute.
In other words: taste, point of view, art direction. The things you do not delegate.
So test it. Take an afternoon to create your avatar and clone your voice. After that, each short will cost you a prompt and a few minutes of attention.



