Tools

Creating a short video with AI, from avatar to editing

The full workflow tested with Claude Code, HeyGen, ElevenLabs and HyperFrames, and the raw result with no retouching.

Louis Graffeuil
Louis Graffeuil
Founder Tandem
October 2, 2026Published
6 minread
Clay-style illustration of a film clapperboard, a smartphone showing an avatar, a sound wave and a coral play button above an editing timeline

People often say editing takes 80% of a video's production time. True, but the other 20% is expensive too: filming yourself, botching the take, starting over, handling sound and lighting.

In this article, I show you the full workflow to create a talking-head AI short video without ever turning on the camera, in a single run. Be forgiving: I share the result further down, with no retouching.

Before getting practical, a word on what has changed lately. LLM performance has improved enormously (Opus 5.5, Astra), and with it the video creation and editing tools (avatar, voice, video).

Here is an example of how video has evolved over the last few years:

Here are the 6 steps to create a short:

Diagram of the six steps to create an AI short video with Claude Code: news and script, ElevenLabs voice, HeyGen avatar, dressing plan, looped production with HyperFrames and fal.ai, 1080x1920 MP4 export

Set up your avatar and cloned voice once

There is a one-time setup: create the avatar, clone the voice, install the tools, store the API keys. It takes some time at first, but once it is done you can focus on the rest: you give a topic, you approve four decisions, you get an MP4. A few minutes of your time.

The AI avatar on HeyGen, filmed once

You film yourself in one go, 30 seconds minimum, 2 minutes recommended, on HeyGen. I kept a free plan for this demo.

Watch your camera quality: I shot with my Mac's camera, at rather low quality. Think about the target format too. The source format is landscape, and for a 9:16 short HeyGen crops the center. You lose sharpness as a result (which is exactly what happened to us).

I covered AI avatars in detail in this comparison of AI avatar tools, so I won't go over it again.

Cloning your voice with ElevenLabs

HeyGen's built-in voice clone is average. So we go through ElevenLabs, then import the audio.

Two ways to clone:

  • Instant clone: 1 to 2 min of audio, ready in a few minutes. Decent likeness, no more. Good for testing.
  • Professional clone: 30 min of audio minimum (1 to 2 h recommended), several hours of training, and a voice check where you read a sentence aloud. That is the one we kept for the short.

Producing the short: the cloned voice sent to HeyGen

Once the cloned voice is ready, Claude writes the script, then goes through ElevenLabs to get an MP3. We send that MP3 straight to HeyGen, which lip-syncs the avatar to it.

On cost, you pay per minute produced: 48 credits/min with Avatar V (about $1.45 on the Creator plan for our short).

Prompt sent to Claude Code to generate my HeyGen avatar from voix.mp3 in 9:16 and 1080p, and the agent's reply listing my three looks before waiting for my confirmation

Video editing with HyperFrames

Here we use HyperFrames, HeyGen's open source framework. The idea: you write HTML, you get a video. Everything runs locally, for free, included in your Claude Code subscription.

I covered it in detail in my article on video editing with Claude Code.

The interesting part: the agent does not jump straight into editing. It transcribes the audio word by word, then writes a dressing plan in a file, with every element of the frame at each moment:

One frame of the short at 23.6 seconds broken into five layers: dotted cream background, fal.ai B-roll, my HeyGen avatar in a window, animated overlays and subtitles, stacked by HyperFrames

It presents the decisions to make, each with its recommendation. Nothing is built and nothing is paid for until we approve.

The verification loop with /goal

We set an end condition with the /goal command. The agent edits, exports, takes screenshots of its own render, looks at them, spots what is wrong and starts again. On its own.

The /goal loop in Claude Code: build a pass, run HyperFrames checks, review screenshots, have an evaluator judge. In the test, 3 passes, 48 moments reviewed and 2 defects fixed

On our short, it noticed the title did not show on the first frame and relaunched the render by itself. It ran for about an hour while we did something else.

Why Claude Code rather than the chat to create a short?

Because the chat does not accept video attachments, and Claude cannot read a video.

And that is the whole problem. A short is not a text, it is a tree of files talking to each other. A voice-over MP3, an avatar MP4, a word-level timestamped transcript, a dressing plan in Markdown, three B-roll shots, sound effects, and the edit file that calls all of them.

Track-by-track timeline of the 37-second short: five-part script, avatar on screen, three B-roll shots, animated scenes, 122 subtitle words, music 19 dB under the voice and 37 sound effects

Each track above is an HTML layer, and each step leaves a file that the next one reads. Schematically:

Files left by each step of the short: script.md, voix.mp3, avatar.mp4, transcript.json, PLAN.md, three B-roll clips, BOUCLE.md and a 37.4-second short-final.mp4

Keep in mind that the real value is no longer in the model. It is in the workflow, the skills and the verification loop.

The raw AI short video, with no retouching

Here is the raw result, launched from Claude Code, with no retouching at all. A 37-second vertical short, with animated subtitles, three generated B-roll shots, music and 37 sound effects. My face and my voice from start to finish, without ever turning on the camera.

So yes, it is a first draft, with limits and things that do not work well. But honestly, on screen it holds up. The visual thread is consistent and the pacing is good. If you did not know it was generated, you might not notice.

How much an AI short video costs, and what went wrong

The real cost of the short: $0.37 in total.

  • B-roll on fal.ai: $0.36, three shots kept on the first try
  • ElevenLabs voice: $0.01
  • HeyGen avatar: free in our case, the Free plan includes 3 videos (otherwise about $1.45)
  • Music and sound effects: €0, HeyGen and HyperFrames libraries
  • Editing, transcription, background removal: €0, all local

What went wrong was the avatar quality. HeyGen's free plan refuses 1080p, so we rendered in 720p with a watermark before upscaling locally. My avatars were filmed horizontally, hence a loss of sharpness when cropping to vertical.

There is also room for improvement on the voice. We did not use the latest ElevenLabs model, v4, which was not available for voice cloning at the time of the test.

The limits, today:

  • Video quality: avatars break down on long shots, and generated B-roll still looks like AI.
  • Taste: without art direction or a visual reference, you quickly get slop. The viral one-shots you see all rely on a skill prepared upfront and an idea that works.
  • Rights: to clone a voice or a face, you need a precise written agreement (medium, purpose, duration). And fully generated content cannot be copyrighted in the United States, according to the Copyright Office report of January 2025.
  • Platforms: YouTube, TikTok and Instagram require labeling AI-generated content, and the AI Act sets metadata and disclosure obligations for deepfakes.

Should you create your short videos with AI?

What struck me during this test was not the technology. It was that the rare skill has moved.

In recent years, knowing how to edit a video was a barrier. It is falling. What remains is knowing what you want to tell, and describing it precisely enough for a machine to execute.

In other words: taste, point of view, art direction. The things you do not delegate.

So test it. Take an afternoon to create your avatar and clone your voice. After that, each short will cost you a prompt and a few minutes of attention.

Frequently asked questions

How long does it take to set up an AI avatar and a cloned voice?

Setting up an AI avatar and a cloned voice is a one-time job. On HeyGen, you film yourself in one go for at least 30 seconds, with 2 minutes recommended. On ElevenLabs, an instant voice clone needs 1 to 2 minutes of audio. A professional clone needs at least 30 minutes of audio and several hours of training. Plan for an afternoon in total.

What is the difference between ElevenLabs instant and professional voice cloning?

An ElevenLabs instant voice clone is built from 1 to 2 minutes of audio and is ready in a few minutes. The likeness is decent but no more, so it is mostly for testing. A professional clone requires at least 30 minutes of audio, ideally 1 to 2 hours, several hours of training and a check of your voice. That is the one to use for a published video.

Are you allowed to publish a video made with an AI avatar?

Yes, as long as you follow two rules. Cloning another person's voice or face requires their precise written consent, covering the medium, the purpose and the duration. YouTube, TikTok and Instagram also require you to label AI-generated content. In Europe, the AI Act sets disclosure and metadata obligations for deepfakes.

Why does a HeyGen avatar lose sharpness in vertical format?

A HeyGen avatar filmed in landscape loses sharpness in short format, because HeyGen crops the center of the frame to get 9:16. The free plan also caps the render at 720p with a watermark. For a sharp short, film the avatar vertically with a good camera, and use a paid plan that allows 1080p.

Read next

All articles →
ToolsTandem illustration of n8n project cost, a small node cluster resting on a stack of coins with a dark calculator leaning against it

What an n8n automation project really costs, from build to monthly bill

By Louis Graffeuil
ToolsTandem illustration of the n8n migration, a pale ramp bridging a light platform to a dark one, with connected nodes crossing it

n8n 2.0, what breaks and how to migrate without losing your workflows

By Louis Graffeuil
ToolsDark clay terminal window on a pale plinth, coral cursor, three small pale robots lined up in front

AI coding agents, where the market actually stands

By Louis Graffeuil