Creative Studio
API
Resources
Features
About Us
Download
Tutorial · Digital Humans & Avatars

A talking-head video, without you on camera

Generate a fully original presenter face — not your own photo — so you never appear on camera at all. For anonymous creators, privacy-conscious educators, or anyone who wants a human presenter without revealing their identity.

A fully original, AI-generated presenter face — not a photo of any real person — used as this guide's real example.
On this page
  1. How to build one
  2. Filming vs. generating
  3. Real example
  4. Honest limits
  5. FAQ
The short answer

Open Kling's video generation, describe the presenter (or add a reference portrait as the start frame), put the words in quotation marks — she says, "your line" — and keep Native Audio on. The clip renders already talking, picture and voice in sync; nobody gets filmed.

How to build one

1

Describe the presenter — or upload a reference image

In Kling's video generation (VIDEO 3.0), either describe the presenter in words — as the real demo below does — or add a portrait as the start frame so a specific face fronts the clip. Use a generated face rather than a real person; nobody has to be filmed.

2

Put the words in quotation marks

The composer's own instruction: use quotation marks for speaking content — she says, "your line here." The quoted words become the actual speech, and the field supports multiple languages, dialects and accents, so you can specify the delivery.

3

Keep Native Audio on and generate

With Native Audio on, the voice renders with the picture — lips, delivery and sound arrive in sync in one output. No separate voice step, no dubbing pass: the clip below came out of the generator already talking.

4

Review the take before publishing

Listen back with sound on — confirm the line is spoken as written, the pacing feels natural and the expression holds for the whole clip, not just the first second. Regenerate with adjusted wording if the delivery is off.

The presenter's face is the one thing this workflow changes completely — everything else about making a talking-head video stays the same.

Filming vs. generating: three honest routes

Three honest options, not a sales pitch for any one of them:

ApproachAnonymityEstimateTrade-off
Filming yourselfNone — you're on cameraFree (just your time and equipment)The real thing — full authenticity, but requires being visible and re-recording every take.
VIDEO 3.0 speaking generation (this page)High — the presenter is described or referenced, not filmedCost shown on the Generate buttonVoice and picture render together in one pass — fastest route for short spoken clips.
Avatar (Digital Human) toolHigh — works from a portrait plus a scriptCost shown before you generate (same mechanism)Built for longer scripts, with a preset voice library and your own audio upload — a separate tool for a different job.

A real example

This presenter was never a real person and was never filmed — the whole clip, including the voice, is one VIDEO 3.0 generation from the prompt below. Play it with sound on:

Script used

A confident woman in a smart blazer sits facing the camera in a bright home office and says, "You can publish a talking video without ever filming yourself." She keeps steady eye contact with natural small hand gestures, camera locked off.

The quoted sentence became the spoken line, and the voice rendered with the picture in the same generation — no real person appears anywhere in this clip, and no separate voice step was involved.

Honest limits

What to expect

A synthetic presenter is not a perfect substitute for real human presence — attentive viewers can often sense subtle unnaturalness in longer clips, and content that depends on deep personal trust or authenticity may not land the same way. Disclosing that a presenter is AI-generated is good practice, not a legal requirement in most contexts, but it costs little and keeps the relationship with your audience honest.

Frequently asked questions

What is a talking head video?
A shot of one person speaking to camera, framed from roughly the chest up — explainers, course modules, product walkthroughs, executive updates. The format carries information through the voice, so the visual only has to be steady and clear.
How do you make a talking head video without filming?
In Kling's VIDEO 3.0 generation, describe the presenter (or add a reference portrait), put the line to be spoken in quotation marks, and keep Native Audio on — the clip renders with the speech already in it, picture and voice in sync. No camera, no studio, no on-camera talent.
How do you make a talking head video interesting?
Cut away from the face. A talking head holds attention for a while on its own, then needs something else on screen — a diagram, a product shot, a short B-roll clip — before it comes back. Script the cutaways at the same time as the words.
Should I disclose that my presenter is AI-generated?
Yes. Platform rules increasingly require it for synthetic media, and it costs nothing in credibility when the content itself is useful. A line in the description or a short on-screen note is enough.
Can I reuse the same synthetic presenter across multiple videos?
Yes — keep the source portrait and reuse it. Consistency across a series comes from reusing the same image, not from re-describing the same person in words each time.
Kling AI Team
Kling AI

Written and tested by the Kling AI team behind Kling's video generation and digital-human features.

Your voice, a face that isn't yours

Generate your presenter right now

Describe an original face and write your first script — no photo of yourself required.