Creative Studio
API
Resources
Features
About Us
Download
Avatar · Text-to-Speech

Make a text to speech video with AI

Type it. Hear it said. Kling can turn a written script into a person speaking it on camera — no recording required. Here's a real one, start to finish.

A preset AI avatar sitting at a table, mid-sentence in a text-to-speech generated talking video.
On this page
  1. A typed script, spoken aloud
  2. How to turn text into a talking video
  3. FAQ
The short answer

Type the line you want spoken in quotation marks into VIDEO 3.0 and the clip renders already speaking it — picture and voice in one pass, no recording needed; that is the route the demo below uses. For longer scripts with a preset presenter, a chosen voice and an adjustable speech rate, Kling's separate Avatar (Digital Human) tool converts a typed script to speech — it's a heavier generation, so check the cost shown on the generate button before confirming a long script.

A typed script, spoken aloud

Nothing was recorded here. This clip took the fastest route: the line was typed in quotation marks straight into VIDEO 3.0, and the video rendered already speaking it — picture and voice in one pass. Real, unedited generation, sound on:

The typed script

“Type your script, and the video speaks it for you.”  ·  typed inside quotation marks in the VIDEO 3.0 prompt

Text in, a talking video out. The voice, the pacing and the lip movement were all generated from the quoted line — for longer scripts with a chosen voice, the dedicated Digital Human tool below is built for exactly that.

How to turn text into a talking video

For scripts longer than a line or two, use Kling’s dedicated Avatar (Digital Human) tool — this is its actual interface: the facial image on the left, the Speech box for your script, and the Voice Selection panel with preview, speech rate and emotion:

The Kling Avatar Digital Human tool: facial-image slot, Speech text box with Upload Audio, and the Voice Selection panel with preset voices, speech rate and emotion options.
The Digital Human tool, as it actually looks: face + typed script + a voice picked from the preset library (or your own uploaded audio). (Interface as of August 2026.)
1

Open the Avatar (Digital Human) tool and pick a face

The tool lives under the Avatar tab of the generator — the dedicated Digital Human workflow that turns typed text into a speaking video. Choose a preset from the Avatar Library, upload a facial image, or generate one with AI Image. No filming or recording required either way.

2

Type the script

Write out exactly what you want said, in the Speech box. There's no length limit beyond practicality — write a sentence or a few paragraphs, whatever the video needs.

3

Choose a voice in Voice Selection

Open Voice Selection and pick from the preset library — organised by profession, gender and age, with a play button to preview each voice before committing. Set the speech rate and an emotion (neutral, happy, serious and more), or upload your own recorded audio instead for full control of the voice and accent.

4

Generate and check the read

Kling turns the text into speech and animates the face to match — lips, expression, small natural movement. Longer scripts mean longer, heavier generations, so check the credit cost shown before confirming.

A real generation cost, worth knowing upfront

The Digital Human tool is one of the heavier generations in Kling — it synthesizes speech and animates a full face to match, not just pixels in a scene. Longer scripts mean longer, heavier runs, and the exact cost is shown on the Generate button before you confirm. Check it before committing a long script — and preview your voice in Voice Selection first, so the read matches what you wrote.

Frequently asked questions

How do you add text to speech to a video?
Generate the speech from your script, then pair it with the visual. In Kling the two arrive together: a portrait plus a script produces a presenter already speaking the words, rather than a voice you have to line up afterwards.
How do you add text to speech in YouTube videos?
Same workflow, then export and upload as normal. What matters for YouTube is that the delivery does not sound flat across a long stretch — break the script into shorter takes and vary the pacing.
Does YouTube monetise text-to-speech videos?
YouTube's policies turn on whether the content is original and adds value, not on whether a synthetic voice was used. Low-effort, mass-produced narration is the thing that gets demonetised — check YouTube's current guidance before building a channel on it.
Can I use my own voice instead?
Yes — record the audio and use it to drive the presenter. A real voice with a generated face is a common and often better combination.
Kling AI Team
Kling AI

Written and tested by the Kling AI team, who build the image, video and effects tools behind Kling.

Skip the camera entirely

Type a script now

Write what should be said, put it in quotation marks, and generate a talking video from text.