XavierFok
← all posts

The voice on my videos is a clone, and the line I hold on it

2026-08-15 · by Xavier Fok

# The voice on my videos is a clone, and the line I hold on it

The narration on my channel comes out of a machine. I recorded a few minutes of myself reading, once, and a model has voiced every script since, in something close enough to my own sound that most listeners never pause on it. If hearing that makes you slightly uneasy, good. That unease is worth keeping, and this piece takes it seriously alongside the practical details.

Why a clone instead of a microphone

For a single video, recording yourself wins on every count. Publishing on a schedule changes the arithmetic.

Consistency drove the decision first. A real voice moves day to day. Tiredness, a cold, a different room, morning energy against late night flatness. The clone sounds identical in every video, which quietly makes a channel feel coherent.

Speed came second. A twelve minute narration generates in under a minute. Recording the same thing cleanly, with the stumbles edited out, used to swallow the better part of an hour.

Editing sealed it. When one sentence changes in a script, the line regenerates in seconds. With a real recording, a one word fix meant setting up the microphone again and hoping the tone matched last week's take. Under a clone, the script is the single source of truth, and everything downstream flows from it.

How mine was made

The process is simpler than people expect, and that simplicity is exactly why the ethics deserve attention.

I read aloud for a few minutes in a quiet room with decent, unremarkable gear. That recording became a reference from which the model captures the character of a voice: the timbre, the rhythm, the way particular sounds land. Since that day there has been no recording at all. Plain text goes in with a pointer at the reference, and audio in my voice comes out.

The amount of source audio required is the part that surprises everyone. No studio hours. A few clean minutes clears the bar for a result most listeners will never question, which is precisely why the tool is powerful and precisely why it deserves care.

The barrier that quietly disappeared

Convincing synthetic speech used to demand hours of studio recording, a real budget, and a specialist team, which kept it inside film and game production. Then models got dramatically better at doing more with less, and the audio requirement collapsed. Hours became minutes. Minutes became, in some systems, seconds.

That collapse is the whole story. A single person at home now does what once took a studio. The technical barrier that used to quietly prevent misuse has simply gone, and the rules people carry in their heads have to catch up in a hurry, because nothing else is standing where that barrier stood.

What it nails, and where it falls down

The overall sound is the strength. Timbre right, general rhythm right, ordinary sentences close to indistinguishable. For narration over footage, with no face on screen, it clears the bar comfortably.

The weaknesses are equally concrete. Emphasis sometimes lands on a word I would have glided past. Very long sentences come out slightly relentless, because the pacing does not breathe like a person. And genuine emotion is beyond it. No real excitement, no knowing pause, none of the things a performance carries.

So the fix is to write for the instrument. Sentences of sensible length. Plain phrasing. Energy carried by the script and the visuals instead of by delivery the model cannot produce. Written that way, the seams mostly vanish.

Where it belongs and where it does not

Faceless narration over visuals is the natural home. Nobody reads lips, the writing carries the energy, and sounding identical across dozens of videos is worth a great deal.

Some uses stay off the table. A heartfelt message to someone I care about should be my actual voice in the actual moment, because the entire value of such a thing is that it is real. Faking a live, off the cuff reaction is out as well, since that would quietly misrepresent what the recording even is. The clone is a production tool for narration. Knowing exactly where that boundary sits is most of using it well.

One document drives everything

A quiet technical benefit turned out to matter more than expected. Because the audio generates directly from the script, the script is the exact record of every spoken word.

Captions pull from that same text rather than from speech recognition guessing at the audio, so spelling holds even on technical terms a transcriber would mangle. Visual timing works by matching phrases from that same script to their position in the audio. One document drives the voice, the captions and the timing of every picture, and an entire category of small errors never gets the chance to exist. That part I did not expect to love. Now it might be my favorite property of the whole setup.

Consent carries all of it

This is my voice. I recorded it, I chose to clone it, and it speaks only on my own channel, saying words I wrote and stand behind. That chain of consent is the entire ethical foundation, and every other consideration sits on top of it.

The same technology, pointed at a few minutes of anyone else's audio, can make that person appear to say things they never said. There is a word for that, and the word is impersonation, with fraud close behind depending on the use. The model has no conscience and no opinion about whose voice it copies. All of the responsibility belongs to the person at the keyboard, and the ease of the tool excuses none of it.

Stated plainly: clone your own voice with your own consent, and never clone anyone else's without theirs. The line is short, and it does not move.

The edge cases people ask about

Cloning your own voice for your own work is fine. Cloning someone's voice with their clear, informed permission, for an agreed purpose, is fine as well, and professional voice work is already heading in that direction. Taking a public figure's voice, or a friend's, from a video online, and generating speech they never agreed to, is out. Audio being easy to find grants no permission to use it.

Disclosure runs softer. I am telling you here, unprompted, that the narration is a clone, because I would rather you know than feel tricked later. No one owes a disclaimer stamped across every frame. What is owed is an honest answer whenever the question comes.

My personal test compresses all of it: would the person whose voice this is be comfortable with exactly this use. Anything short of a clear yes is a no.

The two questions that always come up

Voice work as a livelihood, and the creepiness. Both deserve straight answers.

On the creepiness, my honest position is that the discomfort is healthy and worth keeping. The day this feels completely normal is the day people stop being careful, and careful is what the technology still requires of everyone touching it.

On the livelihood, a tool is no verdict by itself. How it lands depends on how people choose to use it and on what everyone collectively accepts. I can promise only my own conduct: using the tool in a way I can defend out loud, on my own voice, on my own channel, and being honest about what it is.

If you want to try it yourself

Record the reference clean. A quiet room and a decent microphone beat expensive gear, so deal with echo and background hum before anything else. Read naturally, in the tone you want back, because the model learns delivery as much as sound.

Then write for the voice rather than for a performer. Short to medium sentences, plain phrasing, meaning carried by the words. Listen to every full generated track before publishing, since names and technical terms occasionally come out wrong and a spelling tweak in the script usually repairs them. Decide your ethics before the first shortcut tempts you, so the rule already exists at the moment bending it would be easiest. And if anyone asks whether the voice is real, give the honest answer.

The clone lets me publish at a pace hand recording never allowed, with real limits and one non negotiable condition: my voice, my consent, my words. The tools get easier every year, so the honest habits are worth building while they are still cheap. More grounded looks at building with AI live at [xavierfok.com](/).

Get new guides and videos first — join the Telegram channel.