
Training a voice clone means the model's own weights are adjusted to your voice, which lifts the speaker match from roughly 80 percent to about 96 percent. With VoiceClone this runs entirely on your own NVIDIA graphics card, under an hour on a normal card and about 145 minutes on a small 4 GB laptop card, and the trained voice is a file you own.
By Inam, who built VoiceClone and narrates his own videos with it. More of my engineering write ups live at designesh.com. Last updated 4 September 2026.
My instant clone always sounded like a cousin of mine. Close enough that people who know me would nod, far enough that I would wince a little on every listen, because I know exactly where my voice sits and this thing sat two seats away.
The evening I finally trained a proper model of my voice, I pressed the button, went for dinner, came back, generated the same test sentence, and sat there grinning at my own laptop. The cousin was gone. It was me, on a slightly different mic.
That jump is what this article is about, because training is the least understood part of voice cloning and the part most tools quietly do not offer at all.
What does training a voice clone actually mean?
An instant clone never really learns you. The model listens to your reference clip at generation time and imitates it on the fly, the way a good mimic works from a voicemail. That gets you a solid resemblance in about a minute, and it also explains the ceiling: the mimic is still the same generic model underneath, doing an impression.
Training is different machinery. The model's own internal weights are adjusted using your recordings, so your voice stops being an impression it performs and becomes something it knows. The result is a voice model file that exists on your disk, permanently, and loads like any other voice afterward.
The engine underneath VoiceClone is GPT-SoVITS, and its own project page lists these as two separate modes: zero shot from a five second sample, and few shot fine tuning from about a minute of training data. Instant cloning is the first mode. Training is the second.
On my calibrated speaker match scale, my instant clones score around 80 percent against my real voice, and my trained voice scores about 96. The nine year old in my family could not tell the trained one from a real recording, and honestly, on short clips, some days neither can I.
How is a trained clone different from an instant clone?
Same app, same voice, two different machines underneath. These are the differences that actually decide which one you should reach for.
- Instant clone. The model imitates your reference clip at generation time. Ready in about a minute, runs on CPU or GPU, and scores around 80 percent speaker match on my scale.
- Trained clone. The model's own weights are adjusted to your voice and saved as a file. Takes under an hour on a normal NVIDIA card, needs that card, and scores about 96 percent on the same scale.
- Where instant is genuinely better. It costs you a minute instead of an evening, it works on a machine with no NVIDIA card in it at all, and for a draft read or a quick test it is the right tool, not the compromise.
- Where training is genuinely better. Word endings, the weight of pauses, the small habits. Everything that separates "sounds like him" from "that is him".
What do you need before you train a voice at home?
Three things, and I will be exact because vague hardware advice wastes people's money.
First, an NVIDIA graphics card. Training is the one part of this whole field that genuinely wants a GPU, because the training stack runs through CUDA, which is NVIDIA's own GPU interface inside PyTorch. There is no AMD or Intel shortcut around that today.
The good news is how little card it needs. I trained a complete voice on the small 4 GB card inside my work laptop, nothing gaming grade, and it finished in 145 minutes. A normal desktop NVIDIA card does the same run in under an hour.
Before anything starts, VoiceClone shows you which card it found and an honest time estimate for your exact machine, so you commit knowing the price in minutes. I wrote more about the hardware side in GPU or CPU.
Second, clean audio of you. Three to ten minutes of natural talking is the sweet spot, and clean matters more than long, because training studies your recording far more thoroughly than instant cloning does, flaws included. My guide on recording a sample that clones well applies double here.
Third, patience for one evening. Not attention, patience. It runs alone, and after the one time setup it needs no internet at all.
What happens when you press Make it perfect?
In VoiceClone the whole thing is one button called Make it perfect, and I named it that because that is the actual promise. It takes the voice you already cloned, shows you the plan, trains on your card while you live your life, and installs the finished trained voice into your app when it is done.
No account, no cloud, no upload. The training happens on your electricity, and the resulting model belongs to you the way a file belongs to you.
I did not stop at testing it on my own voice. The four ready trained voices included with the license, a British one, an American one, an Indian one, and a Russian one, were all trained with this exact same pipeline, and they score between 60 and 89 percent against their source speakers. I mention this because it is easy to claim a training button works, and harder to ship four voices made by it.
What actually changes in the sound after training?
The measurable part is the speaker match number climbing from 80ish to about 96. The hearable part is more interesting.
Instant clones get the broad shape of a voice but smooth over the small habits, the way certain word endings fall, the exact weight of pauses, the tiny push a person puts on some vowels. Training recovers those, and they turn out to be most of what "sounds like me" means. The synthetic sheen, that faint AI glaze I wrote about here, mostly lives in that gap too.
One honest note: training amplifies the reference, in both directions. Feed it five clean minutes and you get the jump I described. Feed it a noisy phone memo and you get a high fidelity clone of a noisy phone memo.
The quality checks in the app exist exactly to stop that second outcome before it costs you an hour.
FAQ
How long does it take to train a voice clone?
Under an hour on a normal NVIDIA graphics card, and about 145 minutes on a small 4 GB laptop card, measured on my own machine. VoiceClone shows an estimate for your specific GPU before starting.
Can I train a voice clone without a graphics card?
Realistically no, training wants an NVIDIA GPU. Instant cloning and generation still work on CPU only machines, just slower, so you can use the app fully and add training when you have access to a card.
Is my voice uploaded anywhere during training?
Not with VoiceClone. Training runs entirely on your own computer and the finished voice model is a local file you own. Cloud tools train on their servers instead, which means your voice recordings live under their terms.
Is a trained clone really that much better than an instant one?
It is the single biggest quality jump available, bigger than any recording trick. My own voice went from an 80ish percent speaker match to about 96 percent, and the trained version is what narrates my published videos.
Should you train your voice this week?
If you already have the free app, record five clean minutes of yourself, check the quality bars are green, and start Make it perfect right before you eat. The free app includes one complete training run and 3 generations to judge it, so you hear the jump on your own voice instead of taking my word for it.
After those takes, unlimited generation, retraining and exports open with the $49 one time license, and there has never been anything monthly hiding behind it. By the time your tea is done, your voice will be sitting on your disk, trained, and the cousin will be gone.
Hear your own voice come back.
The free Windows app clones and fully trains one voice, yours, on your own PC. No account, no upload, and the download starts right away.