Five minutes here saves you an hour of retries. Everything runs on your computer. Your voice never leaves it.
1 · Recording a voice (do this right and everything else gets easy)
The app copies whatever you give it. Give it a clean, natural recording and you get a clean, natural voice. Give it noise and you get noise.
How long
- 1 to 3 minutes is the sweet spot for an instant clone. 20 seconds is the minimum. More clean talking gives me more good moments to pick from.
- Planning a Pro clone (real training, see section 4)? Record 8 to 20 minutes. That is the single biggest quality lever there is.
- It does not need to be one perfect take. Just keep talking. Anything over 10 minutes is trimmed.
Where and how
- Quiet room. No fan, no AC hum, no traffic, no other people. The app measures this and shows you a Noise check. You want the green chip.
- One person only. A second voice, even far away, confuses the clone.
- No music. Ever. Not even quiet background music.
- Hold the phone or mic about a hand's length from your mouth. Too close pops, too far sounds thin and roomy.
- Avoid empty rooms and bathrooms. Echo gets cloned too.
- Turn OFF noise suppression and enhancement filters if your app has them. They smear the voice. I clean the noise myself, gently.
What to say
- Talk, do not read. Tell a story, describe your day, explain your work. Natural talking carries your real rhythm and melody. Flat reading gives a flat clone.
- Vary it a little: a question, a short sentence, a longer one, a small pause to breathe.
- Speak the way you want the clone to sound. If you whisper, you get a whispering clone.
Good: 60 seconds on your phone, quiet bedroom, telling the story of your week like you would tell a friend.
Bad: 3 hours of a podcast with intro music, two voices, and a coffee machine in the background.
The trim editor (your safety net)
Right after you upload or record, you see the sound drawn as a waveform with a colored strip under it, second by second: green is clear voice, blue is usable, red is music or noise, gray is silence. Nothing is chosen for you: the wave starts dim, and what ships is what YOU pick.
- Drag across the good parts, then press ✓ Use selected. Hold Shift to select more than one part; the total shows as you go, and the meter under the buttons tells you when you have enough.
- ✨ Auto picks for you: full sentences of clear voice only, cut at natural pauses, never half a sentence.
- 🔇 appears ready only when noise was actually measured in your recording, and removes only that. ✦ enhances the sound, checked against your own voice so it can never change who you sound like. Click either again to hear your original.
- ▶ Result plays exactly what the clone will learn from. If it sounds bad, the clone will too.
- A clip from a video with sound effects or a music bed will show a lot of red. Select only the clean talking, and if too little is left, record fresh audio instead. Junk in, junk out. There is no trick around that.
The checks you will see
| Check | What it means |
| Length | How much usable audio you gave after trimming. |
| Volume | How loud your speech is. Quiet is fixable, distorted is not. |
| Clipping | Distortion from recording too loud or too close. Re-record if red. |
| Noise | The gap between your voice and the background. Find a quieter room if red. |
| Speech | Whether I could hear clear words in the best part of your clip. |
Fix the transcript (30 seconds, big payoff)
After creating a voice, I show you what I heard in the chosen reference clip. Read it. If one word is wrong, fix it and save. The voice literally reads along with that transcript on every generation, so one wrong word there teaches wrong sounds everywhere.
2 · Writing text that sounds right
Abbreviations and short forms
- Anything you type in ALL CAPS gets spelled out automatically.
VBA becomes "Vee Bee Aay", HEC becomes "Aitch Ee Cee". Real words in caps (LOVE, YES) are left alone.
- Some are already handled:
SQL is spoken "Sequel", AI is "Aay Eye", ChatGPT is "Chat Jee Pee Tee".
- For anything else, use the My words box on the main page. Type the word, then how it sounds. Example:
Designesh say it like dee zy nesh. Saved once, used forever, per voice.
Your common words (do this once, early)
Every person has 5 to 20 words they say all the time: your name, your brand, your city, your tools. Teach these in the first session. Two ways:
- My words box: spell it the way it sounds. Fast, works for most words.
- Record it: if spelling tricks do not get it right (names are stubborn), select the word after generating, and say it once in the recording panel. The app locks in your real pronunciation and remembers it.
Punctuation and numbers
- Write normal sentences with periods and commas. The voice breathes at commas and rests at periods.
- Question marks give a rising tone. Use them.
- Write numbers the way you want them spoken:
2026 may be read as a number, but "twenty twenty six" is always safe. Same for money and phone numbers.
- Skip emoji and special symbols. They confuse the voice.
Length of one generation
- One to four sentences per Generate gives the best control. You can always generate the next paragraph after.
- Very long text in one go raises the chance of one bad word. That is what the takes and the word checks are for, but shorter is smoother.
4 · Instant clone vs Pro clone (the honest truth about voice cloning)
People ask: ElevenLabs clones a voice in a minute, why is my instant clone not a perfect twin? Here is the honest answer. ElevenLabs runs a giant model on datacenter computers, trained on millions of voices, and even THEIR instant clone is an approximation: their best tier trains on your audio for a while too. An instant clone anywhere carries the accent, tone and pace, not the exact person. To get the exact person, the model has to be TRAINED on that voice. That is the Pro clone, and it runs right here on your own computer.
| Instant clone | Pro clone |
| Wait | about a minute | one training, on your PC |
| Sounds like | the same kind of voice: accent, tone, pace | the person. Close to the real voice |
| Match score | usually 65 to 85 | 90 plus |
| Needs | 20 seconds of clean audio | 8 to 20 minutes of clean audio |
| Runs on | your computer, offline | your computer, offline — your graphics card if you have one |
- Press ★ Make Perfect Clone next to any voice. It shows what it will train on — your NVIDIA graphics card if you have one (fast, under an hour), or your CPU if you do not (slower, a few hours). Then press the button and leave it.
- Nothing leaves your computer. No account, no sign-up, no internet. The training runs on your own machine, start to finish.
- It uses your graphics card fully, so the app pauses making new audio while it trains. When it finishes, the voice switches to its trained Pro version by itself.
- A gaming PC or a laptop with an NVIDIA card trains fast. A plain office PC still works, it just takes longer, so start it before a break.
- This is the exact same training that built the voice this app ships with. It is not a gimmick tier. It is the real thing.