← Back to the Studio

The guide

Five minutes here saves you an hour of retries. Everything runs on your computer. Your voice never leaves it.

1 · Recording a voice (do this right and everything else gets easy)

The app copies whatever you give it. Give it a clean, natural recording and you get a clean, natural voice. Give it noise and you get noise.

How long

Where and how

What to say

Good: 60 seconds on your phone, quiet bedroom, telling the story of your week like you would tell a friend.
Bad: 3 hours of a podcast with intro music, two voices, and a coffee machine in the background.

The trim editor (your safety net)

Right after you upload or record, you see the sound drawn as a waveform with a colored strip under it, second by second: green is clear voice, blue is usable, red is music or noise, gray is silence. Nothing is chosen for you: the wave starts dim, and what ships is what YOU pick.

The checks you will see

CheckWhat it means
LengthHow much usable audio you gave after trimming.
VolumeHow loud your speech is. Quiet is fixable, distorted is not.
ClippingDistortion from recording too loud or too close. Re-record if red.
NoiseThe gap between your voice and the background. Find a quieter room if red.
SpeechWhether I could hear clear words in the best part of your clip.

Fix the transcript (30 seconds, big payoff)

After creating a voice, I show you what I heard in the chosen reference clip. Read it. If one word is wrong, fix it and save. The voice literally reads along with that transcript on every generation, so one wrong word there teaches wrong sounds everywhere.

2 · Writing text that sounds right

Abbreviations and short forms

Your common words (do this once, early)

Every person has 5 to 20 words they say all the time: your name, your brand, your city, your tools. Teach these in the first session. Two ways:

Punctuation and numbers

Length of one generation

3 · How generation works (why this app is different)

4 · Instant clone vs Pro clone (the honest truth about voice cloning)

People ask: ElevenLabs clones a voice in a minute, why is my instant clone not a perfect twin? Here is the honest answer. ElevenLabs runs a giant model on datacenter computers, trained on millions of voices, and even THEIR instant clone is an approximation: their best tier trains on your audio for a while too. An instant clone anywhere carries the accent, tone and pace, not the exact person. To get the exact person, the model has to be TRAINED on that voice. That is the Pro clone, and it runs right here on your own computer.

Instant clonePro clone
Waitabout a minuteone training, on your PC
Sounds likethe same kind of voice: accent, tone, pacethe person. Close to the real voice
Match scoreusually 65 to 8590 plus
Needs20 seconds of clean audio8 to 20 minutes of clean audio
Runs onyour computer, offlineyour computer, offline — your graphics card if you have one

5 · Enhance audio (clean up any recording)

6 · Honest limits (read this so nothing surprises you)

7 · The 5 minute recipe for a new voice

A few clean minutes of audio plus five taught words gives you a voice you can use for months. When you want the perfect twin, press ★ Make Perfect Clone and let it train once.