๐Ÿ† Final voice results

6 models ยท ~95 samples ยท scored by speaker-embedding similarity.
Every clone scored against BOTH her raw voicemail and the cleaned version, so nothing is tuned to an artifact.
Ceiling = 0.9007 โ€” how well her own two recordings match each other.

1. HER REAL voice โ€” the target

โฌ‡ download

2. ๐Ÿ† WINNER โ€” F5-TTS from the raw reference

0.9749 vs her real voice โ€” beats the 0.9007 ceiling by +0.074.

โฌ‡ download

๐Ÿ“Š Every model tested

modelvs her realbalanced
F5-TTS (raw ref)0.97490.9394
MaskGCT (raw ref)0.95530.9153
IndexTTS-2 (raw ref)0.95510.9007
MaskGCT (clean ref)0.89310.9254
IndexTTS-2 (clean ref)0.90900.9155
XTTS-v20.81050.8570
Chatterboxโ€”0.8445
OpenVoice V20.75890.7778

Also swept: 6 reference variants, 3 enhancement settings, 20-sample best-of runs, EQ tilt sweeps, pitch matching.

๐Ÿ”Š Compare the other models

MaskGCT (2nd place)

IndexTTS-2

XTTS-v2

OpenVoice V2

What actually moved the needle

โ€ข Reference domain beat model choice. Same model, raw vs enhanced reference = 0.975 vs 0.894.
โ€ข Scoring against only the cleaned file was misleading โ€” it made an overfit clone look best.
โ€ข Best-of-N is real: seed variance is ยฑ0.015.
โ€ข Denoise-only hurts; bandwidth extension is the useful part.
โ€ข Pitch-shifting always degraded identity.