6 models ยท ~95 samples ยท scored by speaker-embedding similarity.
Every clone scored against BOTH her raw voicemail and the cleaned version, so nothing is tuned to an artifact.
Ceiling = 0.9007 โ how well her own two recordings match each other.
0.9749 vs her real voice โ beats the 0.9007 ceiling by +0.074.
โฌ download| model | vs her real | balanced |
|---|---|---|
| F5-TTS (raw ref) | 0.9749 | 0.9394 |
| MaskGCT (raw ref) | 0.9553 | 0.9153 |
| IndexTTS-2 (raw ref) | 0.9551 | 0.9007 |
| MaskGCT (clean ref) | 0.8931 | 0.9254 |
| IndexTTS-2 (clean ref) | 0.9090 | 0.9155 |
| XTTS-v2 | 0.8105 | 0.8570 |
| Chatterbox | โ | 0.8445 |
| OpenVoice V2 | 0.7589 | 0.7778 |
Also swept: 6 reference variants, 3 enhancement settings, 20-sample best-of runs, EQ tilt sweeps, pitch matching.
MaskGCT (2nd place)
IndexTTS-2
XTTS-v2
OpenVoice V2
โข Reference domain beat model choice. Same model, raw vs enhanced reference = 0.975 vs 0.894.
โข Scoring against only the cleaned file was misleading โ it made an overfit clone look best.
โข Best-of-N is real: seed variance is ยฑ0.015.
โข Denoise-only hurts; bandwidth extension is the useful part.
โข Pitch-shifting always degraded identity.