🏆 Final results — 11 models tested

Every clone scored by three independent speaker-verification models (Resemblyzer, ECAPA-TDNN, WavLM-SV), each against BOTH her raw and cleaned voicemail.
Scores are shown as a ratio to ceiling — the ceiling being how well her own two recordings match each other. Above 1.00 = matches her better than she matches herself.

1. HER REAL voice — the target

⬇ download

2. 🏆 CONSENSUS WINNER — dots.tts (48kHz)

Only candidate above ceiling on all three metrics (1.008 / 1.116 / 1.029). Highest fidelity output too.

⬇ download

3. GLM-TTS — 2nd place (1.043)

⬇ download

4. PilotTTS — 3rd (best on WavLM-SV)

⬇ download

5. F5-TTS — best on Resemblyzer

⬇ download

📊 Consensus scoreboard

modelresemblecapawavlmconsensus
dots.tts0.9080.7420.9181.051
GLM-TTS0.9110.7050.9421.043
PilotTTS0.9110.6840.9451.033
dots.tts (alt)0.9170.6770.9241.024
F5-TTS0.9390.6590.9151.020
MaskGCT0.9150.6380.9391.010
IndexTTS-20.9160.6110.9100.986
ElevenLabs0.8950.3020.8780.811
CosyVoice 20.9110.563
XTTS-v20.8570.437
Chatterbox0.845
OpenVoice V20.7780.170

Why three metrics

Each scorer picked a different winner — F5 (Resemblyzer), dots.tts (ECAPA), PilotTTS (WavLM). Normalizing to each metric's own ceiling and averaging is what breaks the tie honestly, instead of cherry-picking the flattering scorer.

Notably: ElevenLabs, the commercial market leader, finished near the bottom — open models beat it on this reference.