Every clone scored by three independent speaker-verification models (Resemblyzer, ECAPA-TDNN, WavLM-SV), each against BOTH her raw and cleaned voicemail.
Scores are shown as a ratio to ceiling — the ceiling being how well her own two recordings match each other. Above 1.00 = matches her better than she matches herself.
Only candidate above ceiling on all three metrics (1.008 / 1.116 / 1.029). Highest fidelity output too.
⬇ download| model | resembl | ecapa | wavlm | consensus |
|---|---|---|---|---|
| dots.tts | 0.908 | 0.742 | 0.918 | 1.051 |
| GLM-TTS | 0.911 | 0.705 | 0.942 | 1.043 |
| PilotTTS | 0.911 | 0.684 | 0.945 | 1.033 |
| dots.tts (alt) | 0.917 | 0.677 | 0.924 | 1.024 |
| F5-TTS | 0.939 | 0.659 | 0.915 | 1.020 |
| MaskGCT | 0.915 | 0.638 | 0.939 | 1.010 |
| IndexTTS-2 | 0.916 | 0.611 | 0.910 | 0.986 |
| ElevenLabs | 0.895 | 0.302 | 0.878 | 0.811 |
| CosyVoice 2 | 0.911 | 0.563 | — | — |
| XTTS-v2 | 0.857 | 0.437 | — | — |
| Chatterbox | 0.845 | — | — | — |
| OpenVoice V2 | 0.778 | 0.170 | — | — |
Each scorer picked a different winner — F5 (Resemblyzer), dots.tts (ECAPA), PilotTTS (WavLM).
Normalizing to each metric's own ceiling and averaging is what breaks the tie honestly,
instead of cherry-picking the flattering scorer.
Notably: ElevenLabs, the commercial market leader, finished near the bottom —
open models beat it on this reference.