Benchmarking Voice Cloning on Earnings Calls

Side-by-side illustrations of a researcher ranking nine recordings through headphones and an executive seated alongside eight voice-cloned doubles

Eight models, 100 earnings calls, and a comparison of human judgments with automatic scores.

tl;dr: We generated 800 clones from 100 earnings-call answers. In my blind listening comparison, Qwen and IndexTTS2 led the clones, and I ranked a clone above the authentic recording in 36 of the 100 sets. Automatic scores often disagreed with my judgments. Next, we will try changing the delivery while keeping the words and apparent speaker fixed, and test whether this changes the interpretation of an answer.

My voice is my password.

Banks have asked customers to say these words to access telephone banking. HSBC described this process in 2020: after entering their account details, customers spoke the phrase and the system recognized their voice. The attraction is clear. There is no complicated password to remember; the voice identifies the person.

But if a model can reproduce that voice from a short recording, how secure is that identification?

Answering this requires testing the bank’s actual authentication system, including its additional checks. In this blog, we do not attempt to break into bank accounts. Instead, we investigated: how well can we clone an executive’s voice?

We therefore start with a horse race: given the same short reference and the same text, which voice-cloning models produce the most convincing clones? When those clones are mixed with the real recording, can a human listener and automatic methods still distinguish the original?

I ranked all 900 answers (800 clones and 100 authentic recordings) without seeing the candidate identities. Seven automatic methods also ranked the candidates.

These rankings show how convincing baseline cloning can sound. My listening judgments combine resemblance to the speaker with naturalness and recording cues. Different automatic methods score speaker similarity, signs of synthetic speech, recording characteristics or audio quality. Those criteria need not favour the same clone, as we will see below.

The longer-term motivation extends beyond impersonation. If ordinary cloning is convincing enough, controlled synthesis may let us construct alternative versions of the same utterance for social-science experiments: the same words, apparently the same person, but a different delivery. Would that change a human opinion or decision, or a machine’s sentiment or credibility assessment? Those questions matter in finance, politics, management and relationships. We will explore delivery changes in a later post. First, though, we need to know how well the voice itself survives cloning. If the baseline clone sounds robotic or like a different person, those differences could themselves change the listener’s response. We need convincing clones to reduce these confounders before trying to isolate the effect of delivery.

This is Post 3 of the earnings-call audio series. Post 1 introduced the research question; Post 2 described the work needed to build the source dataset.

N.B. This is independent academic research conducted with my colleague Hamdan Al Ahbabi in the context of his PhD thesis on multimodal AI at Khalifa University, under my co-supervision. The research is ongoing.

The experiment

We used the frozen v1.0.0 release of EarningsCallVoice: Core-100. Each item supplies a voice reference from prepared remarks, an analyst’s question, and the executive’s recorded answer from later in the call. There is one item per call. References average 12.7 seconds; recorded answers average 17.6 seconds.

Each cloner received the reference and the same answer text, with the reference transcript where its interface required it. The authentic answer waveform was held out as the comparison. We used the ordinary cloning configuration, with no experimental instruction to sound more confident, hesitant or emotional. The models could still produce their own delivery differences; we were not yet trying to manipulate those differences. These are the eight frozen systems:

System Version tested
Qwen3-TTS 12Hz 0.6B Base
Qwen3-TTS 12Hz 1.7B Base
IndexTTS2 IndexTTS-2
CosyVoice3 Fun-CosyVoice3-0.5B-2512
Chatterbox English 0.5B
Higgs Audio V2 generation-3B-base
F5-TTS F5TTS v1 Base
OpenVoice V2 MeloTTS speech converted to the reference voice

The results apply to these model versions and generation settings. OpenVoice also takes an extra step: MeloTTS first reads the answer in one of its built-in voices, then OpenVoice converts that speech to resemble the executive’s voice in the reference. Its result therefore depends on both stages.

We evaluated one output per model and item. Playable failures were included in the comparison. The analysis notes and downloadable results provide the full tables and model revisions.

For listening and machine ranking, all files were converted to mono, 16 kHz, PCM-16 WAV. This removed clues from file formats and sample rates. We preserved loudness, pauses and background noise, which matters for interpreting the results.

Listening to 900 answers

For each set, I listened to the reference, read the analyst’s question, and compared nine anonymous candidates, A through I. One was authentic and eight were generated. Their identities and order were randomized separately for each set.

The question was: which answers sound most likely to have been spoken by the person in the reference? I could assign ties. The answer transcript was hidden, and I did not separately score word accuracy or naturalness.

In practice, I also considered how natural the speech sounded, whether it contained robotic artifacts, and whether background sounds suggested a real recording. I was not always conscious of how much weight I gave those cues. My ranking therefore combines speaker resemblance, naturalness and recording plausibility in one overall listening judgment.

Having several people label the same sets would help us measure how much listeners agree. Other listeners may notice different cues or give them different weight.

The package contained about 4.2 hours of audio before any replay. I spent several weekends comparing candidates, returning to the reference and assigning ranks: roughly 20 hours of careful listening and labeling, by my estimate.

Which clones were most convincing to the listener?

Human mean listening ranks, combining speaker resemblance with naturalness and recording cues, alongside automatic content-check pass counts

The authentic recording had the best average rank. Among the clones, the two Qwen sizes and IndexTTS2 led the comparison.

Source Mean rank, lower is better In the top tier, out of 100 Best-clone credit, out of 100
Authentic recording 2.2 64 —
Qwen3-TTS 0.6B 4.1 19 22.6
IndexTTS2 4.1 20 19.8
Qwen3-TTS 1.7B 4.3 15 18.8
CosyVoice3 4.7 7 7.9
Chatterbox 4.9 12 11.6
Higgs Audio V2 5.5 11 13.7
F5-TTS 6.4 3 5.7
OpenVoice V2 8.8 0 0.0

Mean ranks include all nine candidates. A tie occupying positions 2, 3 and 4 gives each candidate rank 3. Top-tier counts include ties and therefore need not sum to 100. Best-clone credit removes the authentic answer, finds the best remaining candidate or candidates, and splits one point equally between ties.

Qwen 0.6B and IndexTTS2 were close: comparing their clones of each answer, I ranked Qwen higher in 45 sets, IndexTTS2 higher in 40, and tied them in 15. These ratings do not establish a clear winner between the two. More listeners would help us see whether this close result holds beyond my own preferences.

The larger Qwen model did not improve my average listening rank in this configuration. Higgs occasionally produced my preferred clone but was less consistent across sets. Several Higgs clips sounded very realistic at the beginning, but the speech grew progressively quieter and slower, becoming almost inaudible toward the end. A similar problem has been reported by another user. CosyVoice had a better average rank than Chatterbox, while Chatterbox appeared in the top tier more often.

Could a clone outrank the real recording?

The authentic recording was:

  • alone in first place in 36 sets;
  • tied for first place in 28 sets;
  • below the top tier in 36 sets.

Thus, it appeared in the top tier in 64/100 sets. If a two-way tie gives half a point, a three-way tie one third, and so on, the result becomes 48.7% fractional first-place credit. Uniformly choosing one of nine candidates would yield 11.1% on average.

I found a clone more convincing than the original in 36 sets, and tied a clone with the original at the top in another 28. My task was to rank how likely each answer sounded to come from the reference speaker. Measuring fake detection would require a different task: play a recording and ask, “Is this real or synthetic?” With several listeners, we could then measure how often clones are mistaken for real recordings, and vice versa.

I sometimes used laughter, background noise, a long pause or an imperfect cut as clues to the original recording. In another set, I noted that none of the candidates clearly resembled the reference. Background sounds and the way a clip begins or ends can therefore influence the ranking, alongside the voice itself.

What did the machines hear?

We chose seven automatic methods before the human labeling was complete, and also tested three combinations of their ranks. No method was trained or tuned on my ratings.

Four methods estimate how similar two voices are: Resemblyzer, CAMPPlus, WavLM SV and UniSpeech-SAT SV. Each turns a recording into a numerical description of the voice, called a speaker embedding. Comparing the answer’s embedding with the reference’s gives a speaker-similarity score.

We also used AASIST, which tries to distinguish real from synthetic speech; MFCC similarity, which compares sound characteristics and is sensitive to recording conditions; and DNSMOS P.835, which predicts how people would rate the audio quality.

Heatmap of mean ranks: human speaker-likelihood judgments, four speaker encoders, AASIST, MFCC channel similarity and DNSMOS disagree on several models

CAMPPlus gave IndexTTS2 a mean rank of 1.9, ahead of even the authentic recording at 2.8. It put the two Qwen models much farther down than I did. WavLM’s rankings were more favorable to Qwen. Resemblyzer ranked F5-TTS relatively highly despite its weak human ranking and content problems.

OpenVoice V2 had the best mean DNSMOS rank, 2.5, while my mean listening rank for it was 8.8. The authentic recording had the worst mean DNSMOS rank, 8.1, and never came first. Audio can sound clean while sounding unlike the person we are trying to clone. DNSMOS assesses audio quality, including noise and distortion; it was not designed to identify the speaker.

We also checked how often each method ranked the authentic recording first:

Ranking method What its score targets Authentic first-place credit
Human Speaker likelihood, including naturalness and recording cues 48.7%
CAMPPlus Speaker similarity 29.0%
UniSpeech-SAT SV Speaker similarity 16.0%
WavLM SV Speaker similarity 8.0%
Resemblyzer Speaker similarity 6.0%
AASIST Real-versus-synthetic speech 9.0%
MFCC similarity Voice/channel similarity 8.0%
DNSMOS Perceptual audio quality 0.0%
Identity ensemble Four speaker-method ranks 10.7%
Authenticity ensemble AASIST and MFCC ranks 10.0%
Balanced ensemble Identity and authenticity ensemble ranks 13.6%

Ties share first-place credit, as in the human comparison. Each combination averages its input ranks with equal weights.

AASIST placed a clone above the authentic recording in 91/100 sets. We used its released ASVspoof-2019-LA checkpoint: its training data differs from our earnings calls and predates these cloners. A newer detector could give different results. For the speaker-similarity methods, a high score for a clone is less surprising: reproducing the speaker’s voice is exactly what the cloner is trying to do.

Do any machine methods agree with the human?

The speaker encoders show moderate agreement with my rankings when all candidates are included. CAMPPlus agrees most. However, all four encoders and I consistently put OpenVoice near the bottom. That shared judgment can make agreement look stronger than it is among the other clones.

Speaker method All nine candidates Eight clones only Seven clones, excluding OpenVoice
CAMPPlus 0.47 0.44 0.18
WavLM SV 0.36 0.42 0.16
Resemblyzer 0.33 0.40 0.12
UniSpeech-SAT SV 0.34 0.36 0.05

For each set, we computed the Spearman correlation between my ranking and the machine’s ranking, then averaged across the 100 sets. A correlation of 1 means identical ordering; 0 means no association between the ranks. We added the comparison excluding OpenVoice after seeing how often it came last.

Once the authentic recording and OpenVoice are removed, even CAMPPlus falls to 0.18. The agreement is much weaker among the remaining clones. I would therefore still listen to the outputs when choosing a cloner: a high similarity score from CAMPPlus is only part of what makes a clone convincing.

AASIST shows essentially no agreement with my ordering (−0.05). DNSMOS goes somewhat in the opposite direction (−0.26), which fits its preference for OpenVoice and low ranking of the authentic recordings.

Why does CAMPPlus favor IndexTTS2 so strongly?

Looking at Qwen 0.6B and IndexTTS2 directly makes the disagreement clearer:

Judge Qwen 0.6B preferred Tied IndexTTS2 preferred
Human 45 15 40
CAMPPlus 6 0 94
Resemblyzer 11 0 89
WavLM SV 51 0 49
UniSpeech-SAT SV 45 0 55

WavLM is nearly evenly split. However, it matches my preference on only 44 of the 85 sets where I did not tie the two clones. Similar totals can hide disagreement on individual examples.

IndexTTS2 itself uses CAMPPlus to describe the reference voice before generating speech. Our logs confirm that we used the same version of CAMPPlus to evaluate the clones. This may favor IndexTTS2: the same model helps generate the voice and then judges how closely it matches the reference. Resemblyzer also favors IndexTTS2, so this may only be part of the explanation.

I was also listening for things these methods may miss. A clone can sound like the right person but have awkward pauses or a robotic delivery. I may prefer a slightly less similar voice that sounds more natural. We would need separate human ratings of voice similarity and naturalness to check this explanation. DNSMOS alone cannot answer it: it was developed to evaluate speech quality in noise-suppression systems, and may miss the synthetic-speech problems that bothered me.

It would also be useful to compare voices with other models, such as SpeechBrain’s ECAPA-TDNN and ERes2Net. We did not run them here. Comparing their scores with new listeners’ ratings would help us see whether the Qwen/IndexTTS2 disagreement persists.

Did the clones say the right words?

For later experiments, the clone also needs to say the right words. We used automatic speech recognition to transcribe each answer and compare it with the intended text. We checked for missing beginnings or endings and repeated phrases. Other checks looked for distortion from clipping, long silences, unusual duration, and changes in the voice between sections of a clip.

Source Content checks passed All automatic checks passed
Authentic recording 97/100 94/100
Qwen3-TTS 0.6B 98/100 94/100
IndexTTS2 100/100 99/100
Qwen3-TTS 1.7B 100/100 98/100
CosyVoice3 97/100 97/100
Chatterbox 100/100 100/100
Higgs Audio V2 96/100 94/100
F5-TTS 66/100 65/100
OpenVoice V2 95/100 95/100

These counts come from software checks. A failed check means the software found a possible problem; we need to listen to the clip to confirm it. A transcription error can make a correctly spoken answer fail the content check; even three authentic answers failed it. During labeling, I also heard mispronounced company names and financial acronyms such as EPS and EBITDA. After text-to-speech and speech recognition, the transcript can therefore differ from the input because of a pronunciation error, a recognition error, or both. OpenVoice usually passed these checks while sounding unlike the reference to me.

F5-TTS illustrates the problem particularly well. Resemblyzer gave it a mean rank of 3.4 for voice similarity, but 34 of its 100 outputs failed the word checks. We would need to listen to those clips before using them in an experiment that is supposed to keep the text fixed.

More listeners

You can now try the listening app or download the offline package. Listen to the same anonymous candidates, save your progress, and export your rankings after any complete set. The interface groups the work into sessions of five sets; there is no need to finish all 100 to contribute. Click Submit labels to send your rankings privately to us through the app. No email address or account is needed. You can also download a CSV or JSON backup and continue labeling later.

With more listeners, we can measure how much people agree and whether their preferred models match mine or the machines’.

A further experiment could ask separately about resemblance to the reference voice, naturalness, and which recording the listener thinks is real. This would help us understand why Qwen and IndexTTS2 receive different rankings.

From a convincing voice to a causal experiment

Qwen and IndexTTS2 look like reasonable candidates for experiments on delivery. For example, we could generate a confident and a hesitant version of the same answer, then check that the difference is audible, that both versions sound like the same person and say the same words, and that neither contains obvious audio problems.

We can start with audio models, without recruiting a large pool of participants. We can give a model each version in a separate run, with the same analyst’s question, surrounding information and evaluation prompt. Does it rate the answer as more uncertain, less credible or less responsive when the delivery changes? Repeating each condition would let us compare the effect of delivery with variation between runs.

A human experiment could follow, randomly assigning participants to hear one version. A survey could measure perceived trust or stated intentions; a decision task, such as allocating a hypothetical investment, could measure choices within the experiment.

The difficulty is changing only what we want to study. If the hesitant version also sounds robotic or like a different person, we cannot tell which difference affected the response. An earnings call happens once: we cannot observe how its audience would have reacted to the same answer delivered differently. Synthesis lets us construct those alternative versions and measure responses under controlled conditions. The resulting experiment would tell us how delivery affects the tested models or listeners; it would not establish how the historical call, or its market impact, would have changed.

We still need to check how well each model lets us control delivery. The Qwen Base interface we tested provides voice cloning; its separate model variants provide other forms of voice and instruction control. IndexTTS2 exposes separate emotion controls, which we can test by listening to the resulting clips and checking whether the requested emotion comes through.

Conclusion and next steps

A short voice reference was enough to produce clones that sometimes sounded more real to me than the authentic recording. Qwen and IndexTTS2 were the strongest candidates in this comparison. The automatic methods often disagreed with my rankings, and good voice scores did not guarantee that a clone said the right words. Choosing material for the next experiment will require both listening and content checks.

“My voice is my password” relies on a voice being distinctive and difficult to reproduce. These results show why sounding like someone deserves less trust as evidence of identity, even though we have not tested a bank’s authentication system.

For the next experiment, I would start with Qwen and IndexTTS2. Can we make the same answer sound confident or hesitant, while still sounding like the same executive? If so, does an audio model judge the answer differently? We can test this first, then run a similar experiment with human listeners.