Supporting material for Benchmarking Voice Cloning on Earnings Calls.

Data and scope

The source dataset is EarningsCallVoice: Core-100, revision v1.0.0. It contains authentic source clips, not the 800 clones. The post reports one listener, 100 calls and eight synthesis settings. We chose the machine methods before the human labeling was complete. We wrote the human analysis plan after collecting the labels and before computing the results; it was not preregistered.

Downloadable results

These downloads contain aggregate results; they do not include the cloned audio or individual listening notes.

Uncertainty

Confidence intervals use 10,000 bootstrap samples of the 100 sets, with random seed 20260925. These intervals describe variation across the sets for one listener. The calls are distinct, but we have not checked whether an executive appears in more than one call.

The 95% confidence interval for the mean rank difference between Qwen 0.6B and IndexTTS2 (Qwen minus IndexTTS2) runs from −0.7 to +0.5 rank positions. It includes zero, so these ratings do not establish a winner between the two.

The authentic recording received 48.7% fractional first-place credit, with a 95% bootstrap interval of 40.4–57.1%.