Voice Cloning on Earnings Calls: Analysis Notes
Supporting material for Benchmarking Voice Cloning on Earnings Calls.
Data and scope
The source dataset is EarningsCallVoice: Core-100,
revision v1.0.0. It contains authentic source clips, not the 800 clones.
The post reports one listener, 100 calls and eight synthesis settings. We
chose the machine methods before the human labeling was complete. We wrote
the human analysis plan after collecting the labels and before computing
the results; it was not preregistered.
Downloadable results
- Human rank summary
- Pairwise comparisons
- Machine summary
- Automatic checks
- Model revisions: exact synthesis checkpoints and declared seed.
- Qwen/IndexTTS2 diagnostics: both Qwen sizes and example-level preference agreement summaries.
- Provenance record: checksums linking these files to the labels and audio package used in the analysis.
These downloads contain aggregate results; they do not include the cloned audio or individual listening notes.
Uncertainty
Confidence intervals use 10,000 bootstrap samples of the 100 sets, with random seed 20260925. These intervals describe variation across the sets for one listener. The calls are distinct, but we have not checked whether an executive appears in more than one call.
The 95% confidence interval for the mean rank difference between Qwen 0.6B and IndexTTS2 (Qwen minus IndexTTS2) runs from −0.7 to +0.5 rank positions. It includes zero, so these ratings do not establish a winner between the two.
The authentic recording received 48.7% fractional first-place credit, with a 95% bootstrap interval of 40.4–57.1%.