Anatoly Marchenko


2026

Jitter, shimmer, and harmonics-to-noise ratio (HNR) are often used to detect voice deepfakes, since these features capture biomechanical irregularities of vocal fold vibration that synthetic speech supposedly lacks. We test this assumption on three commodity TTS engines (Google TTS, Microsoft Edge TTS, macOS system voice) with 1,850 samples across 21 languages, measured against 29 emotion corpora in 24 languages (35,091 utterances). Three classifiers (logistic regression, SVM-RBF, Random Forest) all fail to reliably detect Edge TTS: the best result is F1 = 0.78. Effect sizes drop 2.1x-7.4x from Google TTS to Edge TTS. In ablation, no single feature exceeds F1 = 0.60 against Edge TTS. These three perturbation features, taken alone, can no longer separate commodity neural TTS from natural speech.
Search
Co-authors
    Venues
    Fix author