Anatoly Marchenko
2026
How Well Do Commodity Text-to-Speech Systems Evade Acoustic Perturbation Detection? A Multi-Engine Evaluation Across 21 Languages
Anatoly Marchenko
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Anatoly Marchenko
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Jitter, shimmer, and harmonics-to-noise ratio (HNR) are often used to detect voice deepfakes, since these features capture biomechanical irregularities of vocal fold vibration that synthetic speech supposedly lacks. We test this assumption on three commodity TTS engines (Google TTS, Microsoft Edge TTS, macOS system voice) with 1,850 samples across 21 languages, measured against 29 emotion corpora in 24 languages (35,091 utterances). Three classifiers (logistic regression, SVM-RBF, Random Forest) all fail to reliably detect Edge TTS: the best result is F1 = 0.78. Effect sizes drop 2.1x-7.4x from Google TTS to Edge TTS. In ablation, no single feature exceeds F1 = 0.60 against Edge TTS. These three perturbation features, taken alone, can no longer separate commodity neural TTS from natural speech.