[Submitted on 18 Sep 2025 (v1), last revised 21 Jul 2026 (this version, v3)]
Abstract:Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection robustness, existing datasets and methods predominantly rely on single-channel recordings. Moreover, previous studies highlighted that generalization of this attack to new environments is challenging, requiring new methods for generating data encompassing various acoustic conditions. Hence, in this work we introduce an acoustic simulation framework designed to simulate multi-channel replay speech configurations using publicly available resources. Using the framework, we train the state-of-the-art multi-channel replay detector M-ALRAD and evaluate its generalisation on the ReMASC real-recording corpus without any real training data. To improve the exploitation of spatial information, we extend M-ALRAD with inter-channel phase difference features computed for adjacent microphone pairs, augmenting the beamformed representation with directional cues. Synthetic datasets are available at this https URL.
Submission history
From: Michael Neri [view email]
[v1]
Thu, 18 Sep 2025 09:38:58 UTC (309 KB)
[v2]
Fri, 29 May 2026 10:37:14 UTC (520 KB)
[v3]
Tue, 21 Jul 2026 10:44:29 UTC (390 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.