arrow
Return

Moving Speaker Separation via Parallel Spectral-Spatial Processing

delete2026-01-01
delete0
PRE
AI
Y
Yuzhu Wang *
A
Archontis Politis
K
Konstantinos Drossos
T
T. Virtanen
DOI:10.1109/TASLPRO.2026.3671599delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multi-channel speech separation in dynamic environments is challenging as time-varying spatial and spectral features evolve at different temporal scales. Existing methods typically employ sequential architectures, forcing a single network stream to simultaneously model both feature types, creating an inherent modeling conflict. In this paper, we propose a dual-branch parallel spectral-spatial (PS2) architecture that separately processes spectral and spatial features through parallel streams. The spectral branch uses a bi-directional long short-term memory (BLSTM)-based frequency module, a Mamba-based temporal module, and a self-attention module to model spectral features. The spatial branch employs bi-directional gated recurrent unit (BGRU) networks to process spatial features that encode the evolving geometric relationships between sources and microphones. Features from both branches are integrated through a cross-attention fusion mechanism that adaptively weights their contributions. Experimental results demonstrate that the PS2 outperforms existing state-of-the-art (SOTA) methods by 1.6-2.1 dB in scale-invariant signal-to-distortion ratio (SI-SDR) for moving speaker scenarios, with robust separation quality under different reverberation times (RT60), noise levels, and source movement speeds. Even with fast source movements, the proposed model maintains SI-SDR improvements of over 13 dB. These improvements are consistently observed across multiple datasets, including WHAMR! and our generated WSJ0-Demand-6ch-Move dataset.
Keywords:
Microphones
Array signal processing
Time-frequency analysis
Spectrogram
Convolution
Visualization
Time-domain analysis
Tensors
Speech processing
Bidirectional control
Speech separation
multi-channel
speech enhancement
deep neural network
moving source

Journal

I
IEEE Transactions on Audio Speech and Language Processing
IF:
0
Papers:
151
Citations:
0

Organization

N
nokia corporation
Scholars:
1.8K
Papers: 1.5K
Citations: 1
T
tampere university
Scholars:
2.0K
Papers: 927
Citations: 0