Return
Low bit-rate speech coding with predictive multi-level vector quantization
DOI:10.1016/j.apacoust.2025.110538.png)
Abstract
En 中文
During the development of modern communication technology, although wideband speech coding can provide high-fidelity speech transmission, its high bandwidth requirements limit its application in resource-constrained environments. Narrowband speech coding still holds research value. However, traditional narrowband low bit- rate speech coding methods usually cannot generate satisfactory speech quality. To address this issue, this paper proposes a narrowband low bit-rate speech coding architecture called PMVQCodec, with the following major improvements. Firstly, we design a predictive multi-level vector quantization (PMVQ) technique, which employs a predictor to effectively capture the correlations between latent frame vectors and combines it with multilevel vector quantization to enhance quantization efficiency. Additionally, we also introduce a full-band feature extractor to effectively reduce the computational complexity. In our experiments, both subjective and objective evaluations demonstrated the effectiveness of the proposed PMVQCodec architecture. Our proposed method can achieve higher quality reconstructed speech than Encodec and HiFiCodec at 1.2 kbps and 2.4 kbps, and even outperforms LyraV2 at 6 kbps.
Keywords:
Speech coding
Predictive multi-level vector quantization
Full-band feature extractor

