arrow
返回

LMCodec2: Ultra-low bit rate codec with causal multiple transformers

delete2025-03-01
delete0
PRE
AI
D
Dingwei Peng
Q
Qizhen Weng
N
Ningze Zhong
T
Ting Xie
C
Can Gong
X
Xiangwei Zhu
X
Xuelin Yuan
M
Mingjun Ouyang *
DOI:10.1016/j.compeleceng.2024.109960delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In recent years, the bandwidth constraints in satellite Internet of Things (IoT) applications have spurred the development of novel methods for compressing transmitted speech. For satellite voice communications, it is essential to achieve high-quality codecs with a bit rate below 1 kbps, particularly for channels such as Beidou-3, which often operate under such limitations. Neural network-based vocoders have emerged as a promising solution within the AI community, offering high-fidelity audio compression. In this paper, we propose LMCodec2, a causal speech codec designed to operate across a range of bit rates while delivering high-quality audio at extremely low bit rates, specifically tailored for satellite voice transmission. LMCodec2 utilizes a Transformer-based language model to predict tokens frame by frame, achieving a 25 % reduction in bit rate without compromising decoded audio quality. Our experimental evaluations demonstrate that LMCodec2 produces high-quality decoded audio at 0.76 kbps and 1.15 kbps. Notably, at 0.76 kbps, LMCodec2 achieves a MUSHRA (Multi-Stimulus Test with Hidden Reference and Anchor) score that surpasses Encodec's performance at 1.5 kbps. Audio demonstrations, including real-world self-recorded speech datasets, are available at https://dingweipeng.github.io/JACK. github.io. LMCodec2 provides a new way of thinking to addressing the challenges of bandwidth-limited satellite voice communications.
Keyword:
End-to-end codec
VQ-VAE
GAN
Transformer model
Huffman coding

期刊

C
Computers and Electrical Engineering
IF:
4.9
论文数:
6.7K
被引数:
1.3W

机构

S
Sun Yat Sen University
学者数:
9.9W
论文数: 7.2W
被引数: 95
引用论文

引用论文

SoundStream: An End-to-End Neural Audio Codec
err2022-01-01
err178
errOAAI
errZeghidour, Neil; Luebs, Alejandro; Omran, Ahmed; Skoglund, Jan; Tagliasacchi, Marco
err分享
err收藏
Latent-Domain Predictive Neural Speech Coding
err2023-01-01
err9
errOAAI
errJiang, Xue; Peng, Xiulian; Xue, Huaying; Zhang, Yuan; Lu, Yan
err分享
err收藏
AudioLM: A Language Modeling Approach to Audio GenerationAudioLM: 一种音频生成的语言建模方法
err2023-01-01
err79
errOAAI
errBorsos, Zalan; Marinier, Raphael; Vincent, Damien; Kharitonov, Eugene; Pietquin, Olivier; Sharifi, Matt; Roblek, Dominik; Teboul, Olivier; Grangier, David; Tagliasacchi, Marco; Zeghidour, Neil
err分享
err收藏
没有更多内容