arrow
返回

Incremental Syllable-Context Phonetic Vocoding

delete2015-06-01
delete6
delete
OA
AI
M
Miloš Cerňak *
P
Philip N. Garner
A
Alexandros Lazaridis
P
Petr Motlíček
X
Xingyu Na
DOI:10.1109/TASLP.2015.2418577delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Current very low bit rate speech coders are, due to complexity limitations, designed to work off-line. This paper investigates incremental speech coding that operates real-time and incrementally (i.e., encoded speech depends only on already-uttered speech without the need of future speech information). Since human speech communication is asynchronous (i.e., different information flows being simultaneously processed), we hypothesized that such an incremental speech coder should also operate asynchronously. To accomplish this task, we describe speech coding that reflects the human cortical temporal sampling that packages information into units of different temporal granularity, such as phonemes and syllables, in parallel. More specifically, a phonetic vocoder-cascaded speech recognition and synthesis systems-extended with syllable-based information transmission mechanisms is investigated. There are two main aspects evaluated in this work, the synchronous and asynchronous coding. Synchronous coding refers to the case when the phonetic vocoder and speech generation process depend on the syllable boundaries during encoding and decoding respectively. On the other hand, asynchronous coding refers to the case when the phonetic encoding and speech generation processes are done independently of the syllable boundaries. Our experiments confirmed that the asynchronous incremental speech coding performs better, in terms of intelligibility and overall speech quality, mainly due to better alignment of the segmental and prosodic information. The proposed vocoding operates at an uncompressed bit rate of 213 bits/sec and achieves an average communication delay of 243 ms.
Keyword:
Parametric speech synthesis
very low bit rate speech coding
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

暂无机构信息
引用论文

引用论文

Kinetic study of the initial stages of dehydrogenation of cyclohexane on the platinum(111) surface
err2002-05-01
err0
PREAI
errDeborah Holmes Parker; Claire L. Pettiette-Hall; Yunzhi Li; Robert T. McIver; John C. Hemminger
err分享
err收藏
Importance of hydraulic travel time for the evaluation of organic compounds removal in bank filtration
err2023-03-01
err0
errOAAI
errSebastian Handl; Kaan Georg Kutlucinar; Roza Allabashi; Christina Troyer; Ernest Mayr; Günter Langergraber; Stephan Hann; Reinhard Perfler
err分享
err收藏
err分享
err收藏
Parâmetros morfológicos na avaliação de qualidade de mudas de Eucalyptus grandis形态参数在评估Eucalyptus grandis幼苗质量中的应用
err2002-11-01
err0
errOAAI
errJosé Mauro Gomes; Laércio Couto; Helio Garcia Leite; Aloísio Xavier; Silvana Lages Ribeiro Garcia
err分享
err收藏
学者 查看更多内容