arrow
返回

Speech-driven talking face using embedded confusable system for real time mobile multimedia

delete2013-08-17
delete4
PRE
AI
A
Anand Paul *
J
Jhing-Fa Wang
陈
陈宜鸿 (Yi‐Hung Chen)
DOI:10.1007/s11042-013-1609-3delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper presents a real-time speech-driven talking face system which provides low computational complexity and smoothly visual sense. A novel embedded confusable system is proposed to generate an efficient phoneme-viseme mapping table which is constructed by phoneme grouping using Houtgast similarity approach based on the results of viseme similarity estimation using histogram distance, according to the concept of viseme visually ambiguous. The generated mapping table can simplify the mapping problem and promote viseme classification accuracy. The implemented real time speech-driven talking face system includes: 1) speech signal processing, including SNR-aware speech enhancement for noise reduction and ICA-based feature set extractions for robust acoustic feature vectors; 2) recognition network processing, HMM and MCSVM are combined as a recognition network approach for phoneme recognition and viseme classification, which HMM is good at dealing with sequential inputs, while MCSVM shows superior performance in classifying with good generalization properties, especially for limited samples. The phoneme-viseme mapping table is used for MCSVM to classify the observation sequence of HMM results, which the viseme class is belong to; 3) visual processing, arranges lip shape image of visemes in time sequence, and presents more authenticity using a dynamic alpha blending with different alpha value settings. Presented by the experiments, the used speech signal processing with noise speech comparing with clean speech, could gain 1.1 % (16.7 % to 15.6 %) and 4.8 % (30.4 % to 35.2 %) accuracy rate improvements in PER and WER, respectively. For viseme classification, the error rate is decreased from 19.22 % to 9.37 %. Last, we simulated a GSM communication between mobile phone and PC for visual quality rating and speech driven feeling using mean opinion score. Therefore, our method reduces the number of visemes and lip shape images by confusable sets and enables real-time operation.
Keyword:
Real-time speech driven
Lip-synch
Talking face
Hidden markov model (HMM)
Multiclass support vector machine (MCSVM)
Viseme histogram similarity
Confusion matrix
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
2.0W
被引数:
3.2W

机构

N
National Cheng Kung University
学者数:
2.6W
论文数: 2.3W
被引数: 1.7W
引用论文

引用论文

Enantiomeric 2-anilino-2-oxo-1,3,2-oxazaphosphorinanes: Synthesis and NMR-investigation of their non-racemic mixtures
err1995-07-01
err0
PREAI
errAndrei B. Ouryupin; Mikhail I. Kadyko; Pavel V. Petrovskii; Erlen I. Fedin; Andrzej Okruszek; Ryszard Kinas; Wojciech J. Stec
err分享
err收藏
NMR Study on Electronic States in Phosphorus Doped Silicon
err1978-10-01
err0
PREAI
errShun-ichi Kobayashi; Yôichi Fukagawa; Seiichiro Ikehata; Wataru Sasaki
err分享
err收藏
Osteoporosis in β‐thalassaemia major patients: analysis of the genetic background
err2008-08-02
err0
PREAI
errSilverio Perrotta; Maria Domenica Cappellini; Francesco Bertoldo; Veronica Servedio; Giovanni Iolascon; Leonardo D'agruma; Paolo Gasparini; Maria Carmen Siciliani; Achille Iolascon
err分享
err收藏
err分享
err收藏
err分享
err收藏
Different Patterns of Glucose Hypometabolism Underlie Functional Decline in Frontotemporal Dementia and Alzheimer’s Disease: FDG-PET Study
err2018-01-01
err0
errOAAI
errMina Fukai; Tetsu Hirosawa; Mitsuru Kikuchi; Shoryoku Hino; Tatsuru Kitamura; Yasuomi Ouchi; Masamichi Yokokura; Etsuji Yoshikawa; Tomoyasu Bunai; Yoshio Minabe
err分享
err收藏
学者 查看更多内容