arrow
返回

An empirical framework for detecting speaking modes using ensemble classifier

delete2023-05-13
delete4
PRE
AI
S
Sadia Afroze
M
Mohammed Moshiul Hoque *
M
M. Ali Akber Dewan
DOI:10.1007/s11042-023-15254-8delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Detecting the speaking modes of human is an important cue in many applications, including detecting active/inactive participants in video conferencing, monitoring students' attention in classrooms or online, analyzing students' engagement in live video lectures, and identifying drivers' distractions. However, automatically detecting speaking mode from a video is challenging due to the low resolution of the images, noise, illumination change, and unfavorable viewing conditions. This paper proposes a deep learning-based ensemble technique (called V-ensemble) to identify speaking modes, i.e., talking and non-talking, considering low-resolution and noisy images. This work also introduces an automatic algorithm for the video stream-to-image frame acquisition and develops three datasets for this research (LLLR, YawDD-M, and SBD-M). The proposed system integrated mouth region extraction and mouth state detection modules. A multi-task cascaded neural network (MTCNN) is used to extract the mouth region. Eight popular deep learning approaches, such as ResNet18, ResNet35, ResNet50, VGG16, VGG19, CNN, InceptionV3 and SVM have been investigated to select the best models for the mouth state prediction. Experimental results with a rigorous comparative analysis showed that the proposed ensemble classifier achieved the highest accuracy on three datasets: LLLR (96.80%), YawDD-M (96.69%) and SBD-M (96.90%).
Keyword:
Human computer interaction
Computer vision
Speaking mode detection
Lip motion detection
Ensemble-based classification

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
2.0W
被引数:
3.2W

机构

暂无机构信息
引用论文

引用论文

A parallel grid-search-based SVM optimization algorithm on Spark for passenger hotspot prediction
err2022-03-28
err14
PREAI
errXia, Dawen; Zheng, Yongling; Bai, Yu; Yan, Xiaobo; Hu, Yang; Li, Yantao; Li, Huaqing
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
Lip as biometric and beyond: a survey
err2021-11-24
err15
PREAI
errChowdhury, Debbrota P.; Kumari, Ritu; Bakshi, Sambit; Sahoo, Manmath N.; Das, Abhijit
err分享
err收藏
err分享
err收藏
Recent advances in convolutional neural networks卷积神经网络的最新进展
err2018-05-01
err3.8K
errOAAI
errGu, Jiuxiang; Wang, Zhenhua; Kuen, Jason; Ma, Lianyang; Shahroudy, Amir; Shuai, Bing; Liu, Ting; Wang, Xingxing; Wang, Gang; Cai, Jianfei; Chen, Tsuhan
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容