arrow
返回

A Lip Reading Method Based on 3D Convolutional Vision Transformer

delete2022-01-01
delete17
PRE
AI
王会娟 封面图
王会娟 (Huijuan Wang) *
G
Gangqiang Pu
T
Ting‐Yu Chen
DOI:10.1109/ACCESS.2022.3193231delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Lip reading has received increasing attention in recent years. It judges the content of speech based on the movement of the speaker's lips. The rapid development of deep learning has promoted progress in lip reading. However, due to lip reading needs to process the information of continuous video frames, it is necessary to consider the correlation information between adjacent images and the correlation between long-distance images. Moreover, lip reading recognition mainly focuses on the subtle changes of lips and their surrounding environment, and it is necessary to extract the subtle features of small-size images. Therefore, the performance of machine lip reading is generally not high, and the research progress is slow. In order to improve the performance of machine lip reading, we propose a lip reading method based on 3D convolutional vision transformer (3DCvT), which combines vision transformer and 3D convolution to extract the spatio-temporal feature of continuous images, and take full advantage of the properties of convolutions and transformers to extract local and global features from continuous images effectively. The extracted features are then sent to a Bidirectional Gated Recurrent Unit (BiGRU) for sequence modeling. We proved the effectiveness of our method on large-scale lip reading datasets LRW and LRW-1000 and achieved state-of-the-art performance.
Keyword:
Feature extraction
Transformers
Lips
Visualization
Speech recognition
Hidden Markov models
Three-dimensional displays
Lip reading
3D convolution
vision transformer
sequence modeling

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

N
north china institute of aerospace engineering
学者数:
800
论文数: 423
被引数: 1
引用论文

引用论文

Solution structure of the Lewis x oligosaccharide determined by NMR spectroscopy and molecular dynamics simulations
err2002-05-01
err0
PREAI
errKristine E. Miller; Chaitali Mukhopadhyay; Perseveranda Cagas; C. Allen Bush
err分享
err收藏
err分享
err收藏
Single-cell RNA sequencing reveals intrinsic and extrinsic regulatory heterogeneity in yeast responding to stress
err2017-12-14
err0
errOAAI
errAudrey P. Gasch; Feiqiao Brian Yu; James Hose; Leah E. Escalante; Mike Place; Rhonda Bacher; Jad Kanbar; Doina Ciobanu; Laura Sandor; Igor V. Grigoriev; Christina Kendziorski; Stephen R. Quake; Megan N. McClean
err分享
err收藏
DBGC: Dimension-Based Generic Convolution Block for Object Recognition
errSENSORS
IF3.5
err2022-02-24
err39
errOAAI
errPatel, Chirag; Bhatt, Dulari; Sharma, Urvashi; Patel, Radhika; Pandya, Sharnil; Modi, Kirit; Cholli, Nagaraj; Patel, Akash; Bhatt, Urvi; Khan, Muhammad Ahmed; Majumdar, Shubhankar; Zuhair, Mohd; Patel, Khushi; Shah, Syed Aziz; Ghayvat, Hemant
err分享
err收藏
Patient-centered risk stratification of disposition outcomes following radical cystectomy
err2016-05-01
err0
PREAI
errJasmir G. Nayak; John L. Gore; Sarah K. Holt; Jonathan L. Wright; Matthew Mossanen; Atreya Dash
err分享
err收藏
Lip Reading-Based User Authentication Through Acoustic Sensing on Smartphones
err2019-02-01
err75
PREAI
errLu, Li; Yu, Jiadi; Chen, Yingying; Liu, Hongbo; Zhu, Yanmin; Kong, Linghe; Li, Minglu
err分享
err收藏
学者 查看更多内容