arrow
返回

Automatic Disfluency Detection From Untranscribed Speech

delete2024-01-01
delete0
PRE
AI
A
Amrit Romana *
K
Kazuhito Koishida
E
Emily Mower Provost
DOI:10.1109/TASLP.2024.3485465delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Speech disfluencies, such as filled pauses or repetitions, are disruptions in the typical flow of speech. All speakers experience disfluencies at times, and the rate at which we produce disfluencies may be increased by certain speaker or environmental characteristics. Modeling disfluencies has been shown to be useful for a range of downstream tasks, and as a result, disfluency detection has many potential applications. In this work, we investigate language, acoustic, and multimodal methods for frame-level automatic disfluency detection and categorization. Each of these methods relies on audio as an input. First, we evaluate several automatic speech recognition (ASR) systems in terms of their ability to transcribe disfluencies, measured using disfluency error rates. We then use these ASR transcripts as input to a language-based disfluency detection model. We find that disfluency detection performance is largely limited by the quality of transcripts and alignments. We find that an acoustic-based approach that does not require transcription as an intermediate step outperforms the ASR language approach. Finally, we present multimodal architectures which we find improve disfluency detection performance over the unimodal approaches. Ultimately, this work introduces novel approaches for automatic frame-level disfluency and categorization. In the long term, this will help researchers incorporate automatic disfluency detection into a range of applications.
Keyword:
Acoustics
Feature extraction
Location awareness
Speech processing
Bidirectional control
Planning
Encoding
Transformers
Neural networks
Long short term memory
Automatic disfluency detection
multimodal disfluency detection
frame-level disfluency detection
disfluency categorization
automatic speech recognition

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

U
University of Michigan
学者数:
6.4W
论文数: 5.3W
被引数: 124
U
university of michigan system
学者数:
9.1W
论文数: 8.6W
被引数: 133
引用论文

引用论文

Post-artemisinin delayed hemolysis after oral therapy for P. falciparum infection
err2020-01-01
err0
errOAAI
errChristian C. Conlon; Anna Stein; Rhonda E. Colombo; Christina Schofield
err分享
err收藏
err分享
err收藏
“It's almost like they're trying to hide it”: How User-Provided Image Descriptions Have Failed to Make Twitter Accessible
err2019-05-13
err0
errOAAI
errCole Gleason; Patrick Carrington; Cameron Cassidy; Meredith Ringel Morris; Kris M. Kitani; Jeffrey P. Bigham
err分享
err收藏
学者 查看更多内容