arrow
返回

TranStutter: A Convolution-Free Transformer-Based Deep Learning Method to Classify Stuttered Speech Using 2D Mel-Spectrogram Visualization and Attention-Based Feature Representation

delete2023-09-22
delete2
delete
OA
AI
K
Krishna Basak
N
Nilamadhab Mishra
H
Hsien-Tsung Chang *
DOI:10.3390/s23198033delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Stuttering, a prevalent neurodevelopmental disorder, profoundly affects fluent speech, causing involuntary interruptions and recurrent sound patterns. This study addresses the critical need for the accurate classification of stuttering types. The researchers introduce TranStutter, a pioneering Convolution-free Transformer-based DL model, designed to excel in speech disfluency classification. Unlike conventional methods, TranStutter leverages Multi-Head Self-Attention and Positional Encoding to capture intricate temporal patterns, yielding superior accuracy. In this study, the researchers employed two benchmark datasets: the Stuttering Events in Podcasts Dataset (SEP-28k) and the FluencyBank Interview Subset. SEP-28k comprises 28,177 audio clips from podcasts, meticulously annotated into distinct dysfluent and non-dysfluent labels, including Block (BL), Prolongation (PR), Sound Repetition (SR), Word Repetition (WR), and Interjection (IJ). The FluencyBank subset encompasses 4144 audio clips from 32 People Who Stutter (PWS), providing a diverse set of speech samples. TranStutter's performance was assessed rigorously. On SEP-28k, the model achieved an impressive accuracy of 88.1%. Furthermore, on the FluencyBank dataset, TranStutter demonstrated its efficacy with an accuracy of 80.6%. These results highlight TranStutter's significant potential in revolutionizing the diagnosis and treatment of stuttering, thereby contributing to the evolving landscape of speech pathology and neurodevelopmental research. The innovative integration of Multi-Head Self-Attention and Positional Encoding distinguishes TranStutter, enabling it to discern nuanced disfluencies with unparalleled precision. This novel approach represents a substantial leap forward in the field of speech pathology, promising more accurate diagnostics and targeted interventions for individuals with stuttering disorders.
Keyword:
stuttered speech
speech disfluency
multi-head self-attention
transformer
Mel-Spectrogram
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Sensors 封面图
Sensors
IF:
3.5
论文数:
7.2W
被引数:
20.9W

机构

V
vit bhopal university
学者数:
461
论文数: 431
被引数: 11
C
Chang Gung University
学者数:
1.3W
论文数: 1.2W
被引数: 1.2W
引用论文

引用论文

err分享
err收藏
Duke University's Talent Identification Program
err1982-03-01
err0
PREAI
errRobert N. Sawyer; Lynn M. Daggett
err分享
err收藏
PicoServer
err2008-11-07
err0
PREAI
errTaeho Kgil; Ali Saidi; Nathan Binkert; Steve Reinhardt; Krisztian Flautner; Trevor Mudge
err分享
err收藏
err分享
err收藏
A monotonic statistical machine translation approach to speaking style transformation
err2012-10-01
err6
errOAAI
errNeubig, Graham; Akita, Yuya; Mori, Shinsuke; Kawahara, Tatsuya
err分享
err收藏
学者 查看更多内容