返回
Image-based features for speech signal classification
DOI:10.1007/s11042-019-08553-6.png)
摘要
En 中文
Like other applications, under the purview of pattern classification, analyzing speech signals is crucial. People often mix different languages while talking which makes this task complicated. This happens mostly in India, since different languages are used from one state to another. Among many, Southern part of India suffers a lot from this situation, where distinguishing their languages is important. In this paper, we propose image-based features for speech signal classification because it is possible to identify different patterns by visualizing their speech patterns. Modified Mel frequency cepstral coefficient (MFCC) features namely MFCC- Statistics Grade (MFCC-SG) were extracted which were visualized by plotting techniques and thereafter fed to a convolutional neural network. In this study, we used the top 4 languages namely Telugu, Tamil, Malayalam, and Kannada. Experiments were performed on more than 900 hours of data collected from YouTube leading to over 150000 images and the highest accuracy of 94.51% was obtained.
Keyword:
Image-based features
CNN
Speech pattern classification
Language identification
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3
论文数:
2.0W
被引数:
3.2W
机构
引用论文
On the annual and monthly mean diurnal variations of diffuse solar radiation at a meteorological station in west Africa关于西非某气象站散射太阳辐射的年际和月际日变化均值
Inter-Comparisons of Global Surface Albedo and SW Radiation Budgets from Multiple Satellite Missions and Modeling多卫星任务和模型对全球地表反照率与短波辐射预算的相互比较
Changes in Temperature and Precipitation in the Norwegian Arctic during the 20th Century20世纪挪威北极地区的温度和降水变化

