arrow
返回

Machine learning based sample extraction for automatic speech recognition using dialectal Assamese speech

delete2016-06-01
delete36
PRE
AI
S
Swapna Agarwalla *
K
Kandarpa Kumar Sarma
DOI:10.1016/j.neunet.2015.12.010delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Automatic Speaker Recognition (ASR) and related issues are continuously evolving as inseparable elements of Human Computer Interaction (HCI). With assimilation of emerging concepts like big data and Internet of Things (IoT) as extended elements of HCI, ASR techniques are found to be passing through a paradigm shift. Oflate, learning based techniques have started to receive greater attention from research communities related to ASR owing to the fact that former possess natural ability to mimic biological behavior and that way aids ASR modeling and processing. The current learning based ASR techniques are found to be evolving further with incorporation of big data, IoT like concepts. Here, in this paper, we report certain approaches based on machine learning (ML) used for extraction of relevant samples from big data space and apply them for ASR using certain soft computing techniques for Assamese speech with dialectal variations. A class of ML techniques comprising of the basic Artificial Neural Network (ANN) in feedforward (FF) and Deep Neural Network (DNN) forms using raw speech, extracted features and frequency domain forms are considered. The Multi Layer Perceptron (MLP) is configured with inputs in several forms to learn class information obtained using clustering and manual labeling. DNNs are also used to extract specific sentence types. Initially, from a large storage, relevant samples are selected and assimilated. Next, a few conventional methods are used for feature extraction of a few selected types. The features comprise of both spectral and prosodic types. These are applied to Recurrent Neural Network (RNN) and Fully Focused Time Delay Neural Network (FFTDNN) structures to evaluate their performance in recognizing mood, dialect, speaker and gender variations in dialectal Assamese speech. The system is tested under several background noise conditions by considering the recognition rates (obtained using confusion matrices and manually) and computation time. It is found that the proposed ML based sentence extraction techniques and the composite feature set used with RNN as classifier outperform all other approaches. By using ANNin FF form as feature extractor, the performance of the system is evaluated and a comparison is made. Experimental results show that the application of big data samples has enhanced the learning of the ASR system. Further, the ANN based sample and feature extraction techniques are found to be efficient enough to enable application of ML techniques in big data aspects as part of ASR systems. (C) 2015 Elsevier Ltd. All rights reserved.
Keyword:
Automatic Speech Recognition (ASR)
Artificial Neural Network (ANN)
Multi Layer Perceptron (MLP)
Deep Neural Network (DNN)
Recurrent Neural Network (RNN)
Fully Focused Time Delay Neural Network (FFTDNN)
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neural Networks 封面图
Neural Networks
IF:
6.3
论文数:
8.2K
被引数:
3.0W

机构

G
Gauhati University
学者数:
1.8K
论文数: 1.4K
被引数: 11
引用论文

引用论文

Neurosyphilis Showing Transient Global Amnesia-like Attacks and Magnetic Resonance Imaging Abnormalities Mainly in the Limbic System.
err2001-01-01
err0
PREAI
errHiroshi FUJIMOTO; Toshihiro IMAIZUMI; Yasuko NISHIMURA; Yumiko MIURA; Mitsuyoshi AYABE; Hiroshi SHOJI; Toshi ABE
err分享
err收藏
err分享
err收藏
Natural Radiation in Tenerife (Canary Islands)
err1992-12-01
err0
PREAI
errJ.C. Fernández-Aldecoa; B. Robayna; A. Allende; A. Poffijn; J. Hernández-Armas
err分享
err收藏
Natural radioactivity measurements of beach sands in Gran Canaria, Canary Islands (Spain)
err2013-03-17
err0
PREAI
errM. A. Arnedo; A. Tejera; J. G. Rubiano; H. Alonso; J. M. Gil; R. Rodriguez; P. Martel
err分享
err收藏
Acute pancreatitis associated with lisinopril and olanzapine
err2010-02-01
err0
PREAI
errJesse D. Bracamonte; Mike Underhill; Paul Sarmiento
err分享
err收藏
没有更多内容