arrow
Return

Speech-Centric Information Processing: An Optimization-Oriented Approach

delete2013-05-01
delete25
delete
OA
AI
X
Xiaodong He *
L
Li Deng
DOI:10.1109/JPROC.2012.2236631delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Automatic speech recognition (ASR) is a central and common component of voice-driven information processing systems in human language technology, including spoken language translation (SLT), spoken language understanding (SLU), voice search, spoken document retrieval, and so on. Interfacing ASR with its downstream text-based processing tasks of translation, understanding, and information retrieval (IR) creates both challenges and opportunities in optimal design of the combined, speech-enabled systems. We present an optimization-oriented statistical framework for the overall system design where the interactions between the subsystems in tandem are fully incorporated and where design consistency is established between the optimization objectives and the end-to-end system performance metrics. Techniques for optimizing such objectives in both the decoding and learning phases of the speech-centric information processing (SCIP) system design are described, in which the uncertainty in speech recognition subsystem's outputs is fully considered and marginalized. This paper provides an overview of the past and current work in this area. Future challenges and new opportunities are also discussed and analyzed.
Keywords:
Joint optimization
speech recognition
speech-centric information processing (SCIP)
spoken language translation (SLT)
spoken language understanding (SLU)
voice search
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Proceedings of the IEEE cover
Proceedings of the IEEE
IF:
25.9
Papers:
9.9K
Citations:
4.5W

Organization

M
Microsoft
Scholars:
3.0K
Papers: 2.7K
Citations: 7
Cited Papers

Cited Papers

errShare
errSave
Margin-Based Discriminative Training for String Recognition
err2010-12-01
err11
PREAI
errHeigold, Georg; Dreuw, Philippe; Hahn, Stefan; Schlueter, Ralf; Ney, Hermann
errShare
errSave
Research Developments and Directions in Speech Recognition and Understanding, Part 1
err2009-05-01
err154
PREAI
errBaker, Janet M.; Deng, Li; Glass, James; Khudanpur, Sanjeev; Lee, Chin-Hui; Morgan, Nelson; O'Shaughnessy, Douglas
errShare
errSave
Spoken language understanding - Interpreting the signs given by a speech signal
err2008-05-01
err74
PREAI
errDe Mori, Renato; Bechet, Frederic; Hakkani-Tuer, Dilek; McTear, Michael; Riccardi, Giuseppe; Tu, Gokhan
errShare
errSave
Near Ambient Pressure XPS at ALBA
err2013-03-22
err0
errOAAI
errV Pérez-Dieste; L Aballe; S Ferrer; J Nicolàs; C Escudero; A Milán; E Pellegrin
errShare
errSave
researcher View more