arrow
Return

Cross-validation and aggregated EM training for robust parameter estimation

delete2008-04-01
delete16
PRE
AI
T
Takahiro Shinozaki *
M
Mari Ostendorf
DOI:10.1016/j.csl.2007.07.005delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
A new maximum likelihood training algorithm is proposed that compensates for weaknesses of the EM algorithm by using cross-validation likelihood in the expectation step to avoid overtraining. By using a set of sufficient statistics associated with a partitioning of the training data, as in parallel EM, the algorithm has the same order of computational requirements as the original EM algorithm. Another variation uses an approximation of bagging to reduce variance in the E-step but at a somewhat higher cost. Analyses using GMMs with artificial data show the proposed algorithms are more robust to overtraining than the conventional EM algorithm. Large vocabulary recognition experiments on Mandarin broadcast news data show that the methods make better use of more parameters and give lower recognition error rates than standard EM training. (C) 2007 Elsevier Ltd. All rights reserved.
Keywords:
EM training
Overtraining
Cross-validation
Sufficient statistics
HMM
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

C
Computer Speech and Language
IF:
3.4
Papers:
1.5K
Citations:
2.6K

Organization

K
Kyoto University
Scholars:
5.1W
Papers: 4.6W
Citations: 6.1W
U
University of Washington
Scholars:
8.0W
Papers: 7.0W
Citations: 12.5W