arrow
Return

Generating web-based corpora for video transcripts categorization

delete2013-01-01
delete0
PRE
AI
J
José M. Perea‐Ortega *
A
Arturo Montejo‐Ráez
M
María Teresa Martín Valdivia
L
Luís Alfonso Ureña López
DOI:10.1016/j.eswa.2012.07.055delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper proposes the use of Internet as a rich source of information in order to generate learning corpora for video transcripts categorization systems. Our main goal in this work has been to study the behavior of different learning corpora generated from the Internet and analyze some of their features. Specifically, Wikipedia, Google and the blogosphere have been employed to generate these learning corpora, using the VideoCLEF 2008 track as the evaluation framework for the different experiments carried out. Based on this evaluation framework, we conclude that the proposed approach is a promising strategy for the video classification task using the transcripts of the videos. The different sizes of the corpora generated could lead to believe that better results are achieved when the corpus size is larger, but we demonstrate that this feature may not always be a reliable indicator of the behavior of the learning corpus. The obtained results show that the integration of knowledge from the blogosphere or Google allows generating more reliable corpora for this task than those based on Wikipedia. (C) 2012 Elsevier Ltd. All rights reserved.
Keywords:
Video transcripts categorization
Video tagging
Web-based corpora generation
Automatic Speech Recognition (ASR)
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
3.0W
Citations:
10.2W

Organization

U
universidad de jaen
Scholars:
4.5K
Papers: 4.6K
Citations: 4
Cited Papers

Cited Papers

Using information gain to improve multi-modal information retrieval systems
err2008-05-01
err28
PREAI
errMartin-Valdivia, M. T.; Diaz-Galiano, M. C.; Montejo-Raez, A.; Urena-Lopez, L. A.
errShare
errSave
Query expansion with a medical ontology to improve a multimodal information retrieval system
err2009-04-01
err54
PREAI
errDiaz-Galiano, M. C.; Martin-Valdivia, M. T.; Urena-Lopez, L. A.
errShare
errSave
errShare
errSave
Turning building renovation measures into energy saving opportunities
err2015-05-11
err0
PREAI
errNavid Gohardani; Tord Af Klintberg; Folke Björk
errShare
errSave
no more