返回
Generating web-based corpora for video transcripts categorization
DOI:10.1016/j.eswa.2012.07.055.png)
摘要
En 中文
This paper proposes the use of Internet as a rich source of information in order to generate learning corpora for video transcripts categorization systems. Our main goal in this work has been to study the behavior of different learning corpora generated from the Internet and analyze some of their features. Specifically, Wikipedia, Google and the blogosphere have been employed to generate these learning corpora, using the VideoCLEF 2008 track as the evaluation framework for the different experiments carried out. Based on this evaluation framework, we conclude that the proposed approach is a promising strategy for the video classification task using the transcripts of the videos. The different sizes of the corpora generated could lead to believe that better results are achieved when the corpus size is larger, but we demonstrate that this feature may not always be a reliable indicator of the behavior of the learning corpus. The obtained results show that the integration of knowledge from the blogosphere or Google allows generating more reliable corpora for this task than those based on Wikipedia. (C) 2012 Elsevier Ltd. All rights reserved.
Keyword:
Video transcripts categorization
Video tagging
Web-based corpora generation
Automatic Speech Recognition (ASR)
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
Query expansion with a medical ontology to improve a multimodal information retrieval system使用医学本体扩展查询以改善多模态信息检索系统
Transport Properties Investigation of Aqueous Protic Ionic Liquid Solutions through Conductivity, Viscosity, and NMR Self-Diffusion Measurements通过电导率,粘度和NMR自扩散测量研究质子离子液体水溶液的传输特性
没有更多内容

