返回
Learning to adapt cross language information extraction wrapper
DOI:10.1007/s10489-011-0305-0.png)
摘要
En 中文
We propose a framework for adapting a previously learned wrapper from a source Web site to unseen sites in different languages. To achieve this, we exploit the previously learned information extraction knowledge and the previously extracted or collected items in the source Web site. These knowledge and data are automatically translated to the same language as the unseen sites via online Web resources such as online Web dictionaries or maps. Site independent features which capture the characteristics of the content of the data are then derived from the translated information. Several text mining methods are employed to automatically discover a set of machine labeled training examples in the unseen site. Both content oriented features and site dependent features of the machine labeled training examples are used for learning the new wrapper for the new unseen site using our language independent wrapper induction component. We conducted experiments on some real-world Web sites in different languages to demonstrate the effectiveness of our framework.
Keyword:
Web applications
Information extraction
Web mining
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
暂无机构信息
引用论文
Alaskan Lake Sediment Records and Their Implications for the Beringian Standstill Hypothesis
PaleoAmerica
IF0

