arrow
返回

Trinity: On Using Trinary Trees for Unsupervised Web Data Extraction

delete2014-06-01
delete36
delete
OA
AI
R
Rafael Corchuelo
DOI:10.1109/TKDE.2013.161delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Web data extractors are used to extract data from web documents in order to feed automated processes. In this article, we propose a technique that works on two or more web documents generated by the same server-side template and learns a regular expression that models it and can later be used to extract data from similar documents. The technique builds on the hypothesis that the template introduces some shared patterns that do not provide any relevant data and can thus be ignored. We have evaluated and compared our technique to others in the literature on a large collection of web documents; our results demonstrate that our proposal performs better than the others and that input errors do not have a negative impact on its effectiveness; furthermore, its efficiency can be easily boosted by means of a couple of parameters, without sacrificing its effectiveness.
Keyword:
Web data extraction
automatic wrapper generation
wrappers
unsupervised learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Knowledge and Data Engineering 封面图
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
论文数:
6.8K
被引数:
3.2W

机构

U
University of Sevilla
学者数:
1.9W
论文数: 1.7W
被引数: 15
引用论文

引用论文

err分享
err收藏
Peritoneal benign cystic mesothelioma: a case report and review of the literature
err2002-03-01
err0
PREAI
errS. van Ruth; M.W.G.A. Bronkhorst; F. van Coevorden; F.A.N. Zoetmulder
err分享
err收藏
err分享
err收藏
学者 查看更多内容