arrow
返回

On validating web information extraction proposals

delete2022-08-01
delete0
PRE
AI
P
Patricia Jiménez *
R
Rafael Corchuelo
DOI:10.1016/j.eswa.2022.116700delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Many people who have to make informed decisions in today's always-on culture use information extractorsto feed their systems with information that comes from human-friendly documents. Unfortunately, manyproposals that validate information extractors have deficiencies that make it difficult to perform homogeneouscomparisons, confirm or refute performance hypotheses, or draw unbiased conclusions. Consequently, it isvery difficult to select the best-performing proposal on a sound basis. The state-of-the-art validation methodovercomes many deficiencies in the previous proposals, but still overlooks the following issues: completenessof the validation datasets, that is, whether they provide a complete set of annotations or not; structureof the information, that is, whether they check the structure of the record instances extracted or just theattribute instances; and, finally, how extractions and annotations are matched. The decisions made regardingthe previous issues have an impact on the effectiveness results. In this article, we have exhaustively analysedthe literature and we have also highlighted the main weaknesses to tackle. We present a guideline and a methodto compute the effectiveness, which complements and enhances the state-of-the-art validation method.
Keyword:
Web information extractors
Validation method

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

U
University of Sevilla
学者数:
1.9W
论文数: 1.7W
被引数: 15
引用论文

引用论文

err分享
err收藏
Testing the capital structure of Portuguese family businesses
err2021-12-01
err0
errOAAI
errLuciana J. Pestana; Luís Pereira Gomes; Cristina Lopes
err分享
err收藏
学者 查看更多内容