返回
On validating web information extraction proposals
DOI:10.1016/j.eswa.2022.116700.png)
摘要
En 中文
Many people who have to make informed decisions in today's always-on culture use information extractorsto feed their systems with information that comes from human-friendly documents. Unfortunately, manyproposals that validate information extractors have deficiencies that make it difficult to perform homogeneouscomparisons, confirm or refute performance hypotheses, or draw unbiased conclusions. Consequently, it isvery difficult to select the best-performing proposal on a sound basis. The state-of-the-art validation methodovercomes many deficiencies in the previous proposals, but still overlooks the following issues: completenessof the validation datasets, that is, whether they provide a complete set of annotations or not; structureof the information, that is, whether they check the structure of the record instances extracted or just theattribute instances; and, finally, how extractions and annotations are matched. The decisions made regardingthe previous issues have an impact on the effectiveness results. In this article, we have exhaustively analysedthe literature and we have also highlighted the main weaknesses to tackle. We present a guideline and a methodto compute the effectiveness, which complements and enhances the state-of-the-art validation method.
Keyword:
Web information extractors
Validation method
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W

