返回
Merging Web Tables for Relation Extraction With Knowledge Graphs
DOI:10.1109/TKDE.2021.3101479.png)
摘要
En 中文
We propose methods for extracting triples from Wikipedia's HTML tables using a reference knowledge graph. Our methods use a distant-supervision approach to find existing triples in the knowledge graph for pairs of entities on the same row of a table, postulating the corresponding relation for pairs of entities from other rows in the corresponding columns, thus extracting novel candidate triples. Binary classifiers are applied on these candidates to detect correct triples and thus increase the precision of the output triples. We extend this approach with a preliminary step where we first group and merge similar tables, thereafter applying extraction on the larger merged tables. More specifically, we propose an observed schema for individual tables, which is used to group and merge tables. We compare the precision and number of triples extracted with and without table merging, where we show that with merging, we can extract a larger number of triples at a similar precision. Ultimately, from the tables of English Wikipedia, we extract 5.9 million novel and unique triples for Wikidata at an estimated precision of 0.718.
Keyword:
Internet
Encyclopedias
Electronic publishing
Feature extraction
Data mining
Merging
Knowledge engineering
Web tables
relation extraction
information extraction
knowledge graphs
distant supervision
Wikidata
Wikipedia
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Changes in Disparity in County-Level Diagnosed Diabetes Prevalence and Incidence in the United States, between 2004 and 2012
PLOS ONE
IF0
DBpedia - A large-scale, multilingual knowledge base extracted from WikipediaDBpedia-从维基百科中提取的大规模多语言知识库
SEMANTIC WEB
IF2.9

