arrow
Return

WebGCN: Web Information Extraction Algorithm Based on Graph Neural Networks

delete2026-01-01
delete0
PRE
AI
X
Xiaole Wang
D
Dengcheng Yan *
Y
Yuting Wang
Z
Zhang, Heng
X
Xu Wen
F
F. H. Liu
DOI:10.1007/978-981-95-3061-8_5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
E-commerce platforms have complex and dynamically changing webpage structures. Traditional web crawlers and rule-based information extraction methods struggle to adapt to these challenges, resulting in high maintenance costs and low extraction efficiency. To address this issue, this paper proposes a Graph Neural Network (GNN)-based web information extraction algorithm, WebGCN, which effectively leverages the HTML structure by integrating it into the web document representation and incorporates graph attention, sparse attention, and local attention mechanisms to reduce global computational complexity while enhancing extraction accuracy. Experiments on multiple real-world e-commerce datasets demonstrate that WebGCN achieves state-of-the-art performance.
Keywords:
E-commerce Webpages
Information Extraction
Graph Neural Networks

Journal

K
KNOWLEDGE SCIENCE, ENGINEERING AND MANAGEMENT, KSEM 2025, PT V
IF:
0
Papers:
29
Citations:
0

Organization

A
anhui university
Scholars:
1.9W
Papers: 1.2W
Citations: 24