Return
Knowledge graph modeling for data asset pricing
DOI:10.1016/j.aei.2026.104591.png)
Abstract
En 中文
The valuation of data assets is a fundamental problem in data circulation and data trading systems, yet remains challenging due to heterogeneous descriptions, unclear usage contexts, and the lack of standardized comparability across products. To address this issue, we propose a knowledge-graph-based valuation framework that models data assets and their semantic and relational attributes in a structured graph space. Unstructured product descriptions collected from AWS Data Exchange are processed using a large language model (LLM) ensemble with strict-majority voting, which improves triple extraction stability and mitigates hallucination. The constructed knowledge graph encodes data products, providers, industries, and application concepts. To obtain machine-operable representations, we evaluate two embedding strategies: Node2Vec, which captures global semantic proximity, and GraphSAGE, which preserves localized neighborhood patterns. Experimental results show that Node2Vec achieves superior overall performance, especially in concept-rich sub-graph structures, while GraphSAGE performs better in industry-homogeneous scenarios. Using the learned embeddings, we perform clustering to identify comparable data assets and combine these features with supervised learning models to predict pricing levels. Extensive experiments demonstrate that our framework improves robustness, interpretability, and consistency in data asset valuation.
Keywords:
data asset valuation
knowledge graph
embedding strategies
pricing prediction
data circulation
Journal
IF:
9.9
Papers:
4.4K
Citations:
1.7W
Organization
Cited Papers
CARM: Confidence-aware recommender model via review representation learning and historical rating behavior in the online platforms
NEUROCOMPUTING
IF6.5
Data as asset? The measurement, governance, and valuation of digital personal data by Big Tech
BIG DATA & SOCIETY
IF5.9

