1
Return

SR-DAG: a score-based reinforcement learning method for DAG learning

delete2026-03-01
delete0
PRE
AI
Z
Zhou, Yun
H
Hao Zuo *
X
Xuanyi Li
DOI:10.1080/00949655.2026.2647040delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Learning the structure of directed acyclic graphs (DAGs) is the basis for causal discovery. However, directly learning the DAGs from the data is challenging because the graph space increases exponentially with the number of variables. Therefore, people try to model this learning task as a combinatorial optimization problem. The goal is to find the DAG that maximizes the structure score to show how well it fits the data. Recent breakthroughs attempt to solve this problem with neural networks (NNs), which have strong representation power to help explore the graph space, but still have a gap with the optimal solution. Inspired by these methods, we propose a score-based reinforcement learning (RL) model for DAG structure learning. Our approach uses a predefined score function as the reward signal and optimizes the NN parameters using the policy gradient method. Our approach improves the search ability by decomposing the graph search problem into seeking the parent set of each variable. Moreover, with the flexibility of the reward signal of RL, our approach does not rely on the smoothness of score functions. Experimental results on both synthetic and standard benchmark datasets demonstrate that our approach outperforms the state-of-the-art in important metrics for DAG learning.
Keywords:
Directed acyclic graphs
structure learning
reinforcement learning
causal discovery

Journal

J
Journal of Statistical Computation and Simulation
IF:
1.2
Papers:
114
Citations:
4.1K

Organization

N
national university of defense technology - china
Scholars:
1.8W
Papers: 1.4W
Citations: 9
Cited Papers

Cited Papers

Citing Papers

Citing Papers