Return
SR-DAG: a score-based reinforcement learning method for DAG learning
Z
H
X
DOI:10.1080/00949655.2026.2647040.png)
Abstract
En 中文
Learning the structure of directed acyclic graphs (DAGs) is the basis for causal discovery. However, directly learning the DAGs from the data is challenging because the graph space increases exponentially with the number of variables. Therefore, people try to model this learning task as a combinatorial optimization problem. The goal is to find the DAG that maximizes the structure score to show how well it fits the data. Recent breakthroughs attempt to solve this problem with neural networks (NNs), which have strong representation power to help explore the graph space, but still have a gap with the optimal solution. Inspired by these methods, we propose a score-based reinforcement learning (RL) model for DAG structure learning. Our approach uses a predefined score function as the reward signal and optimizes the NN parameters using the policy gradient method. Our approach improves the search ability by decomposing the graph search problem into seeking the parent set of each variable. Moreover, with the flexibility of the reward signal of RL, our approach does not rely on the smoothness of score functions. Experimental results on both synthetic and standard benchmark datasets demonstrate that our approach outperforms the state-of-the-art in important metrics for DAG learning.
Keywords:
Directed acyclic graphs
structure learning
reinforcement learning
causal discovery
Journal
J
IF:
1.2
Papers:
114
Citations:
4.1K
