返回
Structured prediction with reinforcement learning
DOI:10.1007/s10994-009-5140-8.png)
摘要
En 中文
We formalize the problem of Structured Prediction as a Reinforcement Learning task. We first define a Structured Prediction Markov Decision Process (SP-MDP), an instantiation of Markov Decision Processes for Structured Prediction and show that learning an optimal policy for this SP-MDP is equivalent to minimizing the empirical loss. This link between the supervised learning formulation of structured prediction and reinforcement learning (RL) allows us to use approximate RL methods for learning the policy. The proposed model makes weak assumptions both on the nature of the Structured Prediction problem and on the supervision process. It does not make any assumption on the decomposition of loss functions, on data encoding, or on the availability of optimal policies for training. It then allows us to cope with a large range of structured prediction problems. Besides, it scales well and can be used for solving both complex and large-scale real-world problems. We describe two series of experiments. The first one provides an analysis of RL on classical sequence prediction benchmarks and compares our approach with state-of-the-art SP algorithms. The second one introduces a tree transformation problem where most previous models fail. This is a complex instance of the general labeled tree mapping problem. We show that RL exploration is effective and leads to successful results on this challenging task. This is a clear confirmation that RL could be used for large size and complex structured prediction problems.
Keyword:
Structured prediction
Reinforcement learning
Sequence labeling
Tree transformation
HTML to XML
期刊
IF:
2.9
论文数:
2.7K
被引数:
3.4W
机构
引用论文
Learning to match the schemas of data sources: A multistrategy approach学习匹配数据源的模式: 一种多策略方法
MACHINE LEARNING
IF2.9
没有更多内容

