arrow
返回

Reinforcement Learning Trees

delete2016-01-15
delete139
delete
OA
AI
R
Ruoqing Zhu *
Donglin Zeng 封面图
Donglin Zeng (Donglin Zeng)
M
Michael R. Kosorok
DOI:10.1080/01621459.2015.1036994delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
In this article, we introduce a new type of tree-based method, reinforcement learning trees (RLT), which exhibits significantly improved performance over traditional methods such as random forests (Breiman 2001) under high-dimensional settings. The innovations are threefold. First, the new method implements reinforcement learning at each selection of a splitting variable during the tree construction processes. By splitting on the variable that brings the greatest future improvement in later splits, rather than choosing the one with largest marginal effect from the immediate split, the constructed tree uses the available samples in a more efficient way. Moreover, such an approach enables linear combination cuts at little extra computational cost. Second, we propose a variable muting procedure that progressively eliminates noise variables during the construction of each individual tree. The muting procedure also takes advantage of reinforcement learning and prevents noise variables from being considered in the search for splitting rules, so that toward terminal nodes, where the sample size is small, the splitting rules are still constructed from only strong variables. Last, we investigate asymptotic properties of the proposed method under basic assumptions and discuss rationale in general settings. Supplementary materials for this article are available online.
Keyword:
Consistency
Error bound
Random forests
Reinforcement learning
Trees
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

J
Journal of the American Statistical Association
IF:
3
论文数:
5.2K
被引数:
4.8W

机构

U
university of north carolina
学者数:
7.4W
论文数: 6.5W
被引数: 93
引用论文

引用论文

err分享
err收藏
Workplace Aerosol Measurement
err2011-07-07
err0
PREAI
errJon C. Volkwein; Andrew D. Maynard; Martin Harper
err分享
err收藏
Bagging predictorsBagging预测器
err1996-08-01
err1.0W
PREAI
errBreiman, L
err分享
err收藏
学者 查看更多内容