返回
Stochastic dynamic programming with factored representations
DOI:10.1016/S0004-3702(00)00033-3.png)
摘要
En 中文
Markov decision processes (MDPs) have proven to be popular models for decision-theoretic planning, but standard dynamic programming algorithms for solving MDPs rely on explicit, state-based specifications and computations. To alleviate the combinatorial problems associated with such methods, we propose new representational and computational techniques for MDPs that exploit certain types of problem structure. We use dynamic Bayesian networks (with decision trees representing the local families of conditional probability distributions) to represent stochastic actions in an MDP, together with a decision-tree representation rewards. Based on this representation, we develop versions of standard dynamic programming algorithms that directly manipulate decision-tree representations of policies and value functions. This generally obviates the need for state-by-state computation, aggregating states at the leaves of these trees and requiring computations only for each aggregate state. The key to these algorithms is a decision-theoretic generalization of classic regression analysis, in which we determine the features relevant to predicting expected value. We demonstrate the method empirically on several planning problems, showing significant savings for certain types of domains. We also identify certain classes of problems for which this technique fails to perform well and suggest extensions and related ideas that may prove useful in such circumstances. We also briefly describe an approximation scheme based on this approach. (C) 2000 Elsevier Science B.V. All rights reserved.
Keyword:
decision-theoretic planning
Markov decision processes
Bayesian networks
regression
decision trees
abstraction
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
13.9
论文数:
6.1K
被引数:
1.9W
机构
暂无机构信息
引用论文
The training of verb production in Broca's aphasia: A multiple‐baseline across‐behaviours study
Aphasiology
IF0
Explanation-based learning and reinforcement learning: A unified view基于解释的学习和强化学习: 一个统一的观点
MACHINE LEARNING
IF2.9
Phasic alertness can modulate executive control by enhancing global processing of visual stimuli
Cognition
IF0

