返回
LEARNING WHILE EXPERIMENTING
DOI:10.1093/ej/uez043.png)
摘要
En 中文
An agent performing risky experimentation can benefit from suspending it to learn directly about the state. 'Positive' information acquisition seeks news that would confirm the state that favours experimentation. It is used as a last-ditch effort when the agent is pessimistic about the risky arm before abandoning it. 'Negative' information acquisition seeks news that would demonstrate that experimentation is futile. It is used as an insurance strategy to avoid wasteful experimentation when the agent is still optimistic. A higher reward from risky experimentation expands the region of beliefs that the agent optimally chooses information acquisition rather than experimentation.
Keyword:
DEVELOPMENT COMPETITION
DYNAMIC ALLOCATION
MODEL
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
5.5K
被引数:
1.6W
机构
暂无机构信息
引用论文
Transcriptional and posttranscriptional regulation of human androgen receptor expression by androgen

