Return
LEARNING WHILE EXPERIMENTING
DOI:10.1093/ej/uez043.png)
Abstract
En 中文
An agent performing risky experimentation can benefit from suspending it to learn directly about the state. 'Positive' information acquisition seeks news that would confirm the state that favours experimentation. It is used as a last-ditch effort when the agent is pessimistic about the risky arm before abandoning it. 'Negative' information acquisition seeks news that would demonstrate that experimentation is futile. It is used as an insurance strategy to avoid wasteful experimentation when the agent is still optimistic. A higher reward from risky experimentation expands the region of beliefs that the agent optimally chooses information acquisition rather than experimentation.
Keywords:
DEVELOPMENT COMPETITION
DYNAMIC ALLOCATION
MODEL
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.6
Papers:
5.5K
Citations:
1.6W
Organization
No organization information available
Cited Papers
Transcriptional and posttranscriptional regulation of human androgen receptor expression by androgen

