arrow
Return

Temporal Predictive Coding as World Model for Reinforcement Learning

delete2026-01-01
delete0
PRE
AI
A
Artem Prokhorenko *
P
Petr Kuderov
E
Evgenii Dzhivelikian
A
Aleksandr I. Panov
DOI:10.1007/978-3-032-00800-8_13delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Partially observable environments pose a fundamental challenge for reinforcement learning (RL), requiring agents to infer hidden states from incomplete sensory input. We propose incorporating Temporal Predictive Coding (TPC) as a world model within RL agents to address this problem. By continuously predicting future observations, TPC builds robust latent representations that capture essential state information and temporal dependencies. We evaluate this approach in grid-world environments with varying levels of perceptual ambiguity. Across multiple tasks, TPC-augmented agents consistently outperform or match strong baselines, including LSTM, RWKV, Clone-Structured Cognitive Graphs (CSCG), and episodic control agents. Analysis of the learned representations shows that TPC effectively disentangles underlying state structure, resolving perceptual aliasing and supporting generalization across time. These results demonstrate that TPC enables the formation of stable, predictive internal states, improving both sample efficiency and decision-making under uncertainty. Our findings establish predictive coding as a promising framework for model-based RL in partially observable settings.
Keywords:
Temporal Predictive Coding
Model Based Reinforcement Learning
Partially Observable Environments

Journal

A
ARTIFICIAL GENERAL INTELLIGENCE, AGI 2025, PT II
IF:
0
Papers:
31
Citations:
0

Organization

M
moscow institute of physics & technology
Scholars:
4.5K
Papers: 3.0K
Citations: 3