arrow
Return

Refining Evaluation Functions for Game 2048 by Extended Temporal Difference Learning

delete2026-03-01
delete0
PRE
AI
W
Wang, Weikai *
M
Matsuzaki, Kiminori
DOI:10.1109/TG.2025.3606517delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The game 2048 has attracted millions of people with its simple yet challenging gameplay, leading to the development of numerous computer players. Most successful computer players for 2048 use evaluation functions trained through temporal difference learning (TD learning) or its variants. While TD learning is highly effective and can improve evaluation functions quickly, the performance of these functions often plateaus after a certain number of timesteps. Therefore, it is important to refine those evaluation functions to further enhance the performance of computer players. In this article, we extend the conventional TD learning approach and propose two refinement algorithms for 2048. First, we conducted detailed experiments to refine the best open-source neural network, and achieved significant performance improvements, increasing the average score from 2.49 & times;10(5) to 3.37 & times;10(5) in greedy play (1-ply lookahead) and from 4.87 & times;10(5) to 5.45 & times;10(5) with 3-ply expectimax search. We also applied our refinement method to the state-of-the-art N-tuple network, improving the average score from 5.85 & times;10(5) to 6.10 & times;10(5) with 6-ply expectimax search and the tile-downgrading trick.
Keywords:
Games
Temporal difference learning
Training
Neural networks
Artificial intelligence
Data mining
Vectors
Refining
Convolutional neural networks
Video games
Game 2048
n-tuple network
neural network
temporal difference (TD) learning

Journal

I
IEEE Transactions on Games
IF:
2.8
Papers:
45
Citations:
0

Organization

K
kochi university technology
Scholars:
696
Papers: 749
Citations: 4