arrow
Return

Using temporal-difference learning for multi-agent bargaining

delete2008-12-01
delete4
PRE
AI
S
Shiu‐Li Huang *
F
Fu‐ren Lin
DOI:10.1016/j.elerap.2007.04.001delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This research treats a bargaining process as a Markov decision process, in which a bargaining agent's goal is to learn the optimal policy that maximizes the total rewards it receives over the process. Reinforcement learning is an effective method for agents to learn how to determine actions for any time steps in a Markov decision process. Temporal-difference (TD) learning is a fundamental method for solving the reinforcement learning problem, and it can tackle the temporal credit assignment problem. This research designs agents that apply TD-based reinforcement learning to deal with online bilateral bargaining with incomplete information. This research further evaluates the agents' bargaining performance in terms of the average payoff and settlement rate. The results show that agents using TD-based reinforcement learning are able to achieve good bargaining performance. This learning approach is sufficiently robust and convenient, hence it is suitable for online automated bargaining in electronic commerce. (C) 2007 Elsevier B.V. All rights reserved.
Keywords:
Markov decision process
Reinforcement learning
Temporal-difference learning
Risk-attitude
Online bargaining

Journal

Electronic Commerce Research and Applications cover
Electronic Commerce Research and Applications
IF:
6.3
Papers:
2.4K
Citations:
5.9K

Organization

N
National Tsing Hua University
Scholars:
1.6W
Papers: 1.4W
Citations: 1.7W
Ming Chuan University cover
Ming Chuan University
Scholars:
867
Papers: 1.2K
Citations: 781