Return
Mining user behavior patterns based on reinforcement learning algorithm to optimize service robot interaction strategy
G
Z
DOI:10.1051/meca/2026002.png)
Abstract
En 中文
In order to resolve the challenges of reduced adaptability and generalization stemming from complicated user behavior and sparse feedback in service robot interactions, this paper presents a reinforcement learning-based approach to user behavior pattern mining and policy optimization. This approach integrates Bayesian belief updates and automata learning with counterexamples to unify intent modeling and policy iteration: dynamic intent reasoning promotes multi-scale exploration under sparse rewards; and counterexample reasoning reconstructs the rewards function to bolster policy generalization. Experiments revealed that the method led to mean latency standard deviation of 0.031, a task completion rate was 78.14%, and behavior recognition accuracy 0.9, which can contribute to more capable policies and improve user satisfaction, while providing a reference for the design of complex interaction strategies.
Keywords:
Service robot interaction
reinforcement learning
user behavior pattern modeling
dynamic intent recognition
strategy adaptive optimization
Journal
M
IF:
1.2
Papers:
23
Citations:
0
