Return
HandDiff: Spatial-temporal diffusion model for hand pose forecasting
DOI:10.1016/j.inffus.2025.103647.png)
Abstract
En 中文
We propose a novel problem of forecasting future 3D hand pose from a short past sequence. The primary challenge in this task is accurately modeling the stochastic nature of future hand movements. To address this, we propose a diffusion-based hand pose forecasting model designed to generate accurate future hand poses by leveraging spatial–temporal information. Our model incorporates a Spatial-Temporal Attention Module (STAM) to capture correlations between hand joints and time points, and a Coarse Forecasting Module (CFM) to extract limited explicit guidance from the temporal dimension. These features condition the diffusion model to forecast plausible future hand poses. Due to the lack of suitable datasets, we also construct two large-scale datasets based on the existing hand-object interaction datasets HO-3D and HOI4D for benchmarking hand pose forecasting, covering both third-person and egocentric perspectives. Experimental results show that our method HandDiff significantly outperforms other state-of-the-art (SOTA) methods by 16.7% on the HO-3D dataset and 11.1% on the HOI4D dataset in terms of the mean per joint position error (MPJPE), respectively.
Journal
IF:
15.5
Papers:
4.1K
Citations:
2.7W

