Return
GrambaTalk: generating 3D talking head with spatial graph and temporal mamba
X
F
X
J
W
X
DOI:10.1007/s00371-026-04657-x.png)
Abstract
En 中文
Animating 3D head meshes from speech is essential for virtual avatars, digital humans, and immersive VR/AR applications. However, existing methods typically perform spatial reasoning and temporal modeling independently, making it difficult to simultaneously achieve accurate lip articulation, temporal coherence, and geometrically consistent mesh deformation. We present GrambaTalk, a unified spatial–temporal framework for speech-driven 3D head animation. Our key idea is to jointly model speech dynamics and facial geometry throughout the entire generation pipeline by integrating graph-based spatial reasoning with bidirectional Mamba state-space modeling. Specifically, a hierarchical Chebyshev graph encoder learns geometry-aware mesh representations, while a bidirectional Mamba encoder captures long-range speech dynamics. The two modalities are coupled through a vertex-wise FiLM conditioning module and decoded by a dual-branch architecture that combines PCA-based global motion estimation with Graph–Mamba refinement for fine-grained facial articulation. By jointly modeling spatial geometry and temporal dynamics, GrambaTalk generates synchronized, temporally stable, and geometrically consistent facial animations. Experiments on the VOCASET and Multiface benchmarks demonstrate consistent improvements over state-of-the-art methods in both quantitative evaluation and perceptual user studies.
Keywords:
3D talking head
Graph convolutional network
Mamba
Journal
IF:
2.9
Papers:
4.5K
Citations:
6.5K

