1
Return

Bidirectional Feature-Aligned Motion Transformation for Efficient Dynamic Point Cloud Compression

delete2026-03-09
delete0
PRE
AI
X
Xingtao Wang
X
Xiandong Meng
L
Longguang Wang
T
Tiange Zhang
D
Debin Zhao
范晓鹏 (Xiaopeng Fan)
DOI:10.1109/tbc.2026.3668610delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Efficient dynamic point cloud compression (DPCC) critically depends on accurate motion estimation and compensation. However, the inherently irregular structure and substantial local variations of point clouds make this task highly challenging. Existing approaches typically rely on explicit motion estimation, whose encoded motion vectors often fail to capture complex dynamics and inadequately exploit temporal correlations. To address these limitations, we propose a <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Bidirectional Feature-aligned Motion Transformation (Bi-FMT)</b> framework that implicitly models motion in the feature space. Bi-FMT aligns features across both past and future frames to produce temporally consistent latent representations, which serve as predictive context in a conditional coding pipeline, forming a unified “Motion + Conditional” representation. Built upon this bidirectional feature alignment, we introduce a Cross-Transformer Refinement module (CTR) at the decoder side to adaptively refine locally aligned features. By modeling cross-frame dependencies with vector attention, CRT enhances local consistency and restores fine-grained spatial details that are often lost during motion alignment. Moreover, we design a <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Random Access (RA)</b> reference strategy that treats the bidirectionally aligned features as conditional context, enabling frame-level parallel compression and eliminating the sequential encoding. Extensive experiments demonstrate that Bi-FMT surpasses <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">D-DPCC</b> and <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">AdaDPCC</b> in both compression efficiency and runtime, achieving BD-Rate reductions of <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">20% (D1)</b> and <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">9.4% (D1)</b>, respectively. Our code is available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/dovedx/FMT_DPCC</uri>
Keywords:
Dynamic point clouds
compression
deep learning

Journal

IEEE Transactions on Broadcasting cover
IEEE Transactions on Broadcasting
IF:
4.8
Papers:
2.1K
Citations:
3.0K

Organization

S
Sun Yat-Sen University
Scholars:
7.9K
Papers: 2.1K
Citations: 0
P
Peng Cheng Laboratory
Scholars:
1.7K
Papers: 1.7K
Citations: 2.0K
H
Harbin Institute of Technology
Scholars:
1.1W
Papers: 3.8K
Citations: 8.5W
Cited Papers

Cited Papers

Citing Papers

Citing Papers