arrow
Return

Visual–Textual Information-Driven Tactile Data Generation Method

delete2026-06-11
delete0
PRE
AI
R
Rui Song
Y
Yang Xu
Z
Zhangzheng Tu
H
Hongzhou Wang
H
Hongchen Tan
卢湖川 (Huchuan Lu)
DOI:10.1109/TIP.2026.3700922delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Tactile data can enhance the environmental perception and interaction capabilities of intelligent agents, serving as a foundational component for the development of embodied intelligence. Despite its critical role, tactile data acquisition remains cost-prohibitive and labor-intensive, resulting in severe data scarcity. Cross-modal generation offers a promising solution by leveraging abundant visual and textual data. However, effectively aligning heterogeneous visual-textual modalities under data-scarce and sparsely-annotated conditions remains a significant challenge. To address these challenges, a visual-textual information-driven tactile data generation (VTTac) framework is proposed, which features three key innovations. First, a multi-granularity text enhancement strategy is introduced to mitigate annotation sparsity through hierarchical semantic enrichment. Second, a cascaded dual cross-attention mechanism is designed to ensure cross-modal alignment. Third, a condition adapter injects a low-frequency background prior, enabling the generative backbone to focus on high-frequency texture synthesis. Subsequently, a wavelet transform seamlessly fuses these synthesized details with the real background. Extensive evaluations across three datasets demonstrate that VTTac consistently outperforms representative baselines. Furthermore, downstream tasks validate the physical faithfulness of the synthesized data for material classification and semantic reasoning, and zero-shot experiments confirm generalization to unseen objects.
Keywords:
Cross-modal generation
tactile images
visual images
multi-granularity text
diffusion models

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

D
Dalian University of Technology
Scholars:
5.8W
Papers: 4.3W
Citations: 5.5W
J
Jilin University
Scholars:
8.5W
Papers: 5.5W
Citations: 8.9K